Tag
A tweet suggests using Gemini 3.8 Flash for multimodal understanding and references a visual test comparing AI models' ability to name people from a drawing.
FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.
A new stealth multimodal AI model named Space Bunny has been released in OpenCode, available for free trial, with the releasing company unknown.
Xiaomi's MiMo V2.6 Pro is now the top-ranked global open-source AI model on Artificial Analysis, nearly matching GPT-5.6 Sol in intelligence while offering cost-effective pricing for real-world applications.
Jev-Omni is the first open-weight model to extend typed decisions to multimodal, supporting text, images, audio, and video, and directly returning the probability distribution of options without generating explanations.
The paper introduces ReFigBench, a benchmark for evaluating coding agents in reconstructing scientific figures into editable PowerPoint slides, showing that harnesses significantly affect performance across model families.
PixVerse R2 introduces a unified scaling architecture for real-time audiovisual world models, leveraging block-sparse attention and continuous pretraining to enhance video generation and interactive control.
dots3-note Preview is a multimodal Mixture-of-Experts AI model with 280B total parameters and 16B activated, supporting up to 512K token context, released as an open-weight model on Hugging Face.
Xiaomi has open-sourced the MiMo-V2.6 series, introducing two omnimodal AI models, Pro and Flash, that perform competitively with top proprietary models and are available for various applications.
MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.
This paper proposes RRDrive, a risk-aware occupancy representation for register-based end-to-end autonomous driving, which improves trajectory prediction and selection in complex scenes, achieving notable performance gains.
This paper introduces Omni Demand Understanding (ODU), a benchmark to evaluate how well multimodal AI models infer user demands from complex audio-visual interactions, revealing significant performance gaps in current models.
VisPath introduces a visual-intent-guided path reasoning framework for multimodal knowledge graph question answering, achieving significant improvements over baselines on a new benchmark, VisPath-Bench, and existing datasets.
The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.
Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.
A user shares their positive experience switching to ds v4.1 flash from Astra, highlighting its multimodal capabilities and fast speed.
Lillian Ma hosted a large multimodal event in the Bay Area at a museum, showcasing products like Inference, MCP, and AgentBox with over 200 attendees from partner companies.
The paper proposes a discriminative adaptation of SpeechLLMs for emotion recognition, improving performance and interpretability by using a linear classification head on the hidden state of the final prompt token, which removes hallucinations and enhances analysis of emotion directions.
Qwen3.8-Omni-Flash is an omnimodal AI model with a 1M-token context window, supporting text, image, audio, and video inputs, and achieving performance comparable to or better than Gemini 3.8 Flash, now available on the Qianwen AI Platform.
Alibaba has released the Qwen 3.8 Omni Flash AI model, which likely features multimodal capabilities and is optimized for speed.