Tag
ArGuard is a shared task focused on detecting harmful content in Arabic memes and LLM prompts, highlighting challenges in fine-grained classification and releasing datasets for further research.
Gemini 3.8 Live with Live Avatar brings real-time visual presence to conversational AI, allowing enterprises to create more natural and interactive digital experiences with dynamic avatars and multilingual support.
A tweet suggests using Gemini 3.8 Flash for multimodal understanding and references a visual test comparing AI models' ability to name people from a drawing.
FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.
A new stealth multimodal AI model named Space Bunny has been released in OpenCode, available for free trial, with the releasing company unknown.
Xiaomi's MiMo V2.6 Pro is now the top-ranked global open-source AI model on Artificial Analysis, nearly matching GPT-5.6 Sol in intelligence while offering cost-effective pricing for real-world applications.
Jev-Omni is the first open-weight model to extend typed decisions to multimodal, supporting text, images, audio, and video, and directly returning the probability distribution of options without generating explanations.
The paper introduces ReFigBench, a benchmark for evaluating coding agents in reconstructing scientific figures into editable PowerPoint slides, showing that harnesses significantly affect performance across model families.
PixVerse R2 introduces a unified scaling architecture for real-time audiovisual world models, leveraging block-sparse attention and continuous pretraining to enhance video generation and interactive control.
dots3-note Preview is a multimodal Mixture-of-Experts AI model with 280B total parameters and 16B activated, supporting up to 512K token context, released as an open-weight model on Hugging Face.
Xiaomi has open-sourced the MiMo-V2.6 series, introducing two omnimodal AI models, Pro and Flash, that perform competitively with top proprietary models and are available for various applications.
MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.
This paper proposes RRDrive, a risk-aware occupancy representation for register-based end-to-end autonomous driving, which improves trajectory prediction and selection in complex scenes, achieving notable performance gains.
This paper introduces Omni Demand Understanding (ODU), a benchmark to evaluate how well multimodal AI models infer user demands from complex audio-visual interactions, revealing significant performance gaps in current models.
VisPath introduces a visual-intent-guided path reasoning framework for multimodal knowledge graph question answering, achieving significant improvements over baselines on a new benchmark, VisPath-Bench, and existing datasets.
The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.
Jev-Omni is a multimodal decision classifier built on Gemma 4 12B IT, fine-tuned to handle text, images, audio, and video, providing probability scores for options based on questions.
Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.
A user shares their positive experience switching to ds v4.1 flash from Astra, highlighting its multimodal capabilities and fast speed.
Lillian Ma hosted a large multimodal event in the Bay Area at a museum, showcasing products like Inference, MCP, and AgentBox with over 200 attendees from partner companies.