Tag
This paper presents the ICDAR 2026 Competition on Information Extraction from ALD/E Scientific Figures, introducing the Sci-ImageMiner benchmark with four complementary tasks. Results show SOTA multimodal models perform well on classification and summarization but struggle with data extraction and scientific reasoning, especially visual question answering.
The AMD AI DevMaster Hackathon, running July 10–August 6, 2026, features three tracks (multimodal AI, agentic AI, physical AI) and a $30,000 prize pool, with PyTorch as the official framework on ROCm-enabled AMD GPUs.
Mira Murati announces Thinking Machines Lab, a new initiative focused on multimodal AI, custom models, and open science, along with a tool called Tinker for training open-weight models.
This paper introduces Audex, a unified audio-text LLM from NVIDIA that achieves state-of-the-art performance across multiple audio and speech tasks while preserving strong text reasoning capabilities without regression.
Victoria Lin's Stanford CS25 lecture discusses the paradigm shift from language models to native multimodal intelligence, covering tokenization approaches and comparing models like Chameleon and Transfusion.
The tweet highlights the remarkably low cost and fast speed of multimodal AI agents like Nano Banana 2 Lite (~3 cents/image, ~4 seconds) and Gemini Omni Flash ($0.10/sec video with native audio and editing), enabling affordable looping for agents.
A summary of Oriol Vinyals' discussion on Google's Gemini models, world models, multimodal AI, agents, and challenges like continual learning and true innovation.
This article, based on the sharing of researcher Victoria Lin, systematically reviews the mainstream technical approaches of native multimodal large models (Chameleon, Transfusion, MOT) and their pros and cons. It points out that multimodal AI is still in the early exploration stage, with open problems such as gaps in scaling laws, inconsistency between image understanding and generation encoding, and connection with the physical world.
The article discusses the trend of generative AI products evolving from isolated single-capability models into integrated workflow ecosystems that bundle music, video, voice, and editing tools, potentially reducing workflow fragmentation for creators despite trade-offs in model quality.
Magma is an open-source repository from Microsoft Research for building multimodal AI agents that integrate vision, language, and action, providing model links, inference examples, training instructions, and demos.
Mira Murati's startup Thinking Machines Lab previews a new 'interaction model' that natively understands continuous human communication, aiming to keep humans in the loop as AI advances towards superintelligence.
This paper presents a structured framework for benchmarking generative, multimodal, and agentic AI in healthcare, addressing the gap between high benchmark scores and real-world clinical reliability, safety, and relevance.
This article introduces EmoS, a high-fidelity multimodal benchmark designed for fine-grained streaming emotional understanding, addressing limitations in ecological validity and labeling reliability found in existing datasets.
This paper introduces UniCharacter, a two-stage training framework for Customized Multimodal Role-Play (CMRP) that enables unified customization of persona, dialogue style, and visual identity. It presents the RoleScape-20 dataset and demonstrates that the model can achieve coherent cross-modal generation with minimal data.
This paper introduces SenseNova-U1, a unified multimodal architecture that integrates understanding and generation tasks, releasing two variants (8B and 30B) that perform competitively in both perception and image synthesis.
The article discusses how multimodal AI models like GPT-4o and Claude 3.5 Sonnet are overcoming text-only bottlenecks by enabling visual debugging, audio-to-data conversion, and enhanced RAG systems.
This paper introduces Structured Role-Aware Policy Optimization (SRPO), a method that improves multimodal reasoning in Large Vision-Language Models by assigning token-level credit based on distinct perception and reasoning roles within reinforcement learning frameworks.
This paper introduces Positive-and-Negative Decoding (PND), a training-free inference framework that reduces object hallucination in Vision-Language Models by contrasting positive visual evidence with negative counterfactuals during decoding.
Qwen-Image-2.0 is a new image generation foundation model that unifies high-fidelity synthesis and precise editing using Qwen3-VL and a Multimodal Diffusion Transformer. It excels in text-rich content, multilingual typography, and photorealistic generation.
The user shares an impressive example of a 3D biological structure page generated using GPT Image 2 and Gemini 3.1 Pro, noting its potential for AI education and expressing intent to replicate it.