Tag
GAS introduces a generation-guided training framework that improves visual understanding in multimodal models by using auxiliary generation tasks with no inference overhead, via a decoupled mixture-of-transformers architecture.
Kimi K3, a new AI model with 2.8 trillion parameters and 1 million context length, has been released on web and app, featuring leading capabilities in coding, agentic tasks, reasoning, vision, and agent swarms.
Jerry Liu reports updated results for Mistral OCR on ParseBench, showing it outperforms GPT-5.5 and trails only Gemini 3.1 Pro, with strong performance on content faithfulness and semantic formatting.
The Visual Aesthetic Benchmark (VAB) evaluates multimodal models' ability to judge aesthetics through comparative selection, revealing significant gaps versus human experts and showing that fine-tuning on expert examples improves accuracy.
This paper introduces UNO, an Understanding-Oriented Post-Training framework that uses comprehension tasks as supervisory signals to enhance image generation and editing in unified multimodal models.
A comprehensive survey examining image classification into high-level and abstract categories, clarifying the tacit understanding of high-level semantics in computer vision through multidisciplinary analysis of commonsense, emotional, aesthetic, and interpretative semantics. The paper identifies persistent challenges in abstract concept image classification and emphasizes the importance of hybrid AI systems for addressing complex visual reasoning tasks.