Tag
TANGO is a vision-language-action framework for humanoid robots to navigate cluttered environments using simulated training data, demonstrating state-of-the-art performance and zero-shot real-world deployment.
CARDEA is an AI model designed to provide auditable reasoning for coronary angiography interpretation using Chain-of-Box to indicate image regions, enhancing trust in clinical practice.
A shared app for exploring and understanding vision language model evaluations, referencing a thread analyzing popular vision benchmarks.
Jina-OCR-v1 is an efficient end-to-end document parsing model that uses speculative decoding and dense verifiable rewards to achieve high accuracy and speed on low-budget GPUs, scoring 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench.
A coding agent guided by vision-language models generates editable layered designs by synthesizing visual assets and refining HTML/CSS layouts, enabling post-editing with precise control.
WorldReward introduces a vision-language reward model for camera-conditioned world models that unifies action-consistency and visual-quality evaluation through chunk decomposition and preference aggregation, outperforming existing methods like GPT-5.5 on benchmarks.
An announcement that open-source resources, including code, RL environment, and dataset, will be released for using reinforcement learning on a 4B parameter vision-language model to excel at GeoGuesser, enabling full reproduction.
SCAFFOLD is a large-scale dataset of computer science research figures paired with captions, context, QA, and Chain-of-Thought traces, aimed at improving vision-language model understanding of diagrams in CS papers.
This research paper introduces a parametric multimodal user memory system that improves AI agents' ability to recall users by integrating perceptual data like voice and appearance, using vision-language models and dedicated encoders to surpass text-based methods.
SimLoss uses embedding-space contrastive supervision to enable single-pass fine-grained image captioning that matches multi-stage quality at much lower latency, with variants like SimLoss FFT achieving high precision.
FoldingAgent is an agentic framework that uses vision-language models and specialized tools to infer parametric origami folding programs from demonstration videos via sequential reasoning and physical verification.
Cohere introduces Parse, a cost-effective vision language model for processing large volumes of enterprise documents into structured data, offering high performance and scalability for use cases like RAG and document indexing.
The paper introduces Image Bundle Composition (IBC) to shift image retrieval from atomic matching to dynamic composition of cohesive image bundles, addressing limitations in traditional approaches. It proposes a benchmark dataset IBCBench and an agentic framework BundleWeaver that leverages LLMs and VLMs for relational composition.
The article describes the process of fine-tuning a 450 million parameter vision-language model using 50,000 browser screenshots, showing progress from initial to current stages.
The article introduces llava-medical-8B-clip-vit-stage2, a specialized vision-language model fine-tuned for healthcare, enabling AI to interpret medical images like X-rays and MRIs and answer related questions.
QwenMix-3.7 is a weight-interpolation merge of the Qwen3.6-27B and Qwen3.8-27B AI models, created using per-row norm preservation at t=0.5.
SemComp-Bench introduces a benchmark for evaluating semantic task completion in video generation, using a curated dataset and a vision-language model-based evaluation protocol to assess outcome achievement and generation reliability.
ConceptFormer learns continuous latent concept representations for visual document retrieval, bridging visual evidence and semantic relevance without text intermediates, achieving significant improvements over baselines.
MOSS-VL is an open vision-language model family designed for real-time interaction, using gated cross-attention to process vision during generation and achieving top performance in streaming benchmarks among open-source models.
LiquidAI released LFM2.5-VL-3B, a 3.1B local vision-language model that outperforms Gemma-4 E4B and demonstrates a significant jump in screen understanding, making on-device AI more practical.