vision-language

Tag

Cards List
#vision-language

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

arXiv cs.CL · 2026-07-29 Cached

This paper presents VLD-RAG, an agentic multimodal retrieval-augmented generation framework for question answering over long, visually-rich documents. It uses a page-preserving index and a verifier-guided agent workflow to improve cross-page evidence retrieval and reasoning, outperforming prior vision-based baselines on benchmarks like LongDocURL and MMLongBench-Doc.

0 favorites 0 likes
#vision-language

RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

arXiv cs.AI · 2026-07-29 Cached

RRS-10K is a benchmark dataset for evaluating vision-language models on rare remote sensing image interpretation, containing over 10,000 military-related images and multiple task formats. Evaluation of 52 models reveals moderate zero-shot performance and weaknesses in visual grounding and complex reasoning.

0 favorites 0 likes
#vision-language

ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop

arXiv cs.AI · 2026-07-29 Cached

ProcAgent is a fully on-device, agentic, vision-based procedural assistant that uses a propose-and-verify architecture for real-time adaptive guidance on an NVIDIA Jetson AGX Orin. It supports both reactive and proactive modes with human-in-the-loop confirmation, achieving responsive interaction and positive user study ratings.

0 favorites 0 likes
#vision-language

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Hugging Face Daily Papers · 2026-07-28 Cached

This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.

0 favorites 0 likes
#vision-language

Scaling Native Multimodal Pre-Training From Scratch

Hugging Face Daily Papers · 2026-07-24 Cached

This paper investigates the scaling properties of native multimodal pre-training, deriving compute-optimal model sizes and token counts under a fixed budget, and revealing distinct scaling behaviors for language and multimodal objectives.

0 favorites 0 likes
#vision-language

Visual Contrastive Self-Distillation

Hugging Face Daily Papers · 2026-07-23 Cached

VCSD removes the need for external teachers, privileged answers, or visual evidence in on-policy self-distillation by using content-erased control images to produce contrastive signals, consistently outperforming existing methods on vision-language benchmarks.

0 favorites 0 likes
#vision-language

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Hugging Face Daily Papers · 2026-07-23 Cached

Proposes ProVisE, a benchmark-agnostic framework to evaluate spatial cognition in image-generation models using pixel-space outputs, and introduces SpatialGen-Bench for unified evaluation across 14 spatial subtasks.

0 favorites 0 likes
#vision-language

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

arXiv cs.CL · 2026-07-22 Cached

Fusion Embedding introduces a family of models that add audio to a frozen vision-language embedding backbone, enabling a unified space for text, image, video, and audio retrieval. The models train only lightweight adapters and achieve audio-image retrieval without paired audio-visual data.

0 favorites 0 likes
#vision-language

Mage (GitHub Repo)

TLDR AI · 2026-07-22 Cached

Microsoft releases Mage, a family of lightweight 4B-parameter multimodal models for visual understanding and generation, including Mage-VL for image/video understanding and Mage-Flow for text-to-image generation and editing, designed for research and deployment on modest hardware.

0 favorites 0 likes
#vision-language

baseten/GLM-5.2-Vision-NVFP4

Hugging Face Models Trending · 2026-07-20 Cached

Baseten releases GLM-5.2-Vision, a vision-language model that adds MoonViT vision encoder to GLM-5.2 via a trained PatchMerger projector, keeping the text backbone and vision tower frozen. The model is quantized to NVFP4 for efficient inference on Blackwell hardware.

0 favorites 0 likes
#vision-language

Dataset Distillation by Influence Matching

Hugging Face Daily Papers · 2026-07-18 Cached

This paper introduces Influence Matching (Inf-Match), a dataset distillation method that aligns the final training outcome by learning a compact synthetic set whose effect on converged parameters matches that of the full dataset. It achieves state-of-the-art accuracy on classification benchmarks and outperforms strong baselines on vision-language distillation tasks.

0 favorites 0 likes
#vision-language

OpenMOSS-Team/MOSS-VL-Realtime

Hugging Face Models Trending · 2026-07-14 Cached

MOSS-VL-Realtime is a realtime streaming vision-language model that processes continuous video frames, supports interruptible interaction, proactive silence, and dynamic correction, with timestamp-aware encoding and a 256K context window.

0 favorites 0 likes
#vision-language

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

Hugging Face Daily Papers · 2026-07-13 Cached

Introduces SVR-R1, a multi-turn reinforcement learning framework that uses the model's own verification as a learning signal for multi-modal reasoning, achieving significant accuracy improvements over standard GRPO baselines on vision-language reasoning benchmarks.

0 favorites 0 likes
#vision-language

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

arXiv cs.AI · 2026-07-07 Cached

BRAID is a framework that formulates interleaved text-image-text reasoning as a unified Markov decision process, enabling joint optimization of textual and visual generation via reinforcement learning with a VLM judge providing dense turn-level feedback.

0 favorites 0 likes
#vision-language

iFLYTEK-Embodied-Omni Technical Report

arXiv cs.AI · 2026-07-07 Cached

This technical report presents iFLYTEK-Embodied-Omni, a unified multimodal foundation model that jointly models vision, language, and action for embodied agents, using a brain-cerebellum collaboration architecture and a four-stage training strategy.

0 favorites 0 likes
#vision-language

@Prince_Canuma: mlx-vlm v0.6.4 is here! Get started today: > uv pip install -U mlx-vlm 5 new model families — MiniMax M3, Kimi K2.5, Un…

X AI KOLs Timeline · 2026-07-06 Cached

mlx-vlm v0.6.4 is released with support for 5 new model families, TTS/STT endpoints, and significant performance improvements including TurboQuant and continuous batching.

0 favorites 0 likes
#vision-language

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism

arXiv cs.AI · 2026-07-01 Cached

Introduces PEC-CIR, a training-free zero-shot composed image retrieval framework that uses a Planner-Executor-Critic architecture to improve retrieval precision by structuring query construction as a multi-stage reasoning pipeline.

0 favorites 0 likes
#vision-language

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

arXiv cs.AI · 2026-06-30 Cached

IMCBench is a new benchmark for evaluating multimodal LLMs on image-grounded medical conversations, pairing clinical images with synthetic patient profiles. Evaluations across safety, accuracy, and uncertainty show that even strong models like Claude Opus 4.6 have safety issues, highlighting the need for multi-dimensional evaluation.

0 favorites 0 likes
#vision-language

Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature

Hugging Face Daily Papers · 2026-06-29 Cached

Introduces MatMMExtract pipeline to decompose compound scientific figures into panels and annotate them using LLMs, creating the MatSciFig dataset of over 390,000 image-text pairs for vision-language learning in materials science.

0 favorites 0 likes
#vision-language

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Hugging Face Daily Papers · 2026-06-28 Cached

This paper proposes Rank-Aware Hyperbolic Alignment (RAHA), a method for vision-language dataset distillation that leverages hyperbolic geometry and alignment capacity control to efficiently compress large image-text datasets into high-quality synthetic pairs.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback