Tag
LiquidAI announces LFM2.5-VL-3B, an efficient vision-language model for edge hardware with improved screen understanding, grounding, multi-image input, and function calling, trained with 4x more vision data and post-training via SFT and RL.
This paper investigates the compressibility of latent-space communication between vision-language model agents by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations. It demonstrates a 128x reduction in transmitted bytes with minimal accuracy loss, revealing that the dense communication channel is highly redundant.
Microsoft Research introduces CARE-X, a unified chest X-ray vision-language model that combines flexible reasoning, calibrated predictions, and tool-augmented measurement for clinically useful radiology interpretation.
This paper introduces VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model, but reports a negative visual-grounding result, raising architectural questions about NoPE layers and releasing code, benchmarks, and checkpoints.
DocAtlas is a research system that frames long-document understanding as a mutable-state interaction process, using an external document harness with search, reading, note-taking, and review tools. It achieves state-of-the-art results on MMLongBench-Doc with GPT-5.4 and substantially improves compact VLM agents via end-to-end reinforcement learning.
The author trained a 40M-parameter connector on 100K examples to give DeepSeek V4 Flash basic vision, freezing both the language model and MoonViT image encoder, demonstrating a low-cost approach to turning a text-only MoE into a basic VLM.
VLX-Seek-1.5-10B is an open-source 10B vision-language model from omlab, designed for fine-grained visual grounding in embodied scenarios like drones, robots, and surveillance, using region-reference localization instead of coordinate generation.
This paper introduces Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built by injecting structured document parsing capability into a reasoning-oriented multimodal model while preserving VQA ability. It studies parsing-centric SFT, mixed SFT, and DAPO-based parsing-centric RL to improve structured parsing performance.
This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.
Presents VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model coupling a frozen SigLIP encoder with a Spanish decoder via an MLP, yet reports near-zero visual grounding despite functional pipelines, with open-source weights and remediation plans.
EffectLearner is a semantic-reasoning-enhanced framework for real-world video object removal, combining a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser and introducing the EffectWorld dataset to handle complex object-induced effects.
Introduces OncoTriad-QA, a patient-level benchmark integrating radiology, pathology, genomics, and clinical data for pan-cancer reasoning, along with OncoVLM, a reference multimodal model that outperforms existing medical LLMs after fine-tuning.
SkalskiP highlights Qwen3.8-Max, a vision-language model for object detection that can be prompted with positive and negative boxes to generate detections, achieving 60-80% mAP with single or multiple prompts and performing well on diverse image types.
Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.
PosterMELD is a template-conditioned multi-agent pipeline that converts academic papers into editable, print-ready posters, achieving 81.3% Print-Ready Rate with low cost, and is released with code and resources.
Introduces AD-MCQ and DEFT-RLVR, a method for verifiable reasoning in autonomous driving VLMs that defers future trajectory exposure to post-decision verification, improving reasoning faithfulness while reducing hallucinations.
Roomer is a reflective repair framework that identifies and fixes local violations in 3D indoor layouts, using a vision-language model planner and deterministic solver, with new benchmarks Roomer-CC and Roomer-Eval.
This paper performs a forensic reproducibility audit of a radiology vision-language model benchmark, finding divergences between the intended protocol and released artifacts that invalidate the original claims. The authors propose a benchmark contract to expose such failure classes.
Google unveiled Gemini Robotics 2.0, a trio of AI sub-models for robots that improve dexterity, real-time video understanding, and robot collaboration, with the embodied reasoning model publicly available to developers.
NVIDIA unveiled Ising Calibration 1.5, an open-source vision language model that fully automates quantum computer calibration with enhanced in-context learning and improved performance, now deployable on a single GPU.