vision-language-model

Tag

Cards List
#vision-language-model

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

Hugging Face Blog · 20m ago Cached

LiquidAI announces LFM2.5-VL-3B, an efficient vision-language model for edge hardware with improved screen understanding, grounding, multi-image input, and function calling, trained with 4x more vision data and post-training via SFT and RL.

0 favorites 0 likes
#vision-language-model

Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents

arXiv cs.AI · 10h ago Cached

This paper investigates the compressibility of latent-space communication between vision-language model agents by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations. It demonstrates a 128x reduction in transmitted bytes with minimal accuracy loss, revealing that the dense communication channel is highly redundant.

0 favorites 0 likes
#vision-language-model

@MSFTResearch: Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning,…

X AI KOLs Timeline · 22h ago Cached

Microsoft Research introduces CARE-X, a unified chest X-ray vision-language model that combines flexible reasoning, calibrated predictions, and tool-augmented measurement for clinically useful radiology interpretation.

0 favorites 0 likes
#vision-language-model

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

arXiv cs.CL · yesterday Cached

This paper introduces VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model, but reports a negative visual-grounding result, raising architectural questions about NoPE layers and releasing code, benchmarks, and checkpoints.

0 favorites 0 likes
#vision-language-model

DocAtlas: Long-Document Understanding as Mutable-State Interaction

arXiv cs.CL · yesterday Cached

DocAtlas is a research system that frames long-document understanding as a mutable-state interaction process, using an external document harness with search, reading, note-taking, and review tools. It achieves state-of-the-art results on MMLongBench-Doc with GPT-5.4 and substantially improves compact VLM agents via end-to-end reinforcement learning.

0 favorites 0 likes
#vision-language-model

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Reddit r/LocalLLaMA · yesterday

The author trained a 40M-parameter connector on 100K examples to give DeepSeek V4 Flash basic vision, freezing both the language model and MoonViT image encoder, demonstrating a low-cost approach to turning a text-only MoE into a basic VLM.

0 favorites 0 likes
#vision-language-model

omlab/VLX-Seek-1.5-10B · Hugging Face

Reddit r/LocalLLaMA · 2d ago Cached

VLX-Seek-1.5-10B is an open-source 10B vision-language model from omlab, designed for fine-grained visual grounding in embodied scenarios like drones, robots, and surveillance, using region-reference localization instead of coordinate generation.

0 favorites 0 likes
#vision-language-model

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

arXiv cs.CL · 2d ago Cached

This paper introduces Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built by injecting structured document parsing capability into a reasoning-oriented multimodal model while preserving VQA ability. It studies parsing-centric SFT, mixed SFT, and DAPO-based parsing-centric RL to improve structured parsing performance.

0 favorites 0 likes
#vision-language-model

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

arXiv cs.AI · 2d ago Cached

This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.

0 favorites 0 likes
#vision-language-model

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

Hugging Face Daily Papers · 3d ago Cached

Presents VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model coupling a frozen SigLIP encoder with a Spanish decoder via an MLP, yet reports near-zero visual grounding despite functional pipelines, with open-source weights and remediation plans.

0 favorites 0 likes
#vision-language-model

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Hugging Face Daily Papers · 6d ago Cached

EffectLearner is a semantic-reasoning-enhanced framework for real-world video object removal, combining a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser and introducing the EffectWorld dataset to handle complex object-induced effects.

0 favorites 0 likes
#vision-language-model

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

arXiv cs.CL · 2026-08-05 Cached

Introduces OncoTriad-QA, a patient-level benchmark integrating radiology, pathology, genomics, and clinical data for pan-cancer reasoning, along with OncoVLM, a reference multimodal model that outperforms existing medical LLMs after fine-tuning.

0 favorites 0 likes
#vision-language-model

@skalskip92: Qwen3.8-Max can be prompted with positive and negative boxes and use them to generate new detections super useful when …

X AI KOLs Timeline · 2026-08-04 Cached

SkalskiP highlights Qwen3.8-Max, a vision-language model for object detection that can be prompted with positive and negative boxes to generate detections, achieving 60-80% mAP with single or multiple prompts and performing well on diverse image types.

0 favorites 0 likes
#vision-language-model

Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

arXiv cs.CL · 2026-08-04 Cached

Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.

0 favorites 0 likes
#vision-language-model

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

Hugging Face Daily Papers · 2026-08-03 Cached

PosterMELD is a template-conditioned multi-agent pipeline that converts academic papers into editable, print-ready posters, achieving 81.3% Print-Ready Rate with low cost, and is released with code and resources.

0 favorites 0 likes
#vision-language-model

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Hugging Face Daily Papers · 2026-08-03 Cached

Introduces AD-MCQ and DEFT-RLVR, a method for verifiable reasoning in autonomous driving VLMs that defers future trajectory exposure to post-decision verification, improving reasoning faithfulness while reducing hallucinations.

0 favorites 0 likes
#vision-language-model

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

Hugging Face Daily Papers · 2026-08-03 Cached

Roomer is a reflective repair framework that identifies and fixes local violations in 3D indoor layouts, using a vision-language model planner and deterministic solver, with new benchmarks Roomer-CC and Roomer-Eval.

0 favorites 0 likes
#vision-language-model

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

arXiv cs.AI · 2026-07-31 Cached

This paper performs a forensic reproducibility audit of a radiology vision-language model benchmark, finding divergences between the intended protocol and released artifacts that invalidate the original claims. The authors propose a benchmark contract to expose such failure classes.

0 favorites 0 likes
#vision-language-model

Google reveals Gemini Robotics 2.0, promising improved dexterity and safety

Ars Technica · 2026-07-30 Cached

Google unveiled Gemini Robotics 2.0, a trio of AI sub-models for robots that improve dexterity, real-time video understanding, and robot collaboration, with the embodied reasoning model publicly available to developers.

0 favorites 0 likes
#vision-language-model

NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

Reddit r/singularity · 2026-07-27 Cached

NVIDIA unveiled Ising Calibration 1.5, an open-source vision language model that fully automates quantum computer calibration with enhanced in-context learning and improved performance, now deployable on a single GPU.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback