vision-language-model

Tag

Cards List
#vision-language-model

XingChen-AGI/TeleOCR

Hugging Face Models Trending ↗ · 2026-08-14 Cached

TeleOCR is a lightweight open-source Vision-Language Model for document parsing, achieving state-of-the-art performance on both digital and camera-captured documents using techniques like Multi-node Consensus Voting.

0 favorites 0 likes
#vision-language-model

StarDoc-AI/TeleOCR

Hugging Face Models Trending ↗ · 2026-08-14 Cached

TeleOCR is a lightweight open-source Vision-Language Model designed for document parsing, unifying digital and camera-captured documents with state-of-the-art performance on benchmarks.

0 favorites 0 likes
#vision-language-model

@ma_sc_: I've been testing this on many other languages than the 14 officially supported and results have been truly surprising.…

X AI KOLs Following ↗ · 2026-08-12 Cached

A user shares surprising results testing Liquid AI's new LFM2.5-VL-3B vision-language model across many languages, noting strong visual capabilities but weaker instruction following; Liquid AI announces the model can read screens, documents, and ground objects to coordinates.

0 favorites 0 likes
#vision-language-model

LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

Hugging Face Blog ↗ · 2026-08-12 Cached

LiquidAI announces LFM2.5-VL-3B, an efficient vision-language model for edge hardware with improved screen understanding, grounding, multi-image input, and function calling, trained with 4x more vision data and post-training via SFT and RL.

0 favorites 0 likes
#vision-language-model

Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents

arXiv cs.AI ↗ · 2026-08-12 Cached

This paper investigates the compressibility of latent-space communication between vision-language model agents by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations. It demonstrates a 128x reduction in transmitted bytes with minimal accuracy loss, revealing that the dense communication channel is highly redundant.

0 favorites 0 likes
#vision-language-model

@MSFTResearch: Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning,…

X AI KOLs Timeline ↗ · 2026-08-11 Cached

Microsoft Research introduces CARE-X, a unified chest X-ray vision-language model that combines flexible reasoning, calibrated predictions, and tool-augmented measurement for clinically useful radiology interpretation.

0 favorites 0 likes
#vision-language-model

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper introduces VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model, but reports a negative visual-grounding result, raising architectural questions about NoPE layers and releasing code, benchmarks, and checkpoints.

0 favorites 0 likes
#vision-language-model

DocAtlas: Long-Document Understanding as Mutable-State Interaction

arXiv cs.CL ↗ · 2026-08-11 Cached

DocAtlas is a research system that frames long-document understanding as a mutable-state interaction process, using an external document harness with search, reading, note-taking, and review tools. It achieves state-of-the-art results on MMLongBench-Doc with GPT-5.4 and substantially improves compact VLM agents via end-to-end reinforcement learning.

0 favorites 0 likes
#vision-language-model

I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

Reddit r/LocalLLaMA ↗ · 2026-08-11

The author trained a 40M-parameter connector on 100K examples to give DeepSeek V4 Flash basic vision, freezing both the language model and MoonViT image encoder, demonstrating a low-cost approach to turning a text-only MoE into a basic VLM.

0 favorites 0 likes
#vision-language-model

omlab/VLX-Seek-1.5-10B · Hugging Face

Reddit r/LocalLLaMA ↗ · 2026-08-10 Cached

VLX-Seek-1.5-10B is an open-source 10B vision-language model from omlab, designed for fine-grained visual grounding in embodied scenarios like drones, robots, and surveillance, using region-reference localization instead of coordinate generation.

0 favorites 0 likes
#vision-language-model

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper introduces Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built by injecting structured document parsing capability into a reasoning-oriented multimodal model while preserving VQA ability. It studies parsing-centric SFT, mixed SFT, and DAPO-based parsing-centric RL to improve structured parsing performance.

0 favorites 0 likes
#vision-language-model

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

arXiv cs.AI ↗ · 2026-08-10 Cached

This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.

0 favorites 0 likes
#vision-language-model

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

Hugging Face Daily Papers ↗ · 2026-08-09 Cached

Presents VectraYX-Vision-1B, a sub-2B Spanish/LATAM cybersecurity vision-language model coupling a frozen SigLIP encoder with a Spanish decoder via an MLP, yet reports near-zero visual grounding despite functional pipelines, with open-source weights and remediation plans.

0 favorites 0 likes
#vision-language-model

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

EffectLearner is a semantic-reasoning-enhanced framework for real-world video object removal, combining a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser and introducing the EffectWorld dataset to handle complex object-induced effects.

0 favorites 0 likes
#vision-language-model

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

arXiv cs.CL ↗ · 2026-08-05 Cached

Introduces OncoTriad-QA, a patient-level benchmark integrating radiology, pathology, genomics, and clinical data for pan-cancer reasoning, along with OncoVLM, a reference multimodal model that outperforms existing medical LLMs after fine-tuning.

0 favorites 0 likes
#vision-language-model

@skalskip92: Qwen3.8-Max can be prompted with positive and negative boxes and use them to generate new detections super useful when …

X AI KOLs Timeline ↗ · 2026-08-04 Cached

SkalskiP highlights Qwen3.8-Max, a vision-language model for object detection that can be prompted with positive and negative boxes to generate detections, achieving 60-80% mAP with single or multiple prompts and performing well on diverse image types.

0 favorites 0 likes
#vision-language-model

Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

arXiv cs.CL ↗ · 2026-08-04 Cached

Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.

0 favorites 0 likes
#vision-language-model

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

PosterMELD is a template-conditioned multi-agent pipeline that converts academic papers into editable, print-ready posters, achieving 81.3% Print-Ready Rate with low cost, and is released with code and resources.

0 favorites 0 likes
#vision-language-model

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

Introduces AD-MCQ and DEFT-RLVR, a method for verifiable reasoning in autonomous driving VLMs that defers future trajectory exposure to post-decision verification, improving reasoning faithfulness while reducing hallucinations.

0 favorites 0 likes
#vision-language-model

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

Roomer is a reflective repair framework that identifies and fixes local violations in 3D indoor layouts, using a vision-language model planner and deterministic solver, with new benchmarks Roomer-CC and Roomer-Eval.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback