vision-language-model

Tag

Cards List
#vision-language-model

@OpenBMB: Thanks to @_akhaliq for contributing MiniCPM-V 4.6 Hugging Face demo, which allowed us to test the gradio.Server featur…

X AI KOLs Following · 2026-05-23 Cached

OpenBMB thanks @_akhaliq for contributing a Hugging Face demo for MiniCPM-V 4.6, using Gradio server for flexible frontend customization.

0 favorites 0 likes
#vision-language-model

NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction (self-hostable) [P]

Reddit r/MachineLearning · 2026-05-22

Numind released NuExtract3, a 4B open-weight vision-language model based on Qwen3.5-4B, designed for converting document images to Markdown, OCR, and structured data extraction. It is Apache-2.0 licensed and self-hostable with quantized versions for low VRAM.

0 favorites 0 likes
#vision-language-model

AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education

arXiv cs.AI · 2026-05-22 Cached

This paper investigates using vision-language models to assess nursing competency from egocentric video during simulation, finding that recognition accuracy inversely relates to competency level, suggesting a pedagogically informative signal.

0 favorites 0 likes
#vision-language-model

Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task

arXiv cs.CL · 2026-05-21 Cached

University of Florida Gators submission to the AmericasNLP 2026 shared task on cultural image captioning for Indigenous languages, using a two-stage pipeline with Qwen2.5-VL for Spanish captioning and retrieval-augmented Gemini 2.5 Flash for target-language translation, achieving significant improvements over the baseline.

0 favorites 0 likes
#vision-language-model

SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents

arXiv cs.AI · 2026-05-20 Cached

SimGym is a framework that simulates A/B tests on e-commerce storefronts using vision-language model agents, reducing experimental cycles from weeks to under an hour while achieving 77% directional alignment with real buyer behavior.

0 favorites 0 likes
#vision-language-model

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

Hugging Face Daily Papers · 2026-05-20 Cached

AutoRubric-T2I automatically generates and selects explicit rubrics to guide Vision-Language Model judges for text-to-image generation, achieving high-quality reward signals with minimal human annotation and improving generation quality in downstream tasks.

0 favorites 0 likes
#vision-language-model

Aurora: Unified Video Editing with a Tool-Using Agent

Hugging Face Daily Papers · 2026-05-18 Cached

Aurora is an agentic video editing framework that pairs a tool-augmented vision-language model agent with a diffusion transformer to automatically resolve textual and visual underspecification in user requests, enabling unified video editing tasks like replacement, removal, style transfer, and reference-driven insertion.

0 favorites 0 likes
#vision-language-model

PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control

Hugging Face Daily Papers · 2026-05-15 Cached

This paper introduces PAGER, a topology-aware agent that bridges the semantic-execution gap in point-precise GUI control, achieving 4.1x higher task success than baselines on the new PAGE Bench.

0 favorites 0 likes
#vision-language-model

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

Hugging Face Daily Papers · 2026-05-13 Cached

RoboEvolve is a framework that co-evolves a VLM planner and VGM simulator for robotic manipulation, achieving data efficiency with only 500 unlabeled seed images and robust continual learning.

0 favorites 0 likes
#vision-language-model

Can a Language Model Paint?

Hacker News Top · 2026-05-12 Cached

The author explores whether language models can create art through an iterative painting process rather than one-shot generation, building an app that uses a vision-language model to apply strokes one at a time. The experiment highlights the fragility of LLM-generated artefacts and reflects on artistic sincerity.

0 favorites 0 likes
#vision-language-model

MiniCPM-V 4.6

Product Hunt · 2026-05-12

MiniCPM-V 4.6 is an ultra-efficient 1.3B vision-language model optimized for mobile devices.

0 favorites 0 likes
#vision-language-model

@Prince_Canuma: Congratulations to @OpenBMB on the launch of MiniCPM-V 4.6! We have Day-0 support for it on MLX-VLM h/t Magic Yang Runs…

X AI KOLs Timeline · 2026-05-11 Cached

OpenBMB has launched the MiniCPM-V 4.6 vision language model, which features immediate day-0 support on the MLX-VLM package for high-speed inference on Apple Silicon Macs.

0 favorites 0 likes
#vision-language-model

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Hugging Face Daily Papers · 2026-05-11 Cached

RoboMemArena introduces a large-scale benchmark for evaluating robotic memory across 26 complex tasks with real-world validation, alongside PrediMem, a dual-system vision-language-action model that improves memory management through predictive coding.

0 favorites 0 likes
#vision-language-model

@tom_doerr: Converts images and PDFs to Markdown without OCR https://github.com/NanoNets/docext

X AI KOLs Timeline · 2026-05-08 Cached

docext is an on-premises toolkit that converts images and PDFs to markdown without OCR, leveraging vision-language models. It also introduces Nanonets-OCR-s, a compact 3B parameter model for efficient image-to-markdown conversion.

0 favorites 0 likes
#vision-language-model

MolmoAct 2

Product Hunt · 2026-05-05

MolmoAct 2 is an open robotics model that reasons in 3D space before taking actions, developed by the Allen Institute for Artificial Intelligence.

0 favorites 0 likes
#vision-language-model

@techNmak: A lightweight VLM that beats the giants at OCR. (1.7B parameters, SOTA on OmniDocBench) dots. ocr is a new multilingual…

X AI KOLs Timeline · 2026-04-20 Cached

dots.ocr is a new lightweight 1.7B parameter multilingual vision-language model that achieves state-of-the-art performance on OmniDocBench, outperforming much larger models (72B+) at document parsing and OCR tasks.

0 favorites 0 likes
#vision-language-model

SGOCR: A Spatially-Grounded OCR-focused Pipeline & V1 Dataset [P]

Reddit r/MachineLearning · 2026-04-20

SGOCR is an open-source dataset pipeline for generating spatially-grounded, OCR-focused visual question answering (VQA) tuples with rich metadata to support diverse VLM training. The pipeline uses a multi-stage approach combining models like Nvidia's nemotron-ocr-v2, Gemma4, Qwen3-VL, and Gemini-2.5-Flash, along with an agentic optimization loop.

0 favorites 0 likes
#vision-language-model

Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents

Hugging Face Blog · 2026-03-31 Cached

IBM releases Granite 4.0 3B Vision, a compact vision-language model designed for enterprise document understanding, featuring specialized capabilities for table extraction, chart interpretation via ChartNet, and key-value pair grounding.

0 favorites 0 likes
#vision-language-model

dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

Papers with Code Trending · 2025-12-02 Cached

This paper presents dots.ocr, a unified Vision-Language Model that jointly learns layout detection, text recognition, and relational understanding for multilingual document layout parsing. It achieves state-of-the-art results on OmniDocBench and introduces the XDocParse benchmark spanning 126 languages.

0 favorites 0 likes
#vision-language-model

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Papers with Code Trending · 2025-10-09 Cached

Hulu-Med is a transparent medical vision-language model that unifies understanding across text, 2D/3D images, and video, achieving state-of-the-art performance on 30 benchmarks while being fully open-source.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback