vision-language-model

Tag

Cards List
#vision-language-model

Learning to Detect UI Principle Violations via Reinforcement Learning

arXiv cs.CL · 2026-07-24 Cached

This paper presents a method for training a lightweight vision-language model via reinforcement learning to detect violations of 19 UI quality principles in generated web pages, achieving 84% micro-F1. The model serves as a critic for auditing generated interfaces and providing design-aware feedback.

0 favorites 0 likes
#vision-language-model

SceneActBench: Can Agents Act on the 3D Scenes They See?

Hugging Face Daily Papers · 2026-07-24 Cached

SceneActBench is a benchmark for evaluating VLM agents on acting in complete multi-object 3D scenes, using task-specific geometric metrics across five tasks.

0 favorites 0 likes
#vision-language-model

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

Hugging Face Daily Papers · 2026-07-22 Cached

SeededGrasp proposes a data-efficient framework that uses a vision-language model to predict a seed point for a lightweight grasp generator, enabling language-guided grasping in complex scenes with multiple robot embodiments. The method outperforms baselines with 72% simulation and 78% real-world success, and includes a new large-scale multi-embodiment grasping dataset.

0 favorites 0 likes
#vision-language-model

Robostral Navigate

Hugging Face Daily Papers · 2026-07-22 Cached

Robostral Navigate is an 8B vision-language model that uses only monocular RGB images for robot navigation, achieving state-of-the-art on R2R-CE and RxR-CE benchmarks with 77.4% and 75.1% success rates, respectively.

0 favorites 0 likes
#vision-language-model

Introducing Cosmos 3 Edge

Hugging Face Blog · 2026-07-20 Cached

NVIDIA released Cosmos 3 Edge, a 4-billion-parameter open world model for edge devices that helps robots and vision AI agents understand surroundings, reason in real time, and generate actions. It achieves best-in-class throughput and accuracy among similar-sized models.

0 favorites 0 likes
#vision-language-model

Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

arXiv cs.AI · 2026-07-20 Cached

This paper introduces MAR-12, a framework using Vision-Language Models and multi-angle reasoning to detect and explain harmful humor in memes, achieving state-of-the-art accuracy on PrideMM and Memotion datasets.

0 favorites 0 likes
#vision-language-model

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

arXiv cs.AI · 2026-07-15 Cached

This paper investigates whether GRPO post-training improves a small (4B-8B) language and vision-language model web agent. It finds a controlled null result: no configuration yields credible gains on mastered tasks, and moderate-to-high learning rates cause degradation or collapse, revealing a double dissociation between degrade and collapse regimes.

0 favorites 0 likes
#vision-language-model

I trained a vision-language model to play Snake, and so can you. [P]

Reddit r/MachineLearning · 2026-07-14

A demonstration of training a vision-language model to play Snake using the FeynRL framework, illustrating the full training pipeline in an accessible manner.

0 favorites 0 likes
#vision-language-model

Prompt-Driven Exploration

arXiv cs.LG · 2026-07-13 Cached

The paper introduces Prompt-Driven Exploration (PDE), a method that uses a vision-language model to iteratively refine natural language prompts for reinforcement learning policies, enabling global exploration and successful policy learning even from zero-reward starts.

0 favorites 0 likes
#vision-language-model

moondream3.1-9B-A2B

Reddit r/LocalLLaMA · 2026-07-12 Cached

Moondream 3.1 is a vision language model with mixture-of-experts architecture (9B total parameters, 2B active), delivering state-of-the-art visual reasoning, detection, pointing, and captioning, deployable locally via the Photon inference engine or through the Moondream Cloud API.

0 favorites 0 likes
#vision-language-model

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

arXiv cs.AI · 2026-07-10 Cached

Introduces OmniFood-Bench, a benchmark for evaluating vision-language models on nutrient reasoning and personalized health advice. Experiments show VLMs struggle with mass estimation and safety-critical recommendations.

0 favorites 0 likes
#vision-language-model

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Hugging Face Daily Papers · 2026-07-06 Cached

HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model that improves efficiency via DFlash (6.37x inference speedup) and capability via Agentic Data Flow, achieving top-tier performance on document parsing, OCR, and multilingual tasks.

0 favorites 0 likes
#vision-language-model

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Hugging Face Daily Papers · 2026-07-06 Cached

InternVLA-A1.5 integrates pretrained vision-language models with future prediction in latent space to enable efficient robot manipulation with compositional generalization and long-horizon execution, achieving state-of-the-art results on simulation benchmarks.

0 favorites 0 likes
#vision-language-model

SINA: A Fully Automated Circuit Schematic Image to Netlist Generator Using Artificial Intelligence

arXiv cs.LG · 2026-07-03 Cached

This paper presents SINA, an open-source AI pipeline that automatically converts circuit schematic images into netlists with 96.67% accuracy, significantly improving over prior methods by integrating deep learning, OCR, and vision-language models.

0 favorites 0 likes
#vision-language-model

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

Hugging Face Daily Papers · 2026-07-02 Cached

Rank-Then-Act (RTA) is a framework for learning control policies from expert video demonstrations without environment rewards, using a Vision-Language Model as a progress-based ordinal scorer with correlation-based rewards. It achieves stable cross-task transfer and outperforms prior methods on discrete and continuous control benchmarks.

0 favorites 0 likes
#vision-language-model

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

arXiv cs.AI · 2026-07-01 Cached

This paper introduces Agentic RAG-VLM, a unified framework that integrates retrieval-augmented generation with vision-language models and self-reflective planning for generalizable robotic grasping in cluttered environments, achieving 78.3% success rate.

0 favorites 0 likes
#vision-language-model

EVLA: An Electro-Aware Multimodal Assistant for Physically-Grounded Driving Reasoning and Control

arXiv cs.CL · 2026-06-30 Cached

Introduces EVLA, a framework that enhances vision-language driving assistants with real-time awareness of electrified powertrain states, enabling energy-optimal and physically grounded decisions.

0 favorites 0 likes
#vision-language-model

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Hugging Face Daily Papers · 2026-06-30 Cached

3D HAMSTER enhances robot manipulation by using a vision-language model with depth encoding to generate 3D trajectories for point cloud-based control, outperforming 2D-guided baselines.

0 favorites 0 likes
#vision-language-model

InstanceControl: Controllable Complex Image Generation without Instance Labeling

Hugging Face Daily Papers · 2026-06-30 Cached

InstanceControl enables multi-instance controllable image generation without manual instance labeling by leveraging a vision-language model to establish instance-level correspondences between text prompts and visual conditions, with adaptive mask refinement for improved accuracy.

0 favorites 0 likes
#vision-language-model

Xiaomi-GUI-0 Technical Report

Hugging Face Daily Papers · 2026-06-30 Cached

This technical report presents Xiaomi-GUI-0, a native multimodal GUI agent trained and evaluated in real-device environments using a hybrid infrastructure and a progressive training pipeline, achieving high success rates and improved stability on real-world mobile tasks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback