Tag
This paper presents a method for training a lightweight vision-language model via reinforcement learning to detect violations of 19 UI quality principles in generated web pages, achieving 84% micro-F1. The model serves as a critic for auditing generated interfaces and providing design-aware feedback.
SceneActBench is a benchmark for evaluating VLM agents on acting in complete multi-object 3D scenes, using task-specific geometric metrics across five tasks.
SeededGrasp proposes a data-efficient framework that uses a vision-language model to predict a seed point for a lightweight grasp generator, enabling language-guided grasping in complex scenes with multiple robot embodiments. The method outperforms baselines with 72% simulation and 78% real-world success, and includes a new large-scale multi-embodiment grasping dataset.
Robostral Navigate is an 8B vision-language model that uses only monocular RGB images for robot navigation, achieving state-of-the-art on R2R-CE and RxR-CE benchmarks with 77.4% and 75.1% success rates, respectively.
NVIDIA released Cosmos 3 Edge, a 4-billion-parameter open world model for edge devices that helps robots and vision AI agents understand surroundings, reason in real time, and generate actions. It achieves best-in-class throughput and accuracy among similar-sized models.
This paper introduces MAR-12, a framework using Vision-Language Models and multi-angle reasoning to detect and explain harmful humor in memes, achieving state-of-the-art accuracy on PrideMM and Memotion datasets.
This paper investigates whether GRPO post-training improves a small (4B-8B) language and vision-language model web agent. It finds a controlled null result: no configuration yields credible gains on mastered tasks, and moderate-to-high learning rates cause degradation or collapse, revealing a double dissociation between degrade and collapse regimes.
A demonstration of training a vision-language model to play Snake using the FeynRL framework, illustrating the full training pipeline in an accessible manner.
The paper introduces Prompt-Driven Exploration (PDE), a method that uses a vision-language model to iteratively refine natural language prompts for reinforcement learning policies, enabling global exploration and successful policy learning even from zero-reward starts.
Moondream 3.1 is a vision language model with mixture-of-experts architecture (9B total parameters, 2B active), delivering state-of-the-art visual reasoning, detection, pointing, and captioning, deployable locally via the Photon inference engine or through the Moondream Cloud API.
Introduces OmniFood-Bench, a benchmark for evaluating vision-language models on nutrient reasoning and personalized health advice. Experiments show VLMs struggle with mass estimation and safety-critical recommendations.
HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model that improves efficiency via DFlash (6.37x inference speedup) and capability via Agentic Data Flow, achieving top-tier performance on document parsing, OCR, and multilingual tasks.
InternVLA-A1.5 integrates pretrained vision-language models with future prediction in latent space to enable efficient robot manipulation with compositional generalization and long-horizon execution, achieving state-of-the-art results on simulation benchmarks.
This paper presents SINA, an open-source AI pipeline that automatically converts circuit schematic images into netlists with 96.67% accuracy, significantly improving over prior methods by integrating deep learning, OCR, and vision-language models.
Rank-Then-Act (RTA) is a framework for learning control policies from expert video demonstrations without environment rewards, using a Vision-Language Model as a progress-based ordinal scorer with correlation-based rewards. It achieves stable cross-task transfer and outperforms prior methods on discrete and continuous control benchmarks.
This paper introduces Agentic RAG-VLM, a unified framework that integrates retrieval-augmented generation with vision-language models and self-reflective planning for generalizable robotic grasping in cluttered environments, achieving 78.3% success rate.
Introduces EVLA, a framework that enhances vision-language driving assistants with real-time awareness of electrified powertrain states, enabling energy-optimal and physically grounded decisions.
3D HAMSTER enhances robot manipulation by using a vision-language model with depth encoding to generate 3D trajectories for point cloud-based control, outperforming 2D-guided baselines.
InstanceControl enables multi-instance controllable image generation without manual instance labeling by leveraging a vision-language model to establish instance-level correspondences between text prompts and visual conditions, with adaptive mask refinement for improved accuracy.
This technical report presents Xiaomi-GUI-0, a native multimodal GUI agent trained and evaluated in real-device environments using a hybrid infrastructure and a progressive training pipeline, achieving high success rates and improved stability on real-world mobile tasks.