vision-language-model

Tag

Cards List
#vision-language-model

Self-Evolving Visual Questioner

Hugging Face Daily Papers · 2026-06-11 Cached

This paper introduces a self-evolving framework for vision-language models to improve their question-generation capabilities without external supervision, enhancing both question quality and answerer performance.

0 favorites 0 likes
#vision-language-model

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

arXiv cs.AI · 2026-06-10 Cached

This paper presents Architect-Ant, an editable automatic furnishing framework for architectural floor plans, together with a curated dataset (AntPlan-270) of 270 floor plans with furniture annotations. The method uses a fine-tuned vision-language model and a domain-specific language to generate geometrically valid and functionally plausible furniture layouts that can be rasterized into blueprint-style images.

0 favorites 0 likes
#vision-language-model

World Model Self-Distillation: Training World Models to Solve General Tasks

Hugging Face Daily Papers · 2026-06-10 Cached

A scalable framework combines self-distillation and reinforcement learning to transfer task-solving abilities from vision-language models to video diffusion models without requiring labeled task-video data.

0 favorites 0 likes
#vision-language-model

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

Hugging Face Daily Papers · 2026-06-08 Cached

OmniGameArena introduces a unified benchmark for evaluating VLM agents in diverse Unreal Engine 5 game environments, featuring an Improvement Dynamics Curve for tracking skill evolution across reflection rounds.

0 favorites 0 likes
#vision-language-model

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

Hugging Face Daily Papers · 2026-06-05 Cached

VoLoAgent integrates vision-language models with robot capabilities for open-vocabulary long-horizon manipulation tasks, introducing a physical orchestrator that plans, monitors, and recovers using interruptible tools, and a benchmark called RoboVoLo for evaluation.

0 favorites 0 likes
#vision-language-model

@_philschmid: We released Gemma 4 12B yesterday. Here is a visual guide that explains the full architecture. → How encoders typically…

X AI KOLs Following · 2026-06-04 Cached

A visual guide explaining the full architecture of Gemma 4 12B, covering how it handles text, images, and audio without separate encoder models by removing traditional vision and audio encoders.

0 favorites 0 likes
#vision-language-model

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Hugging Face Daily Papers · 2026-06-04 Cached

This paper introduces Structured Defect Grounding (SDG), a method that models text-to-image defects as structured (location, type, reason, importance) tuples and uses VLMs for detection, along with a 30K-image dataset SDG-30K and a diagnosis-to-alignment framework called BoxFlow-GRPO.

0 favorites 0 likes
#vision-language-model

Holo3.1 35B/9B/4B/0.8B (Qwen 3.5 finetunes)

Reddit r/LocalLLaMA · 2026-06-03

H Company releases Holo3.1, a family of Vision-Language Models (0.8B to 35B) for computer use agents, supporting web, desktop, and mobile automation with native function calling and optimized quantized checkpoints for local deployment.

0 favorites 0 likes
#vision-language-model

MAOAM: Unified Object and Material Selection with Vision-Language Models

Hugging Face Daily Papers · 2026-06-02 Cached

This paper presents MAOAM, a unified vision-language model framework that enables precise object and material selection through text or click interactions for interactive image editing. It introduces a scalable data generation pipeline and shows emergent improvement when combining text and clicks at inference.

0 favorites 0 likes
#vision-language-model

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

Hugging Face Daily Papers · 2026-06-02 Cached

PaddleOCR-VL-1.6 improves document parsing by identifying and refining under-optimized regions via targeted data optimization and progressive post-training, achieving state-of-the-art 96.33% on OmniDocBench v1.6.

0 favorites 0 likes
#vision-language-model

@Prince_Canuma: Today we're shipping our biggest MLX-VLM release yet: v0.6.0 ...and we are raising This one's about turning your Apple …

X AI KOLs Following · 2026-06-01 Cached

MLX-VLM v0.6.0 is released, adding speculative decoding, an agent-ready server compatible with Anthropic's API, new models (DeepSeek V4, ZAYA1-VL, etc.), image generation/editing, and audio input support, enabling local AI agents on Apple devices.

0 favorites 0 likes
#vision-language-model

MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?

Hugging Face Daily Papers · 2026-06-01 Cached

MMG2Skill converts web-based procedural guides into executable skills for agents through closed-loop learning, improving performance across GUI control, gameplay, and card play tasks with macro-average gains of +12.8 to +25.3 percentage points.

0 favorites 0 likes
#vision-language-model

PhyDrawGen: Physically Grounded Diagram Generation from Natural Language

arXiv cs.AI · 2026-06-01 Cached

PhyDrawGen is a neuro-symbolic pipeline that generates physically accurate diagrams from natural language by combining LLM-based scene understanding with a deterministic constraint solver and a VLM-based verify loop, outperforming existing models on a benchmark of physics problems.

0 favorites 0 likes
#vision-language-model

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

Hugging Face Daily Papers · 2026-05-31 Cached

This paper introduces 3DCodeBench, a benchmark for evaluating vision-language models on procedural 3D modeling via code, and 3DCodeArena, a ranking platform based on pairwise human preferences.

0 favorites 0 likes
#vision-language-model

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

arXiv cs.AI · 2026-05-29 Cached

This paper introduces the PiSAR benchmark for screen-conditioned action prediction and compares supervised fine-tuned models against frontier zero-shot baselines. Key findings show a fine-tuned Qwen3-VL-8B achieves 0.783 semantic similarity, significantly outperforming Claude Opus 4.7 and GPT-5.5 (0.459 and 0.482), but the same fine-tuning recipe on a larger reasoning-tuned Gemma model yields only 0.441, indicating a model-recipe mismatch.

0 favorites 0 likes
#vision-language-model

VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis

arXiv cs.AI · 2026-05-29 Cached

This paper proposes VFEAgent, a multi-agent system that automates finite element analysis by integrating vision-language models with a verification-first code synthesis framework, enabling end-to-end simulation from images and problem descriptions.

0 favorites 0 likes
#vision-language-model

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

Hugging Face Daily Papers · 2026-05-28

Stable-Layers is a reinforcement learning framework that fine-tunes a pretrained image layer decomposition model using VLM feedback instead of paired supervision, employing Flow-GRPO with LoRA and a two-stage reward calibration pipeline to improve layer quality on the Crello dataset.

0 favorites 0 likes
#vision-language-model

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

Hugging Face Daily Papers · 2026-05-28 Cached

This paper introduces VisAnomReasoner, a parameter-efficient vision-language model fine-tuned on a novel benchmark (VisAnomBench) with natural-language rationales, achieving over 21pp improvement in precision and F1 for time-series anomaly detection and strong cross-benchmark generalization.

0 favorites 0 likes
#vision-language-model

MedExpMem: Adapting Experience Memory for Differential Diagnosis

arXiv cs.LG · 2026-05-25 Cached

Proposes MedExpMem, an experience memory framework that enables medical vision-language models to accumulate and retrieve discriminative diagnostic experience from past cases, improving differential diagnosis accuracy by up to 7.0% on a radiology benchmark.

0 favorites 0 likes
#vision-language-model

InstructSAM: Segment Any Instance with Any Instructions

Hugging Face Daily Papers · 2026-05-25 Cached

InstructSAM presents a unified framework for multi-instance segmentation using instruction-driven queries that bridge vision-language models and SAM3, achieving strong results across complex benchmarks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback