vision-language-model

Tag

Cards List
#vision-language-model

InstanceControl: Controllable Complex Image Generation without Instance Labeling

Hugging Face Daily Papers · 2026-06-30 Cached

InstanceControl enables multi-instance controllable image generation without manual instance labeling by leveraging a vision-language model to establish instance-level correspondences between text prompts and visual conditions, with adaptive mask refinement for improved accuracy.

0 favorites 0 likes
#vision-language-model

Xiaomi-GUI-0 Technical Report

Hugging Face Daily Papers · 2026-06-30 Cached

This technical report presents Xiaomi-GUI-0, a native multimodal GUI agent trained and evaluated in real-device environments using a hybrid infrastructure and a progressive training pipeline, achieving high success rates and improved stability on real-world mobile tasks.

0 favorites 0 likes
#vision-language-model

Aloe-Vision: Robust Vision-Language Models for Healthcare

arXiv cs.CL · 2026-06-29 Cached

Aloe-Vision introduces a family of open medical Vision-Language Models trained on a quality-filtered mixture of medical and general data, along with a new benchmark CareQA-Vision for reliable evaluation. The models demonstrate competitive performance while highlighting vulnerabilities to adversarial inputs.

0 favorites 0 likes
#vision-language-model

@GithubProjects: Chunkr is an open-source document intelligence service that converts PDFs, PPTs, Word docs, and images into structured …

X AI KOLs Timeline · 2026-06-27 Cached

Chunkr is an open-source document intelligence service that converts PDFs, PPTs, Word docs, and images into structured chunks for RAG and LLM pipelines. It features layout analysis with OCR, structured HTML/Markdown output, vision-language model processing, and self-hosted deployment via Docker Compose with configurable LLM providers.

0 favorites 0 likes
#vision-language-model

Can Qwen3.6-35B-A3B on an RTX 3060 Replace Google Vision for Receipt-to-JSON Extraction?

Reddit r/LocalLLaMA · 2026-06-26

A developer shares their experience using a local Qwen VL model on an RTX 3060 to parse Japanese receipts into JSON, replacing Google Vision, with results showing accurate extraction of key fields at ~31 seconds per receipt.

0 favorites 0 likes
#vision-language-model

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

Hugging Face Daily Papers · 2026-06-26 Cached

This paper presents ZooClaw-FashionSigLIP2, a fashion-specialized vision-language model that achieves superior retrieval performance through full fine-tuning with knowledge distillation and weight interpolation, outperforming larger backbones and LoRA. It also introduces a new high-quality benchmark, ZooClaw-Fashion, and analyzes structural biases in existing datasets.

0 favorites 0 likes
#vision-language-model

The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents

arXiv cs.AI · 2026-06-24 Cached

The paper introduces the Latent Bridge, a trainable continuous channel that couples a slow reasoning VLM (Qwen3-VL-8B-Thinking) and a fast reactive VLM (MiniCPM-o 4.5) for real-time game agents. Experiments on Atari games and MetaDrive show it matches or outperforms the text-based bridge while avoiding destructive interference when used alone.

0 favorites 0 likes
#vision-language-model

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

Hugging Face Daily Papers · 2026-06-24 Cached

Physics Question Scene Graph (PQSG) is a hierarchical question-based pipeline using VLMs to evaluate video generation models' physical plausibility with fine-grained violation detection. It introduces the FinePhyEval dataset and shows higher correlation with human judgments than prior work.

0 favorites 0 likes
#vision-language-model

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

Hugging Face Daily Papers · 2026-06-23 Cached

This paper introduces CF-World, a counterfactual benchmark to evaluate whether text-to-image models rely on causal reasoning or mere pattern matching. Experiments show all models degrade sharply in counterfactual settings, suggesting their understanding is limited to tightly coupled visual-textual patterns rather than genuine causal reasoning.

0 favorites 0 likes
#vision-language-model

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

Hugging Face Daily Papers · 2026-06-22 Cached

ABACUS is a unified vision-language model that handles multiple counting tasks and count-faithful image generation without benchmark-specific training, achieving state-of-the-art results across seven benchmarks.

0 favorites 0 likes
#vision-language-model

Semantic Browsing: Controllable Diversity for Image Generation

Hugging Face Daily Papers · 2026-06-22 Cached

Semantic Browsing introduces a method for controlled diversity in text-to-image generation by using a Vision Language Model with an agentic workflow to generate structured, interpretable variations based on semantic decisions.

0 favorites 0 likes
#vision-language-model

A satellite is now running Google's Gemma 3 vision-language model in orbit, doing onboard inference instead of downlinking everything first

Reddit r/singularity · 2026-06-19

Loft Orbital's YAM-9 satellite runs Google's Gemma 3 vision-language model onboard for real-time image analysis, reducing downlink bandwidth and latency by deciding what data to send to Earth.

0 favorites 0 likes
#vision-language-model

@andimarafioti: Can a VLM see without a vision encoder? We trained one for $100, inspired by Gemma 4 12B. Latency on an M3 Pro MacBook:…

X AI KOLs Timeline · 2026-06-18 Cached

Researchers trained a vision-language model without a vision encoder for only $100, inspired by Gemma 4 12B, achieving a 30% reduction in end-to-end latency on an M3 Pro MacBook.

0 favorites 0 likes
#vision-language-model

NAVI-Orbital: First In-Orbit Demonstration of a Zero-Shot Vision-Language Model for Autonomous Earth Observation

arXiv cs.AI · 2026-06-18 Cached

NAVI-Orbital demonstrates the first in-orbit deployment of a zero-shot vision-language model (Gemma 3) on a LEO satellite, enabling autonomous scene classification and semantic compression of Earth observation data without fine-tuning.

0 favorites 0 likes
#vision-language-model

FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness

arXiv cs.AI · 2026-06-17 Cached

FinAcumen is a framework that accumulates reasoning experience from prior trajectories into a persistent memory bank for financial multimodal reasoning, improving performance across four benchmarks while maintaining a frozen 8B vision-language model.

0 favorites 0 likes
#vision-language-model

@atomic_chat_hq: Open-weight MiniMax M3 filled out a US customs form from a driver's license photo For this test we deployed MiniMax M3 …

X AI KOLs Timeline · 2026-06-15 Cached

A test of the open-weight MiniMax M3 model using MLX-VLM on a Mac Studio shows it can autonomously fill out a US customs form from a driver's license photo and a scanned document, using tool calls for fields, checkboxes, and signature.

0 favorites 0 likes
#vision-language-model

A satellite just learned to find things on its own — here’s what that means

TechCrunch AI · 2026-06-15 Cached

A satellite called Yam-9 used Google DeepMind's Gemma 3 vision-language model in orbit to autonomously identify areas of interest based on natural language queries, marking the first reported use of a VLM in space and signaling a shift toward more autonomous satellite operations.

0 favorites 0 likes
#vision-language-model

@jiqizhixin: What if your AI could “see” video like a streaming codec—spending tokens only on the most important moments? Introducin…

X AI KOLs Timeline · 2026-06-15 Cached

LLaVA-OneVision-2 introduces codec-stream tokenization for efficient video understanding, significantly outperforming Qwen3-VL-8B on temporal and spatial benchmarks. The model, data, and code are open-sourced.

0 favorites 0 likes
#vision-language-model

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

Hugging Face Daily Papers · 2026-06-14 Cached

SciOrch presents an 8B vision-language model trained with MCTS to coordinate multiple expert LLMs for multimodal scientific reasoning, achieving superior performance while reducing API costs.

0 favorites 0 likes
#vision-language-model

Detecting AI-Generated Content on Social Media with Multi-modal Language Models

arXiv cs.CL · 2026-06-11 Cached

This paper from Meta and Carnegie Mellon presents a multi-modal vision-language model pipeline for detecting AI-generated content on social media, achieving state-of-the-art performance and positive downstream impacts on user engagement.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback