vision-language

Tag

Cards List
#vision-language

Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions

arXiv cs.CL ↗ · 13h ago Cached

This paper evaluates the causal link between explanations and model predictions in vision-language reasoning through generation order interventions, finding that larger models are required for rationale-first reasoning and that answer-first generation reduces format-related errors.

0 favorites 0 likes
#vision-language

Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

arXiv cs.AI ↗ · yesterday Cached

This paper examines how the order of conflicting evidence in multimodal large language models affects judgments, revealing cross-modal evidence noncommutativity where placing perceptual evidence later increases model reliance on it.

0 favorites 0 likes
#vision-language

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

arXiv cs.AI ↗ · 2026-09-12 Cached

This paper proposes a trajectory-aware decoding control method for diffusion vision-language models to address reasoning-budget mismatch by routing examples based on their decoding state, improving robustness across benchmarks.

0 favorites 0 likes
#vision-language

Damage-Aware Bandit Pruning for Vision and Language Transformers

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper proposes a damage-aware multi-armed bandit method for structured post-training pruning of vision and language transformers, showing reduced performance degradation compared to baseline approaches in experiments across various models and datasets.

0 favorites 0 likes
#vision-language

@AdinaYakup: Ling-3.0-flash-VL just dropped from @inclusionAI 🔥 - Native image + video: understand > reason > act > verify - 124B/5…

X AI KOLs Timeline ↗ · 2026-09-09 Cached

Ling-3.0-flash-VL is a newly released multimodal AI model from inclusionAI, featuring native image and video understanding, a 124B parameter architecture with 5.5B active parameters, a 1M token context, and an MIT license.

0 favorites 0 likes
#vision-language

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Hugging Face Daily Papers ↗ · 2026-09-09 Cached

LLaDA-UI is a 16.7B-parameter mixture-of-experts diffusion vision-language agent that achieves strong multimodal GUI performance with block-parallel decoding efficiency, outperforming existing models on benchmarks.

0 favorites 0 likes
#vision-language

Qwen-Drive (GitHub Repo)

TLDR AI ↗ · 2026-09-08 Cached

Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that integrates 3D perception, visual question answering, and motion planning in a unified framework, achieving strong performance on benchmarks.

0 favorites 0 likes
#vision-language

TaichuAI/ZDTaichu5.0-9B

Hugging Face Models Trending ↗ · 2026-09-04 Cached

ZDTaichu5.0-9B is a multimodal foundation model that combines a Qwen3.5-9B language backbone with a C-RADIOv4-H vision encoder, excelling in general visual understanding, spatial reasoning, and agent tasks among 10B-scale VLMs.

0 favorites 0 likes
#vision-language

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

LLaDA-Image presents a unified framework that combines a 6B diffusion transformer with a frozen vision-language module for generating photorealistic images with precise editing, achieving state-of-the-art results among open-source models through efficient training and fast inference.

0 favorites 0 likes
#vision-language

SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning

arXiv cs.AI ↗ · 2026-09-02 Cached

SlideBank is a training-free framework for whole-slide image reasoning in pathology, using a persistent hierarchical evidence bank to improve consistency and reduce inference costs.

0 favorites 0 likes
#vision-language

MedTVL: Harnessing Vision and Language for Medical Time Series Classification

arXiv cs.AI ↗ · 2026-09-01 Cached

MedTVL is a text-guided dual-pathway architecture for medical time series classification that synergizes temporal and visual modalities with textual semantics, demonstrating superiority in various learning settings.

0 favorites 0 likes
#vision-language

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.

0 favorites 0 likes
#vision-language

orcarouter/GLM-5.3-Flash-Uncensored-FP8

Hugging Face Models Trending ↗ · 2026-08-29 Cached

This article describes an uncensored version of Z.ai's GLM-5.3-Flash model, with safety alignments removed via abliteration, released as a block-FP8 checkpoint for research purposes.

0 favorites 0 likes
#vision-language

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), showing that current LVLMs perform below human baselines and struggle with proactive question-driven grounding.

0 favorites 0 likes
#vision-language

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

arXiv cs.AI ↗ · 2026-08-25 Cached

This paper compares prior injection methods for sparse-reward reinforcement learning in vision-language math reasoning, finding that effectiveness depends on delivery and that certain evaluation slices can mislead generalization assessments.

0 favorites 0 likes
#vision-language

@p_misirov: Qwen3.8-27B-Uncensored really has no filter. https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

The Qwen3.8-27B-Uncensored is a 27B-parameter AI model with safety alignment removed via abliteration, released on Hugging Face for research purposes without built-in guardrails.

0 favorites 0 likes
#vision-language

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

arXiv cs.CL ↗ · 2026-08-18 Cached

The paper identifies the Ghost Anchor phenomenon in multilingual MLLMs, where visual signals are underutilized during early alignment, and proposes the ANCHOR training framework to improve visual semantic emergence and performance across languages.

0 favorites 0 likes
#vision-language

Qwen3.8-27B-Uncensored-MLX (4 minute read)

TLDR AI ↗ · 2026-08-18 Cached

An uncensored MLX build of Qwen's Qwen3.8-27B model, quantized for Apple Silicon, with safety alignment removed for research purposes.

0 favorites 0 likes
#vision-language

orcarouter/Qwen3.8-27B-Uncensored-FP8

Hugging Face Models Trending ↗ · 2026-08-15 Cached

A modified version of Qwen3.8-27B with safety refusal removed and FP8 quantization, designed for research in AI safety and interpretability.

0 favorites 0 likes
#vision-language

@TheAhmadOsman: Qwen 3.8 27B in NVFP4 would fit on a single RTX 5090 btw

X AI KOLs Following ↗ · 2026-08-14 Cached

The Qwen 3.8 27B model has been released in an NVFP4 quantized version, enabling it to run on a single RTX 5090 GPU with enhancements in coding, agentic tasks, and vision-language understanding.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback