Tag
This paper evaluates the causal link between explanations and model predictions in vision-language reasoning through generation order interventions, finding that larger models are required for rationale-first reasoning and that answer-first generation reduces format-related errors.
This paper examines how the order of conflicting evidence in multimodal large language models affects judgments, revealing cross-modal evidence noncommutativity where placing perceptual evidence later increases model reliance on it.
This paper proposes a trajectory-aware decoding control method for diffusion vision-language models to address reasoning-budget mismatch by routing examples based on their decoding state, improving robustness across benchmarks.
This paper proposes a damage-aware multi-armed bandit method for structured post-training pruning of vision and language transformers, showing reduced performance degradation compared to baseline approaches in experiments across various models and datasets.
Ling-3.0-flash-VL is a newly released multimodal AI model from inclusionAI, featuring native image and video understanding, a 124B parameter architecture with 5.5B active parameters, a 1M token context, and an MIT license.
LLaDA-UI is a 16.7B-parameter mixture-of-experts diffusion vision-language agent that achieves strong multimodal GUI performance with block-parallel decoding efficiency, outperforming existing models on benchmarks.
Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that integrates 3D perception, visual question answering, and motion planning in a unified framework, achieving strong performance on benchmarks.
ZDTaichu5.0-9B is a multimodal foundation model that combines a Qwen3.5-9B language backbone with a C-RADIOv4-H vision encoder, excelling in general visual understanding, spatial reasoning, and agent tasks among 10B-scale VLMs.
LLaDA-Image presents a unified framework that combines a 6B diffusion transformer with a frozen vision-language module for generating photorealistic images with precise editing, achieving state-of-the-art results among open-source models through efficient training and fast inference.
SlideBank is a training-free framework for whole-slide image reasoning in pathology, using a persistent hierarchical evidence bank to improve consistency and reduce inference costs.
MedTVL is a text-guided dual-pathway architecture for medical time series classification that synergizes temporal and visual modalities with textual semantics, demonstrating superiority in various learning settings.
Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.
This article describes an uncensored version of Z.ai's GLM-5.3-Flash model, with safety alignments removed via abliteration, released as a block-FP8 checkpoint for research purposes.
This paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), showing that current LVLMs perform below human baselines and struggle with proactive question-driven grounding.
This paper compares prior injection methods for sparse-reward reinforcement learning in vision-language math reasoning, finding that effectiveness depends on delivery and that certain evaluation slices can mislead generalization assessments.
The Qwen3.8-27B-Uncensored is a 27B-parameter AI model with safety alignment removed via abliteration, released on Hugging Face for research purposes without built-in guardrails.
The paper identifies the Ghost Anchor phenomenon in multilingual MLLMs, where visual signals are underutilized during early alignment, and proposes the ANCHOR training framework to improve visual semantic emergence and performance across languages.
An uncensored MLX build of Qwen's Qwen3.8-27B model, quantized for Apple Silicon, with safety alignment removed for research purposes.
A modified version of Qwen3.8-27B with safety refusal removed and FP8 quantization, designed for research in AI safety and interpretability.
The Qwen 3.8 27B model has been released in an NVFP4 quantized version, enabling it to run on a single RTX 5090 GPU with enhancements in coding, agentic tasks, and vision-language understanding.