Tag
The Qwen 3.8 27B model has been released in an NVFP4 quantized version, enabling it to run on a single RTX 5090 GPU with enhancements in coding, agentic tasks, and vision-language understanding.
Qwen releases Qwen3.8-27B, an open-weights 27B dense vision-language model with major gains in coding, professional work, and long-horizon agentic tasks, available in FP8 with flexible thinking control.
Qwen3.8-27B is a new AI model with enhanced capabilities in coding, agentic tasks, and vision-language understanding, offering flexible thinking control and long context lengths. It is available on Hugging Face and designed for deployment-friendly use.
This paper proposes a YOLO- and CLIP-based vision-language framework to classify mosquito flight frames for Dengue virus detection, achieving 98.54% accuracy and 99.91% sensitivity at frame level, with complete video-level performance after temporal aggregation.
A Meta/Oxford study finds that multimodal models need surprisingly little image-generation data if language and visual understanding are trained together, suggesting an optimal 70/25/5 split of language, image understanding, and image generation tokens. It also warns that delaying vision training causes 'vision laziness' where models ignore images.
Qwen3.8-27B, a compact 27B dense vision-language model with flexible thinking control and long-context support, is released as the most capable Qwen open model to date, available soon via Hugging Face and Qwen Cloud.
Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.
Cohere released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model with native-resolution image support, multilingual and multi-image capabilities, released under Apache 2.0.
LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.
Introduces PoVisLE, a Polish vision-language evaluation benchmark with 1,117 images and 2,366 manually annotated VQA pairs, designed to assess culturally grounded multimodal understanding beyond surface-level recognition.
An empirical attribution study of failures in multi-page visually rich document understanding (MP-VRDU), isolating representation, selection, and reasoning failure modes and providing guidance for building such systems under fixed compute budgets.
Introduces Promptable Gaze Target Estimation (PGE), an end-to-end concept-driven paradigm using text or visual prompts for gaze analysis, along with the Gaze-Co dataset and the GazeAnywhere model.
This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.
This paper proposes a commutation theory for label-free reliability in vision-language figure reading, showing that consistency-based methods have a computable blind spot and introducing an Equivariance-Consistency Score enhanced by cyclic relabeling.
NVIDIA releases Nemotron Parse 2.0, a document image parsing model that converts scanned PDFs and images into structured text with layout, bounding boxes, and reading order, adding multilingual OCR improvements and chart-aware parsing.
Introduces CARGO-VL, a group-relative optimization framework for vision-language models that improves handling of conflicting image-text evidence and unsupported-answer avoidance via counterfactual consistency and risk-constrained control, along with the XMC conflict training resource.
ChronoVision is a multimodal framework that improves temporal reasoning in vision-language models by aligning visual logic with latent imagery, using a Reconstructive Visual Head, ROI Attention Locating module, and reinforcement learning. It introduces the Vbvr-VQA dataset and achieves SOTA accuracy on temporal tracking benchmarks.
This repository provides INT8 ConvRot-quantized ComfyUI safetensors of Qwen3-VL-32B, including a MiniMax-H3 conditioning encoder with layers 0-49 and an optional prompt-enhancement tail for layers 50-63, designed for use in ComfyUI on 32GB GPUs.
This paper introduces an information-asymmetric spot-the-difference task to measure epistemic vigilance in vision-language models, finding that models often overlook private evidence to agree with partners. Model steering to reduce sycophancy improves reliability in cooperative tasks.
This paper describes a two-stage vision-language adaptation system for Nepali meme classification, using Qwen3-VL-8B-Instruct with LoRA fine-tuning and contrastive learning. The system achieved 2nd place in hate speech detection and 4th in sentiment analysis at the CHiPSAL 2026 shared task.