vision-language

Tag

Cards List
#vision-language

@TheAhmadOsman: Qwen 3.8 27B in NVFP4 would fit on a single RTX 5090 btw

X AI KOLs Following · 22h ago Cached

The Qwen 3.8 27B model has been released in an NVFP4 quantized version, enabling it to run on a single RTX 5090 GPU with enhancements in coding, agentic tasks, and vision-language understanding.

0 favorites 0 likes
#vision-language

Qwen 3.8 27B is out: open weights, best local dense model yet

Hacker News Top · 23h ago Cached

Qwen releases Qwen3.8-27B, an open-weights 27B dense vision-language model with major gains in coding, professional work, and long-horizon agentic tasks, available in FP8 with flexible thinking control.

0 favorites 0 likes
#vision-language

@no_stp_on_snek: Just a few hours away from Qwen 3.8! Clear your benches! I’ll be working on behavioral tests and comparisons against 3.…

X AI KOLs Timeline · yesterday Cached

Qwen3.8-27B is a new AI model with enhanced capabilities in coding, agentic tasks, and vision-language understanding, offering flexible thinking control and long context lengths. It is available on Hugging Face and designed for deployment-friendly use.

0 favorites 0 likes
#vision-language

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

arXiv cs.AI · yesterday Cached

This paper proposes a YOLO- and CLIP-based vision-language framework to classify mosquito flight frames for Dengue virus detection, achieving 98.54% accuracy and 99.91% sensitivity at frame level, with complete video-level performance after temporal aggregation.

0 favorites 0 likes
#vision-language

@rohanpaul_ai: The Meta/Oxford study finds, a multimodal model may need surprisingly little image-generation data if language and visu…

X AI KOLs Following · yesterday Cached

A Meta/Oxford study finds that multimodal models need surprisingly little image-generation data if language and visual understanding are trained together, suggesting an optimal 70/25/5 split of language, image understanding, and image generation tokens. It also warns that delaying vision training causes 'vision laziness' where models ignore images.

0 favorites 0 likes
#vision-language

@AdinaYakup: Open source summer party is not over yet https://huggingface.co/Qwen/Qwen3.8-27B…

X AI KOLs Timeline · 2d ago Cached

Qwen3.8-27B, a compact 27B dense vision-language model with flexible thinking control and long-context support, is released as the most capable Qwen open model to date, available soon via Hugging Face and Qwen Cloud.

0 favorites 0 likes
#vision-language

unsloth/Qwen3.8-27B-NVFP4

Hugging Face Models Trending · 2d ago Cached

Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.

0 favorites 0 likes
#vision-language

CohereLabs/North-Micro-Vision-Instruct · Hugging Face

Reddit r/LocalLLaMA · 2d ago Cached

Cohere released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model with native-resolution image support, multilingual and multi-image capabilities, released under Apache 2.0.

0 favorites 0 likes
#vision-language

LiquidAI/LFM2.5-VL-3B · Hugging Face

Reddit r/LocalLLaMA · 3d ago Cached

LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.

0 favorites 0 likes
#vision-language

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

arXiv cs.CL · 4d ago Cached

Introduces PoVisLE, a Polish vision-language evaluation benchmark with 1,117 images and 2,366 manually annotated VQA pairs, designed to assess culturally grounded multimodal understanding beyond surface-level recognition.

0 favorites 0 likes
#vision-language

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

arXiv cs.AI · 4d ago Cached

An empirical attribution study of failures in multi-page visually rich document understanding (MP-VRDU), isolating representation, selection, and reasoning failure modes and providing guidance for building such systems under fixed compute budgets.

0 favorites 0 likes
#vision-language

Gaze Target Estimation Anywhere with Concepts

Hugging Face Daily Papers · 4d ago Cached

Introduces Promptable Gaze Target Estimation (PGE), an end-to-end concept-driven paradigm using text or visual prompts for gaze analysis, along with the Gaze-Co dataset and the GazeAnywhere model.

0 favorites 0 likes
#vision-language

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

arXiv cs.CL · 2026-08-07 Cached

This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.

0 favorites 0 likes
#vision-language

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

arXiv cs.LG · 2026-08-07 Cached

This paper proposes a commutation theory for label-free reliability in vision-language figure reading, showing that consistency-based methods have a computable blind spot and introducing an Equivariance-Consistency Score enhanced by cyclic relabeling.

0 favorites 0 likes
#vision-language

nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face

Reddit r/LocalLLaMA · 2026-08-06 Cached

NVIDIA releases Nemotron Parse 2.0, a document image parsing model that converts scanned PDFs and images into structured text with layout, bounding boxes, and reading order, adding multilingual OCR improvements and chart-aware parsing.

0 favorites 0 likes
#vision-language

CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

arXiv cs.AI · 2026-08-06 Cached

Introduces CARGO-VL, a group-relative optimization framework for vision-language models that improves handling of conflicting image-text evidence and unsupported-answer avoidance via counterfactual consistency and risk-constrained control, along with the XMC conflict training resource.

0 favorites 0 likes
#vision-language

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Hugging Face Daily Papers · 2026-08-06 Cached

ChronoVision is a multimodal framework that improves temporal reasoning in vision-language models by aligning visual logic with latent imagery, using a Reconstructive Visual Head, ROI Attention Locating module, and reinforcement learning. It introduces the Vbvr-VQA dataset and achieves SOTA accuracy on temporal tracking benchmarks.

0 favorites 0 likes
#vision-language

ethanfel/Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot

Hugging Face Models Trending · 2026-08-03 Cached

This repository provides INT8 ConvRot-quantized ComfyUI safetensors of Qwen3-VL-32B, including a MiniMax-H3 conditioning encoder with layers 0-49 and an optional prompt-enhancement tail for layers 50-63, designed for use in ComfyUI on 32GB GPUs.

0 favorites 0 likes
#vision-language

Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks

arXiv cs.CL · 2026-08-03 Cached

This paper introduces an information-asymmetric spot-the-difference task to measure epistemic vigilance in vision-language models, finding that models often overlook private evidence to agree with partners. Model steering to reduce sycophancy improves reliability in cooperative tasks.

0 favorites 0 likes
#vision-language

ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification

arXiv cs.CL · 2026-08-03 Cached

This paper describes a two-stage vision-language adaptation system for Nepali meme classification, using Qwen3-VL-8B-Instruct with LoRA fine-tuning and contrastive learning. The system achieved 2nd place in hate speech detection and 4th in sentiment analysis at the CHiPSAL 2026 shared task.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback