vision-language-model

Tag

Cards List
#vision-language-model

Holo4

Reddit r/LocalLLaMA ↗ · 23h ago

H Company has released two vision-language models, Holo4-27B-GGUF and Holo4-35B-A3B-GGUF, designed for computer use automation. These models, built on Qwen architectures, work with the hai-agents harness to execute tasks like clicks, typing, and code.

0 favorites 0 likes
#vision-language-model

@lillyguisnet: A quality metric is a little difficult here since I haven't had the patience to make ground truth labels, but a 3-way a…

X AI KOLs Following ↗ · yesterday Cached

The tweet discusses the difficulty of setting a quality metric for polygon output without ground truth labels, but proposes a 3-way agreement method as reliable. It highlights impressive results for a non-specialized Vision-Language Model despite false positives/negatives, with improvements noted using a grid prompt.

0 favorites 0 likes
#vision-language-model

I trained a 500M VLM that answers typed questions about an image (choice / score / yes-no) with calibrated probabilities. ~400 ms on an M1 Pro, no text generation [P]

Reddit r/MachineLearning ↗ · 2d ago Cached

A small vision-language model that answers typed questions about images with calibrated probabilities, optimized for fast inference on consumer hardware.

0 favorites 0 likes
#vision-language-model

Accelerating vision-language models with LFM2.5-VL-DSpark

Hugging Face Blog ↗ · 4d ago Cached

Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.

0 favorites 0 likes
#vision-language-model

@yoheinakajima: glance-vlm speedlab is now open source! read: https://glance.yohei.me/speed/ try: https://github.com/yoheinakajima/glan…

X AI KOLs Timeline ↗ · 5d ago Cached

The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.

0 favorites 0 likes
#vision-language-model

apple/LensVLM-9B · Hugging Face

Reddit r/LocalLLaMA ↗ · 5d ago Cached

LensVLM-9B is a 9B-parameter Vision Language Model from Apple that scans compressed images of text and selectively expands relevant pages using learned tools, with paper, code, and usage instructions provided.

0 favorites 0 likes
#vision-language-model

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

HARMONY is an open-source framework using hierarchical agentic reasoning with VLMs to reconstruct compositional 3D scenes from single indoor images, competing with GPT-6 Astra in visual quality and outperforming in geometric alignment.

0 favorites 0 likes
#vision-language-model

@yoheinakajima: Jev-style logit read on a 4B open VLM, measured: http://glance.yohei.me vs the same model writing JSON: ~1/3 less time …

X AI KOLs Timeline ↗ · 2026-09-21 Cached

Yohei Nakajima presents a method to read typed visual judgements from a frozen open vision-language model using logits, achieving similar accuracy to hosted models with reduced time and GPU cost.

0 favorites 0 likes
#vision-language-model

Toward individual-level calibration in affect recognition with perceptual adjustment queries

arXiv cs.LG ↗ · 2026-09-21 Cached

The paper introduces a framework for individual-level calibration in facial affect recognition using perceptual adjustment queries to normalize perceptual difficulty, validated in a behavioral study.

0 favorites 0 likes
#vision-language-model

PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking

arXiv cs.AI ↗ · 2026-09-21 Cached

This paper introduces PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a reasoning problem for VLSI physical design, achieving significant improvements in timing and wirelength. It also presents PlaceReasoner-Bench, an open benchmark for evaluating methods using routed PPA and DRC.

0 favorites 0 likes
#vision-language-model

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

X-Planner introduces an event-structured planning front-end for embodied AI, improving task planning in long-horizon manipulation through supervised data and a VLM backbone with discrete and latent interfaces.

0 favorites 0 likes
#vision-language-model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

RULER introduces instance-aware rubric rewards for SVG generation, using a vision-language judge to optimize reinforcement learning and significantly improve performance over previous methods.

0 favorites 0 likes
#vision-language-model

@HuggingModels: OCR just got a major upgrade. jina-ocr-v1 is a multimodal vision language model built for document intelligence. It rea…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.

0 favorites 0 likes
#vision-language-model

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Hugging Face Daily Papers ↗ · 2026-09-16 Cached

The paper presents PANORAMA, a vision-language model for panoptic grounded captioning that uses mask proposal selection to ground captions with pixel-level masks, and introduces the PanoCaps benchmark for evaluation.

0 favorites 0 likes
#vision-language-model

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

ModaLens is a paired image-swap audit that measures how report availability reduces image sensitivity in medical vision-language models, demonstrated using MedGemma-27B on the MIMIC-CXR dataset.

0 favorites 0 likes
#vision-language-model

AI-Powered Flare Combustion Efficiency Estimation

arXiv cs.AI ↗ · 2026-09-12 Cached

This academic paper proposes an AI-based system that uses thermal video footage to estimate combustion efficiency in flare stacks, integrated into a graphical user interface with high uptime and low maintenance.

0 favorites 0 likes
#vision-language-model

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

Hugging Face Daily Papers ↗ · 2026-09-10 Cached

A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits, reaching 89% accuracy of the larger pipeline with only 1.1% of parameters.

0 favorites 0 likes
#vision-language-model

@AdinaYakup: This is new 👀 @Alibaba_Qwen just released an open VLM foundation model for autonomous driving on @huggingface - 4B / A…

X AI KOLs Following ↗ · 2026-09-09 Cached

Alibaba Qwen has released an open-source 4B parameter vision-language model for autonomous driving on Hugging Face, featuring 3D detection, occupancy prediction, and BEV capabilities while preserving strong VL abilities.

0 favorites 0 likes
#vision-language-model

Show-Harness: Just a VLM Agent Can Play Robots

Hugging Face Daily Papers ↗ · 2026-09-09 Cached

Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.

0 favorites 0 likes
#vision-language-model

@HuggingModels: Ever seen an AI that reads images AND writes text? DeepSeek-V4-Flash-Vision-Exp does exactly that. It's a vision-langua…

X AI KOLs Timeline ↗ · 2026-09-08

DeepSeek-V4-Flash-Vision-Exp is a vision-language model that processes images and generates text, with over 313k downloads, useful for tasks like image description and visual question answering.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback