Tag
H Company has released two vision-language models, Holo4-27B-GGUF and Holo4-35B-A3B-GGUF, designed for computer use automation. These models, built on Qwen architectures, work with the hai-agents harness to execute tasks like clicks, typing, and code.
The tweet discusses the difficulty of setting a quality metric for polygon output without ground truth labels, but proposes a 3-way agreement method as reliable. It highlights impressive results for a non-specialized Vision-Language Model despite false positives/negatives, with improvements noted using a grid prompt.
A small vision-language model that answers typed questions about images with calibrated probabilities, optimized for fast inference on consumer hardware.
Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B, which adds a speculative decoding path for faster inference with minimal memory cost.
The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.
LensVLM-9B is a 9B-parameter Vision Language Model from Apple that scans compressed images of text and selectively expands relevant pages using learned tools, with paper, code, and usage instructions provided.
HARMONY is an open-source framework using hierarchical agentic reasoning with VLMs to reconstruct compositional 3D scenes from single indoor images, competing with GPT-6 Astra in visual quality and outperforming in geometric alignment.
Yohei Nakajima presents a method to read typed visual judgements from a frozen open vision-language model using logits, achieving similar accuracy to hosted models with reduced time and GPU cost.
The paper introduces a framework for individual-level calibration in facial affect recognition using perceptual adjustment queries to normalize perceptual difficulty, validated in a behavioral study.
This paper introduces PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a reasoning problem for VLSI physical design, achieving significant improvements in timing and wirelength. It also presents PlaceReasoner-Bench, an open benchmark for evaluating methods using routed PPA and DRC.
X-Planner introduces an event-structured planning front-end for embodied AI, improving task planning in long-horizon manipulation through supervised data and a VLM backbone with discrete and latent interfaces.
RULER introduces instance-aware rubric rewards for SVG generation, using a vision-language judge to optimize reinforcement learning and significantly improve performance over previous methods.
Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.
The paper presents PANORAMA, a vision-language model for panoptic grounded captioning that uses mask proposal selection to ground captions with pixel-level masks, and introduces the PanoCaps benchmark for evaluation.
ModaLens is a paired image-swap audit that measures how report availability reduces image sensitivity in medical vision-language models, demonstrated using MedGemma-27B on the MIMIC-CXR dataset.
This academic paper proposes an AI-based system that uses thermal video footage to estimate combustion efficiency in flare stacks, integrated into a graphical user interface with high uptime and low maintenance.
A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits, reaching 89% accuracy of the larger pipeline with only 1.1% of parameters.
Alibaba Qwen has released an open-source 4B parameter vision-language model for autonomous driving on Hugging Face, featuring 3D detection, occupancy prediction, and BEV capabilities while preserving strong VL abilities.
Show-Harness is a method that enables vision-language models to control robots through discrete semantic actions, allowing zero-shot deployment and efficient fine-tuning across different robots and GUIs.
DeepSeek-V4-Flash-Vision-Exp is a vision-language model that processes images and generates text, with over 313k downloads, useful for tasks like image description and visual question answering.