Tag
HomeBody is a humanoid robot system that explores environments, builds digital twins using Real2Sim, and leverages VLMs to perform autonomous long-horizon tasks like tidying kitchens.
The paper presents World Action Agent (WAA), a multi-agent system that uses Vision-Language Models (VLMs) for robot manipulation through a visual action workspace with contact views, action rehearsal, and in-view correction, achieving state-of-the-art results on benchmarks.
The paper proposes FedRepRAG, a decentralized retrieval-augmented generation framework that exchanges latent representations instead of raw content to reduce computational overhead and improve privacy in sensitive domains like healthcare.
CUSP is a training-free framework for quantifying collective uncertainty in multi-agent multimodal systems using semantic opinion pooling, improving reliability and accuracy over baseline methods.
The author seeks advice on choosing OCR models for a multi-document ERP system that handles mixed languages, handwritten text, and complex scan formats.
CoVeR is a training-free spatial token selector that improves 3D reasoning in Vision-Language Models by enforcing exact token budgets and full scene coverage, outperforming prior state-of-the-art methods.
A new paradigm enables a coding agent to use a VLM as a creative brain and an image model as a visual simulator to generate posters and infographics featuring real text, decoupled layers, and full editability.
A shared app for exploring and understanding vision language model evaluations, referencing a thread analyzing popular vision benchmarks.
LlamaIndex has significantly improved LlamaParse over the past three months, achieving 10-20% better accuracy on complex tables, charts, and grounding while maintaining costs below 0.4 cents per page, as measured against their ParseBench benchmark.
SVG-Score introduces a human-aligned evaluation framework for text-to-SVG generation, addressing limitations of current metrics like CLIPScore by developing specialized evaluators and benchmarks.
The paper introduces RoboSPA, a large-scale benchmark for evaluating vision-language-action models on fine-grained spatial reasoning and long-horizon procedural planning in robotic manipulation.
Radixark shares a blog post about how Miles supports multimodal learning for AI models with a shared post-training design for vision-language models and diffusion models.
The paper proposes credit-addressable reasoning with executable code traces and localized reinforcement learning to enhance multimodal geometry reasoning, achieving significant accuracy improvements over baseline models like Qwen3-VL-8B.
This paper presents a unified taxonomy for investigating multilingual multimodal misinformation on social media, using a large-scale dataset and automated annotation with a Vision-Language Model to uncover insights for detection and mitigation.
A new benchmark dataset and evaluation methodology for 52 text-to-image models has been published, including results, a leaderboard, and a gallery to assess performance on challenging prompts.
RubSE is a framework that uses rubric-guided self-evolution to enhance the stability of UI-to-code generation by mitigating visual repair coupling and trajectory collapse.
The article discusses a two-pass document processing trend for AI agents, where a fast OSS pass enables efficient retrieval and a VLM-based pass ensures accuracy, promoting Llama Index's LiteParse and LlamaParse tools to enhance cost and performance.
The article describes the process of fine-tuning a 450 million parameter vision-language model using 50,000 browser screenshots, showing progress from initial to current stages.
LFM2.5 DSpark by liquidai is integrated into mlx-vlm v0.6.16, enabling up to 3.7× faster speculative decoding on M5 Max with zero output drift for on-device VLM inference.
This paper benchmarks open-source OCR, LLM, and VLM systems for structured information extraction in a high-risk public sector application, finding that VLMs generally outperform OCR+LLM pipelines but most configurations struggle in zero-shot settings, emphasizing the critical role of input quality.