Articles from HuggingFace
GGUF quantizations of the Qwen-Image-2.1 model for local image generation using ComfyUI, with recommended quantizations and setup instructions for deployment.
Hemmingway-1 is a 27B-parameter open-source AI model specialized for everyday writing tasks, outperforming leading models on benchmarks for human-like communication.
This paper proposes Grounded Action Models (GAMs), a new paradigm for robot foundation models that integrates 3D grounding, achieving state-of-the-art performance on manipulation tasks.
Mira-Scene introduces a compositional 3D scene reconstruction framework using pixel-aligned canonical coordinate maps for accurate object layouts, achieving significant improvements in layout accuracy over existing methods.
This paper introduces a category-aware expert training framework for software engineering agents to mitigate uneven progress across task categories, using iterative training and multi-teacher distillation, with significant performance gains on Pro-618 and SWE-bench Multilingual benchmarks.
This paper presents an interpretability study on video diffusion models, revealing that Rotary Position Embedding (RoPE) induces excessive spatial attention decay, causing physics violations, and proposes a lightweight architectural modification to enhance physical coherence in generated videos.
The paper proposes Calibrated Clipping to stabilize FP8 quantization in reinforcement learning for LLMs by aligning clipping bounds with high-precision distributions, eliminating entropy surges and restoring performance.
UltraTex is an efficient framework for high-resolution multi-view diffusion-based 3D texturing, introducing techniques to reduce redundancy and achieve significant speedups in training and inference.
OmniEdu introduces an open family of foundation models for K-12 education, trained on a curated corpus to improve problem-solving, curriculum grounding, and pedagogical tutoring capabilities.
This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.
Laya is an open-source non-autoregressive decision model that provides typed answers with calibrated probabilities, designed for tasks like email triage and conversational AI, showing significant performance improvements over existing models.
This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.
The paper presents PARTS, a real-world subtask reinforcement learning framework that improves long-horizon manipulation tasks by focusing on bottleneck subtasks with minimal human intervention, achieving higher success rates in experiments.
This paper presents APort Vault, a benchmark for evaluating payment authorization in AI agents, featuring over 225,000 evaluations across 14 models to test security policies and the Open Agent Passport specification.
Gricea is an open-science platform for conversational AI research that enables configurable and deployable research artifacts to support replication, extension, and cumulative knowledge building.
This paper introduces OmniVChat, a task for native audio-visual dialogue, and presents a data engine, benchmark, and reinforcement learning method to train and evaluate omni models, demonstrating improvements on synthesized and human-recorded data.
MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, achieving state-of-the-art performance on various benchmarks.
The paper introduces OmniVBench, a comprehensive benchmark for omni reference-to-video generation, and the Omni-R2VDataset, a large-scale training dataset, to evaluate and improve R2V models.
RecreationWorld introduces a scalable and verifiable framework for hybrid computer-use agents, providing environments across five platforms and a benchmark for evaluation.
GraphSkillEvo is an evolutionary optimization framework that represents agent skills as graph-structured artifacts to improve LLM performance, outperforming baselines on multiple benchmarks.