HuggingFace

Articles from HuggingFace

Cards List

Harness-Zero: Harness Distillation via Agent-as-Harness

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Harness-Zero, a method for distilling optimized agent harnesses into large language models via agent-as-harness, significantly boosting task performance in knowledge work, tool use, and science domains even after specialized harnesses are removed.

0 favorites 0 likes

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Hugging Face Daily Papers · 2d ago Cached

onPanda is an interactive tool that uses token-level correction to efficiently annotate LLM alignment data and agent trajectories, reducing median annotation time by 52%.

0 favorites 0 likes

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

Hugging Face Daily Papers · 2d ago Cached

The paper introduces an information-efficiency ratio (IER) for optimizing token selection in sparse on-policy distillation, demonstrating that using only 1% of tokens can achieve performance comparable to full supervision.

0 favorites 0 likes

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Hugging Face Daily Papers · 2d ago Cached

CARE is a framework for Vision-Language-Action policies that enhances robotic manipulation by learning from execution failures to generate corrective actions, demonstrating improved success rates in simulations and real-world tasks.

0 favorites 0 likes

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

Hugging Face Daily Papers · 2d ago Cached

Document Retrieval-Aware Chunking (D-RAC) is a method that normalizes enterprise documents to PDF, converts them to retrieval-optimized Markdown using a multimodal LLM, and chunks them efficiently, significantly reducing token usage and costs compared to agentic chunking.

0 favorites 0 likes

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Hugging Face Daily Papers · 2d ago Cached

WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory for consistent and camera-controllable streaming scene exploration from a single image or text prompt.

0 favorites 0 likes

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.

0 favorites 0 likes

tokenizers v1: encode, decode and scaling, measured

Hugging Face Blog · 2d ago Cached

Hugging Face releases tokenizers v1, a major performance update for the tokenization library, with benchmarks showing significant speed improvements over previous versions.

0 favorites 0 likes

abenzerps/Qwen-Image-2.1-Uncensored-GGUF

Hugging Face Models Trending · 2d ago Cached

GGUF quantizations of the Qwen-Image-2.1 model for local image generation using ComfyUI, with recommended quantizations and setup instructions for deployment.

0 favorites 0 likes

Altworld/Hemmingway-1

Hugging Face Models Trending · 2d ago Cached

Hemmingway-1 is a 27B-parameter open-source AI model specialized for everyday writing tasks, outperforming leading models on benchmarks for human-like communication.

0 favorites 0 likes

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Hugging Face Daily Papers · 3d ago Cached

This paper proposes Grounded Action Models (GAMs), a new paradigm for robot foundation models that integrates 3D grounding, achieving state-of-the-art performance on manipulation tasks.

0 favorites 0 likes

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene

Hugging Face Daily Papers · 3d ago Cached

Mira-Scene introduces a compositional 3D scene reconstruction framework using pixel-aligned canonical coordinate maps for accurate object layouts, achieving significant improvements in layout accuracy over existing methods.

0 favorites 0 likes

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

Hugging Face Daily Papers · 3d ago Cached

This paper introduces a category-aware expert training framework for software engineering agents to mitigate uneven progress across task categories, using iterative training and multi-teacher distillation, with significant performance gains on Pro-618 and SWE-bench Multilingual benchmarks.

0 favorites 0 likes

Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

Hugging Face Daily Papers · 3d ago Cached

This paper presents an interpretability study on video diffusion models, revealing that Rotary Position Embedding (RoPE) induces excessive spatial attention decay, causing physics violations, and proposes a lightweight architectural modification to enhance physical coherence in generated videos.

0 favorites 0 likes

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Hugging Face Daily Papers · 4d ago Cached

The paper proposes Calibrated Clipping to stabilize FP8 quantization in reinforcement learning for LLMs by aligning clipping bounds with high-precision distributions, eliminating entropy surges and restoring performance.

0 favorites 0 likes

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Hugging Face Daily Papers · 4d ago Cached

UltraTex is an efficient framework for high-resolution multi-view diffusion-based 3D texturing, introducing techniques to reduce redundancy and achieve significant speedups in training and inference.

0 favorites 0 likes

OmniEdu: Open Foundation Models for Learning and Teaching

Hugging Face Daily Papers · 4d ago Cached

OmniEdu introduces an open family of foundation models for K-12 education, trained on a curated corpus to improve problem-solving, curriculum grounding, and pedagogical tutoring capabilities.

0 favorites 0 likes

Transferring the Intelligence of VLMs to Robotic Control

Hugging Face Daily Papers · 4d ago Cached

This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.

0 favorites 0 likes

convaiinnovations/laya

Hugging Face Models Trending · 5d ago Cached

Laya is an open-source non-autoregressive decision model that provides typed answers with calibrated probabilities, designed for tasks like email triage and conversational AI, showing significant performance improvements over existing models.

0 favorites 0 likes

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Hugging Face Daily Papers · 5d ago Cached

This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback