HuggingFace

Articles from HuggingFace

Cards List

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Hugging Face Blog · yesterday Cached

UK AISI and EvalEval are collaborating to openly share AI evaluation results using a standardized schema and platform, enhancing reproducibility and transparency in benchmarking for AI models.

0 favorites 0 likes

Transformers now runs llama.cpp quants

Hugging Face Blog · yesterday Cached

Hugging Face's transformers library now supports GGUF models from llama.cpp, enabling efficient local inference on consumer hardware through familiar APIs.

0 favorites 0 likes

Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community

Hugging Face Blog · yesterday Cached

Jun Kim, creator of oMLX, joins Hugging Face to support the MLX community, enhancing stability and development for local AI on Apple Silicon.

0 favorites 0 likes

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Hugging Face Blog · yesterday Cached

The paper presents a physics-inspired approach to pruning LLM blocks by modeling block removal as a constrained binary optimization problem mapped to an Ising glass, achieving significant compression gains without benchmarking each configuration.

0 favorites 0 likes

ACLArena: Agent Continue Learning in Multi-stage Post-training

Hugging Face Daily Papers · 2d ago Cached

The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.

0 favorites 0 likes

Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Complex KDA, an enhanced version of Kimi Delta Attention that combines a delta-rule transformation with a reflection to achieve greater expressivity, outperforming Transformers in some tasks while maintaining efficiency, with open-source code and models available.

0 favorites 0 likes

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

Hugging Face Daily Papers · 2d ago Cached

EdgeGen is a synthetic task generation framework that creates database-grounded edge-case tasks to improve tool-calling agents through fine-tuning and harness optimization, demonstrating consistent performance improvements.

0 favorites 0 likes

Streaming Video Editing with Easy Adaptation

Hugging Face Daily Papers · 2d ago Cached

This paper introduces SVEET, a framework for high-quality streaming video editing that leverages a pretrained video diffusion model to enable auto-regressive editing with real-time performance on a single GPU.

0 favorites 0 likes

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

Hugging Face Daily Papers · 2d ago Cached

The paper introduces THAW-VLA, a method that distills world-model representations into Vision-Language-Action models for robotics, enhancing robustness and performance on simulation and real hardware without additional inference overhead.

0 favorites 0 likes

VideoGen-Agent: Reinforcing Video Generation Agents

Hugging Face Daily Papers · 2d ago Cached

The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.

0 favorites 0 likes

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Hugging Face Daily Papers · 2d ago Cached

Jev-Mem introduces an agentic memory architecture inspired by System-One/System-Two cognition, enhancing efficiency and effectiveness for long-horizon AI agents with improved scores and faster operations.

0 favorites 0 likes

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Hugging Face Daily Papers · 2d ago Cached

The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.

0 favorites 0 likes

Harness-Zero: Harness Distillation via Agent-as-Harness

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Harness-Zero, a method for distilling optimized agent harnesses into large language models via agent-as-harness, significantly boosting task performance in knowledge work, tool use, and science domains even after specialized harnesses are removed.

0 favorites 0 likes

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Hugging Face Daily Papers · 2d ago Cached

onPanda is an interactive tool that uses token-level correction to efficiently annotate LLM alignment data and agent trajectories, reducing median annotation time by 52%.

0 favorites 0 likes

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

Hugging Face Daily Papers · 2d ago Cached

The paper introduces an information-efficiency ratio (IER) for optimizing token selection in sparse on-policy distillation, demonstrating that using only 1% of tokens can achieve performance comparable to full supervision.

0 favorites 0 likes

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Hugging Face Daily Papers · 2d ago Cached

CARE is a framework for Vision-Language-Action policies that enhances robotic manipulation by learning from execution failures to generate corrective actions, demonstrating improved success rates in simulations and real-world tasks.

0 favorites 0 likes

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

Hugging Face Daily Papers · 2d ago Cached

Document Retrieval-Aware Chunking (D-RAC) is a method that normalizes enterprise documents to PDF, converts them to retrieval-optimized Markdown using a multimodal LLM, and chunks them efficiently, significantly reducing token usage and costs compared to agentic chunking.

0 favorites 0 likes

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Hugging Face Daily Papers · 2d ago Cached

WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory for consistent and camera-controllable streaming scene exploration from a single image or text prompt.

0 favorites 0 likes

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.

0 favorites 0 likes

tokenizers v1: encode, decode and scaling, measured

Hugging Face Blog · 2d ago Cached

Hugging Face releases tokenizers v1, a major performance update for the tokenization library, with benchmarks showing significant speed improvements over previous versions.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback