HuggingFace

Articles from HuggingFace

Cards List

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Hugging Face Daily Papers · 2d ago Cached

The paper proposes SpeakerMem-R1, a speaker-centered dual-track memory system for multi-party dialogue, addressing bottlenecks in message attribution and relational understanding. It achieves state-of-the-art results on benchmarks like EverMemBench and LoCoMo.

0 favorites 0 likes

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Hugging Face Daily Papers · 2d ago Cached

This paper introduces JEV-as-a-Judge, a cost-effective evaluation method for LLMs that uses a decision-only judge with confidence thresholds to accept certain verdicts and escalate uncertain ones, achieving comparable accuracy to state-of-the-art models at significantly lower cost.

0 favorites 0 likes

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Hugging Face Daily Papers · 2d ago Cached

Agensh is a scalable self-organized multi-agent system without a central orchestrator that improves performance on complex tasks by scaling the number of agents, showing significant test-pass rate increases on benchmarks like ProgramBench and pandoc.

0 favorites 0 likes

Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament

Hugging Face Daily Papers · 2d ago Cached

This study examines blame attribution in the Danish Parliament from 1997 to 2026 using the BlameBERT classifier, revealing ideological asymmetries and a banana-shaped trajectory in political discourse.

0 favorites 0 likes

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

Hugging Face Daily Papers · 2d ago Cached

RoboFollow introduces a diagnostic benchmark to expose the illusion of instruction-following in embodied agents by analyzing high scene entropy and perturbations, revealing gaps in current models despite strong initial performance.

0 favorites 0 likes

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training

Hugging Face Daily Papers · 2d ago Cached

StableVQ proposes practical guidelines to stabilize the training of vector-quantized tokenizers by decoupling encoder-decoder and codebook training, improving stability and codebook utilization for image generation models.

0 favorites 0 likes

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Taste-Bench, a benchmark for measuring taste in LLM agents' long-horizon decisions, finding that frontier models have low accuracy and that taste can be improved through distillation training.

0 favorites 0 likes

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Hugging Face Daily Papers · 2d ago Cached

Flash-dLLM is a training-free inference acceleration framework for diffusion LLMs that uses IO-aware KV caching and parallel decoding to achieve significant speedups and memory efficiency improvements.

0 favorites 0 likes

Recursive self-improvement of AI research agents

Hugging Face Daily Papers · 2d ago Cached

This paper introduces AIDE^2, a system that enables AI research agents to autonomously improve their own code through recursive self-improvement, leading to performance gains across various AI research tasks.

0 favorites 0 likes

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Hugging Face Blog · 2d ago Cached

UK AISI and EvalEval are collaborating to openly share AI evaluation results using a standardized schema and platform, enhancing reproducibility and transparency in benchmarking for AI models.

0 favorites 0 likes

Transformers now runs llama.cpp quants

Hugging Face Blog · 2d ago Cached

Hugging Face's transformers library now supports GGUF models from llama.cpp, enabling efficient local inference on consumer hardware through familiar APIs.

0 favorites 0 likes

Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community

Hugging Face Blog · 2d ago Cached

Jun Kim, creator of oMLX, joins Hugging Face to support the MLX community, enhancing stability and development for local AI on Apple Silicon.

0 favorites 0 likes

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Hugging Face Blog · 2d ago Cached

The paper presents a physics-inspired approach to pruning LLM blocks by modeling block removal as a constrained binary optimization problem mapped to an Ising glass, achieving significant compression gains without benchmarking each configuration.

0 favorites 0 likes

HappyWorld-Bench

Hugging Face Daily Papers · 3d ago Cached

HappyWorld-Bench is a comprehensive benchmark that evaluates the reliability of world models under interaction and modification across video, spatial, and embodied tracks.

0 favorites 0 likes

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

Hugging Face Daily Papers · 3d ago Cached

ImIR adapts a pretrained image-editing model for six image restoration tasks using image-derived instructions, enabling efficient and task-agnostic restoration.

0 favorites 0 likes

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Hugging Face Daily Papers · 3d ago Cached

The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.

0 favorites 0 likes

Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Hugging Face Daily Papers · 3d ago Cached

LeanPool is a repository of formalized mathematics that is grown, maintained, and optimized by AI agents.

0 favorites 0 likes

Emergent Collusion in Long-Horizon LLM Agent Interaction

Hugging Face Daily Papers · 3d ago Cached

This paper studies the emergence of collusion in long-horizon multi-agent environments with LLM agents, finding that agents increasingly deviate from verification protocols over repeated interactions, posing safety risks.

0 favorites 0 likes

RULER: Instance-aware Rubric Rewards for SVG Generation

Hugging Face Daily Papers · 3d ago Cached

RULER introduces instance-aware rubric rewards for SVG generation, using a vision-language judge to optimize reinforcement learning and significantly improve performance over previous methods.

0 favorites 0 likes

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

Hugging Face Daily Papers · 3d ago Cached

The paper introduces Ovis-Embedding, a state-of-the-art omni-modal embedding model that uses a shared backbone to encode text, image, video, and audio in a common representation space, achieving top performance on benchmarks like MMEB-v3 and MVEB.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback