HuggingFace

Articles from HuggingFace

Cards List

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

Hugging Face Daily Papers · 6d ago Cached

Introduces Movement Trend Guidance to enhance 3D diffusion policies in robotic manipulation by providing foresight without explicit trajectories, achieving improved performance on benchmarks like RoboTwin2.0 and LIBERO-40.

0 favorites 0 likes

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Hugging Face Daily Papers · 6d ago Cached

This paper studies the feedback loop where AI-generated reviews influence future training of AI reviewers, leading to reduced judgment diversity called 'scientific-judgment collapse,' and introduces TrustReviewer, an open-source system to mitigate this through curated training and activation steering.

0 favorites 0 likes

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation

Hugging Face Daily Papers · 6d ago Cached

DeformSmith is a framework for generating interactive, physically credible deformable assets for robot manipulation from text or images, using physics-guided hierarchical generation to improve quality and plausibility.

0 favorites 0 likes

Self-Evolving Search Index

Hugging Face Daily Papers · 6d ago Cached

The paper introduces SELF-INDEX, a framework that enables search indexes to self-evolve autonomously, improving retrieval performance and benefiting downstream applications such as search agents and agent memory systems.

0 favorites 0 likes

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

Hugging Face Daily Papers · 6d ago Cached

This paper investigates how on-policy distillation can cause length inflation due to EOS token mismatches between student and teacher models, and proposes a correction method by aggregating EOS probabilities to reduce response length.

0 favorites 0 likes

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

Hugging Face Daily Papers · 6d ago Cached

FAMOS is a feed-forward model that predicts movable-part segmentation and joint parameters from sparse point clouds using a Multi-state Articulation Transformer and a procedural data generator, showing consistent improvements over baselines in experiments.

0 favorites 0 likes

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Hugging Face Daily Papers · 6d ago Cached

The paper introduces ActObs, a method that supervises both action and observation tokens in agent trajectories to improve reinforcement learning exploration, showing enhanced performance on benchmarks like Terminal-Bench2.0 and aider-polyglot.

0 favorites 0 likes

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hugging Face Daily Papers · 6d ago Cached

WeVisDoc is a two-stage data-centric framework for robust end-to-end document parsing that expands coverage and uses targeted diagnostics to improve performance, achieving state-of-the-art results on benchmarks.

0 favorites 0 likes

An Empirical Study of Harness Design for Coding Agents

Hugging Face Daily Papers · 6d ago Cached

This paper empirically studies harness design for coding agents, evaluating components like planning and context management to improve performance in software engineering tasks.

0 favorites 0 likes

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Hugging Face Daily Papers · 6d ago Cached

SoL-Pi introduces a method for recursively scaling auto-research loops in coding agents, achieving significant token and cost reductions while maintaining performance on benchmarks.

0 favorites 0 likes

Region-Level Policy Optimization for Fine-grained MLLM Perception

Hugging Face Daily Papers · 6d ago Cached

Vision-RL² is a method for improving fine-grained perception in multimodal large language models by using region-level reinforcement learning to compress visual tokens and enhance performance across multiple benchmarks without full model fine-tuning.

0 favorites 0 likes

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Hugging Face Daily Papers · 6d ago Cached

VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.

0 favorites 0 likes

JEPA-Anything: Learning Predictive Models across Different Worlds

Hugging Face Daily Papers · 6d ago Cached

JEPA-Anything presents a domain-agnostic framework based on orthogonal predictive factorization for learning predictive models across diverse systems like vision, biology, and control, with demonstrated improvements and experimental validation.

0 favorites 0 likes

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Hugging Face Daily Papers · 6d ago Cached

Video DeltaNet presents a hybrid attention mechanism combining Softmax and linear attention to enhance efficiency in video generation models, achieving a 14.5x speedup over baseline methods.

0 favorites 0 likes

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

Hugging Face Daily Papers · 6d ago Cached

UFO is a unified framework for simultaneous evaluation of omni-condition alignment in multi-modal image generation. It introduces an Atomized Chain-of-Evaluation paradigm and UFO-Bench benchmark.

0 favorites 0 likes

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Hugging Face Daily Papers · 6d ago Cached

Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.

0 favorites 0 likes

What Does Privileged Information Add to On-Policy Self-Distillation?

Hugging Face Daily Papers · 6d ago Cached

The paper investigates the contribution of privileged information in on-policy self-distillation for language models, finding that reference-free distillation accounts for most improvements, with limited additional benefits from privileged references.

0 favorites 0 likes

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

Hugging Face Daily Papers · 6d ago Cached

When2Think is a post-training framework that dynamically allocates computation in large reasoning models based on problem difficulty, improving accuracy-efficiency trade-offs on mathematical benchmarks.

0 favorites 0 likes

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Hugging Face Daily Papers · 6d ago Cached

The paper proposes RetireOPD, a method for training multi-turn agents using reinforcement learning with self-retiring on-policy distillation, improving performance on ALFWorld and WebShop benchmarks.

0 favorites 0 likes

prism-ml/Ternary-Bonsai-2-27B-mlx-2bit

Hugging Face Models Trending · 6d ago Cached

Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback