HuggingFace

Articles from HuggingFace

Cards List

MintAct: A Unified Visual Agent for Digital Environments

Hugging Face Daily Papers · 6d ago Cached

MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, achieving state-of-the-art performance on various benchmarks.

0 favorites 0 likes

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Hugging Face Daily Papers · 6d ago Cached

The paper introduces OmniVBench, a comprehensive benchmark for omni reference-to-video generation, and the Omni-R2VDataset, a large-scale training dataset, to evaluate and improve R2V models.

0 favorites 0 likes

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Hugging Face Daily Papers · 6d ago Cached

RecreationWorld introduces a scalable and verifiable framework for hybrid computer-use agents, providing environments across five platforms and a benchmark for evaluation.

0 favorites 0 likes

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

Hugging Face Daily Papers · 6d ago Cached

GraphSkillEvo is an evolutionary optimization framework that represents agent skills as graph-structured artifacts to improve LLM performance, outperforming baselines on multiple benchmarks.

0 favorites 0 likes

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Hugging Face Daily Papers · 6d ago Cached

CodeMidas is an agentic pipeline that creates reinforcement learning environments from source code, scaling the training of coding agents and showing performance improvements on benchmarks like issue repair and program construction.

0 favorites 0 likes

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

Hugging Face Daily Papers · 2026-09-17 Cached

The paper proposes prediction-powered smoothing and validation methods for disaggregated AI evaluation, enhancing accuracy in point and interval estimates across domains with limited labeled data.

0 favorites 0 likes

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

Hugging Face Daily Papers · 2026-09-17 Cached

RefineEdit is a training-free prompt-to-prompt image editing method that uses a generative refinement network to enhance edit localization and background preservation, achieving top benchmark scores.

0 favorites 0 likes

Retention-Constrained Post-Training Quantization of Cellpose-SAM for Stem Cell Microscopy

Hugging Face Daily Papers · 2026-09-17 Cached

The paper evaluates compression schemes for Cellpose-SAM in stem cell microscopy, demonstrating that mixed-precision quantization achieves 6.76x reduction without catastrophic failures while maintaining segmentation accuracy.

0 favorites 0 likes

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Hugging Face Daily Papers · 2026-09-17 Cached

Paint-Anything introduces a unified hex-prompt interface for color control in image generation and editing, trained on a custom dataset and evaluated on a new benchmark showing significant improvements in color fidelity.

0 favorites 0 likes

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

Hugging Face Daily Papers · 2026-09-17 Cached

This paper proposes a training-adaptive convolutional sparse coding framework that leverages information bottleneck principles for robust visual representation, achieving improved performance on CIFAR and ImageNet under input perturbations.

0 favorites 0 likes

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

Hugging Face Daily Papers · 2026-09-17 Cached

The paper presents TeleAntiFraud 2.0, an audio-based benchmark for evaluating telecom fraud detection models using a mixed-tree generation pipeline and frozen monthly sets to address evolving fraud scripts and near-domain negatives.

0 favorites 0 likes

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

Hugging Face Daily Papers · 2026-09-17 Cached

Introduces Movement Trend Guidance to enhance 3D diffusion policies in robotic manipulation by providing foresight without explicit trajectories, achieving improved performance on benchmarks like RoboTwin2.0 and LIBERO-40.

0 favorites 0 likes

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Hugging Face Daily Papers · 2026-09-17 Cached

This paper studies the feedback loop where AI-generated reviews influence future training of AI reviewers, leading to reduced judgment diversity called 'scientific-judgment collapse,' and introduces TrustReviewer, an open-source system to mitigate this through curated training and activation steering.

0 favorites 0 likes

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation

Hugging Face Daily Papers · 2026-09-17 Cached

DeformSmith is a framework for generating interactive, physically credible deformable assets for robot manipulation from text or images, using physics-guided hierarchical generation to improve quality and plausibility.

0 favorites 0 likes

Self-Evolving Search Index

Hugging Face Daily Papers · 2026-09-17 Cached

The paper introduces SELF-INDEX, a framework that enables search indexes to self-evolve autonomously, improving retrieval performance and benefiting downstream applications such as search agents and agent memory systems.

0 favorites 0 likes

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

Hugging Face Daily Papers · 2026-09-17 Cached

This paper investigates how on-policy distillation can cause length inflation due to EOS token mismatches between student and teacher models, and proposes a correction method by aggregating EOS probabilities to reduce response length.

0 favorites 0 likes

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

Hugging Face Daily Papers · 2026-09-17 Cached

FAMOS is a feed-forward model that predicts movable-part segmentation and joint parameters from sparse point clouds using a Multi-state Articulation Transformer and a procedural data generator, showing consistent improvements over baselines in experiments.

0 favorites 0 likes

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Hugging Face Daily Papers · 2026-09-17 Cached

The paper introduces ActObs, a method that supervises both action and observation tokens in agent trajectories to improve reinforcement learning exploration, showing enhanced performance on benchmarks like Terminal-Bench2.0 and aider-polyglot.

0 favorites 0 likes

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hugging Face Daily Papers · 2026-09-17 Cached

WeVisDoc is a two-stage data-centric framework for robust end-to-end document parsing that expands coverage and uses targeted diagnostics to improve performance, achieving state-of-the-art results on benchmarks.

0 favorites 0 likes

An Empirical Study of Harness Design for Coding Agents

Hugging Face Daily Papers · 2026-09-17 Cached

This paper empirically studies harness design for coding agents, evaluating components like planning and context management to improve performance in software engineering tasks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback