HuggingFace

Articles from HuggingFace

Cards List

harshatheg/Qwen-2.5-1B-RLCD

Hugging Face Models Trending · 2026-09-16 Cached

A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.

0 favorites 0 likes

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

Hugging Face Daily Papers · 2026-09-16 Cached

ALPINE introduces an ultra-lightweight spatial-relational architecture for few-shot image classification that achieves accuracy gains with fewer parameters, faster convergence, and better robustness compared to baselines like Prototypical Networks and MAML.

0 favorites 0 likes

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Hugging Face Daily Papers · 2026-09-16 Cached

This paper introduces GAVEL, a framework that uses graph world models to verify and repair long-horizon LLM planning for robotic tasks, significantly improving success rates and efficiency in simulations.

0 favorites 0 likes

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

Hugging Face Daily Papers · 2026-09-16 Cached

FRAUDSkill is a structured frozen-weight adaptation framework for audio anti-fraud detection that optimizes external skill programs without modifying the underlying audio-language model, achieving higher accuracy and reduced invalid outputs.

0 favorites 0 likes

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

Hugging Face Daily Papers · 2026-09-16 Cached

This paper evaluates MiniMax-H3, an omni-modal generative model, by introducing a comprehensive framework to assess its reasoning about the physical world through multimodal inputs. The evaluation reveals that video-based decision reasoning performs best, while audio-based disambiguation reasoning is the weakest.

0 favorites 0 likes

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

Hugging Face Daily Papers · 2026-09-16 Cached

PACT is a benchmark for assessing how LLM-based AI assistants comply with rules under pressure, covering 12 regulated enterprise domains and 48 realistic scenarios.

0 favorites 0 likes

In-Context Robot Learning with VLM Agents

Hugging Face Daily Papers · 2026-09-16 Cached

This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.

0 favorites 0 likes

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents

Hugging Face Daily Papers · 2026-09-16 Cached

CERA-MoA introduces a co-evolving framework for mixture-of-agents systems that uses reinforcement learning to dynamically route queries and adapt agent capabilities, enhancing task performance and efficiency.

0 favorites 0 likes

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Hugging Face Daily Papers · 2026-09-16 Cached

The paper presents PANORAMA, a vision-language model for panoptic grounded captioning that uses mask proposal selection to ground captions with pixel-level masks, and introduces the PanoCaps benchmark for evaluation.

0 favorites 0 likes

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Hugging Face Daily Papers · 2026-09-16 Cached

ScienceIDE introduces infrastructure for converting scientific code repositories into programmable environments for scientific agents, enabling task generation, execution, and verification, with trained models showing improvements in scientific code repair and general capabilities.

0 favorites 0 likes

Agora: Git as Shared Memory for Collective AutoResearch

Hugging Face Daily Papers · 2026-09-16 Cached

Agora is a shared memory system for autonomous AI research agents that uses Git to record research as an append-only DAG, enabling collaborative discovery. In a 12-day experiment, 13 agents improved a model's performance by 62% towards a trained baseline.

0 favorites 0 likes

A Zeroth-Order Paradigm for LLM Preference Alignment

Hugging Face Daily Papers · 2026-09-16 Cached

The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method for LLMs that uses comparison oracles to avoid likelihood displacement. It includes theoretical guarantees and experimental improvements over existing methods.

0 favorites 0 likes

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Hugging Face Daily Papers · 2026-09-16 Cached

This paper introduces ActionPiece, a novel action tokenization method for autoregressive vision-language-action models that uses physical rank consistency to improve the fidelity of action relationships, evaluated on benchmarks like LIBERO and SimplerEnv.

0 favorites 0 likes

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

Hugging Face Daily Papers · 2026-09-16 Cached

The paper identifies Value Flattening as a failure mode in PPO critic learning for LLMs and introduces SP3O, a sparse supervision method, to mitigate it, showing consistent improvements in experiments.

0 favorites 0 likes

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

Hugging Face Daily Papers · 2026-09-16 Cached

ProgramDistill introduces a scalable benchmark for evaluating coding agents by having them implement features discovered through interaction with fully functional reference web applications, using an automated pipeline to create verifiable tasks.

0 favorites 0 likes

Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Hugging Face Daily Papers · 2026-09-16 Cached

This paper compares gaze behavior in the MapTask and MUNDEX corpora to understand common grounding in collaborative tasks, finding that task-directed gaze is associated with aligned references and understood judgments, but the effects are modest.

0 favorites 0 likes

Comfy-Org/Qwen-Image-2.1

Hugging Face Models Trending · 2026-09-15 Cached

Repackaged model files for Qwen-Image-2.1 optimized for ComfyUI workflows, including text-to-image and image edit capabilities.

0 favorites 0 likes

Your Agent Aced the Task. Will It Do It Again?

Hugging Face Blog · 2026-09-15 Cached

This article addresses the consistency problem in AI agents, where tasks may fail on repeated attempts, and introduces ALTK-Evolve's Consistency Analyzer to diagnose and improve reliability, reducing the consistency gap from 24.4pp to 12.0pp without losing average accuracy.

0 favorites 0 likes

Embedding Physics Priors in Robot Learning: A Survey

Hugging Face Daily Papers · 2026-09-15 Cached

This survey paper reviews methods for embedding physics priors in robot learning, providing a unified taxonomy and discussing open challenges and future research directions in the field.

0 favorites 0 likes

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

Hugging Face Daily Papers · 2026-09-15 Cached

The paper introduces TAPe+MLv3, a compact computer vision system using structured representation for multi-task tasks, achieving competitive performance on benchmarks like COCO with fewer than 100,000 parameters.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback