HuggingFace

Articles from HuggingFace

Cards List

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

Hugging Face Daily Papers · 2026-09-17 Cached

SoL-Pi introduces a method for recursively scaling auto-research loops in coding agents, achieving significant token and cost reductions while maintaining performance on benchmarks.

0 favorites 0 likes

Region-Level Policy Optimization for Fine-grained MLLM Perception

Hugging Face Daily Papers · 2026-09-17 Cached

Vision-RL² is a method for improving fine-grained perception in multimodal large language models by using region-level reinforcement learning to compress visual tokens and enhance performance across multiple benchmarks without full model fine-tuning.

0 favorites 0 likes

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Hugging Face Daily Papers · 2026-09-17 Cached

VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.

0 favorites 0 likes

JEPA-Anything: Learning Predictive Models across Different Worlds

Hugging Face Daily Papers · 2026-09-17 Cached

JEPA-Anything presents a domain-agnostic framework based on orthogonal predictive factorization for learning predictive models across diverse systems like vision, biology, and control, with demonstrated improvements and experimental validation.

0 favorites 0 likes

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Hugging Face Daily Papers · 2026-09-17 Cached

Video DeltaNet presents a hybrid attention mechanism combining Softmax and linear attention to enhance efficiency in video generation models, achieving a 14.5x speedup over baseline methods.

0 favorites 0 likes

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

Hugging Face Daily Papers · 2026-09-17 Cached

UFO is a unified framework for simultaneous evaluation of omni-condition alignment in multi-modal image generation. It introduces an Atomized Chain-of-Evaluation paradigm and UFO-Bench benchmark.

0 favorites 0 likes

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Hugging Face Daily Papers · 2026-09-17 Cached

Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.

0 favorites 0 likes

What Does Privileged Information Add to On-Policy Self-Distillation?

Hugging Face Daily Papers · 2026-09-17 Cached

The paper investigates the contribution of privileged information in on-policy self-distillation for language models, finding that reference-free distillation accounts for most improvements, with limited additional benefits from privileged references.

0 favorites 0 likes

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

Hugging Face Daily Papers · 2026-09-17 Cached

When2Think is a post-training framework that dynamically allocates computation in large reasoning models based on problem difficulty, improving accuracy-efficiency trade-offs on mathematical benchmarks.

0 favorites 0 likes

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Hugging Face Daily Papers · 2026-09-17 Cached

The paper proposes RetireOPD, a method for training multi-turn agents using reinforcement learning with self-retiring on-policy distillation, improving performance on ALFWorld and WebShop benchmarks.

0 favorites 0 likes

prism-ml/Ternary-Bonsai-2-27B-mlx-2bit

Hugging Face Models Trending · 2026-09-16 Cached

Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.

0 favorites 0 likes

prism-ml/Ternary-Bonsai-2-27B-gguf

Hugging Face Models Trending · 2026-09-16 Cached

Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.

0 favorites 0 likes

AlexWortega/openjev

Hugging Face Models Trending · 2026-09-16 Cached

openjev is a model based on Qwen3.5 trained for entailment tasks, enabling applications in reranking, grading, and real-time game playing, with the v2 version adding multi-modal capabilities and improved zero-shot performance.

0 favorites 0 likes

XingChen-AGI/Xing4.0-29B-A4B

Hugging Face Models Trending · 2026-09-16 Cached

Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.

0 favorites 0 likes

Cactus-Compute/needle3

Hugging Face Models Trending · 2026-09-16 Cached

Needle 3 is a compact AI foundation model optimized for edge devices like mobiles and wearables, offering tool calling, structured extraction, and text embedding in a single 8-29 MB file.

0 favorites 0 likes

harshatheg/Qwen-2.5-1B-RLCD

Hugging Face Models Trending · 2026-09-16 Cached

A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.

0 favorites 0 likes

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

Hugging Face Daily Papers · 2026-09-16 Cached

ALPINE introduces an ultra-lightweight spatial-relational architecture for few-shot image classification that achieves accuracy gains with fewer parameters, faster convergence, and better robustness compared to baselines like Prototypical Networks and MAML.

0 favorites 0 likes

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Hugging Face Daily Papers · 2026-09-16 Cached

This paper introduces GAVEL, a framework that uses graph world models to verify and repair long-horizon LLM planning for robotic tasks, significantly improving success rates and efficiency in simulations.

0 favorites 0 likes

FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

Hugging Face Daily Papers · 2026-09-16 Cached

FRAUDSkill is a structured frozen-weight adaptation framework for audio anti-fraud detection that optimizes external skill programs without modifying the underlying audio-language model, achieving higher accuracy and reduced invalid outputs.

0 favorites 0 likes

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

Hugging Face Daily Papers · 2026-09-16 Cached

This paper evaluates MiniMax-H3, an omni-modal generative model, by introducing a comprehensive framework to assess its reasoning about the physical world through multimodal inputs. The evaluation reveals that video-based decision reasoning performs best, while audio-based disambiguation reasoning is the weakest.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback