Articles from HuggingFace
SoL-Pi introduces a method for recursively scaling auto-research loops in coding agents, achieving significant token and cost reductions while maintaining performance on benchmarks.
Vision-RL² is a method for improving fine-grained perception in multimodal large language models by using region-level reinforcement learning to compress visual tokens and enhance performance across multiple benchmarks without full model fine-tuning.
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.
JEPA-Anything presents a domain-agnostic framework based on orthogonal predictive factorization for learning predictive models across diverse systems like vision, biology, and control, with demonstrated improvements and experimental validation.
Video DeltaNet presents a hybrid attention mechanism combining Softmax and linear attention to enhance efficiency in video generation models, achieving a 14.5x speedup over baseline methods.
UFO is a unified framework for simultaneous evaluation of omni-condition alignment in multi-modal image generation. It introduces an Atomized Chain-of-Evaluation paradigm and UFO-Bench benchmark.
Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.
The paper investigates the contribution of privileged information in on-policy self-distillation for language models, finding that reference-free distillation accounts for most improvements, with limited additional benefits from privileged references.
When2Think is a post-training framework that dynamically allocates computation in large reasoning models based on problem difficulty, improving accuracy-efficiency trade-offs on mathematical benchmarks.
The paper proposes RetireOPD, a method for training multi-turn agents using reinforcement learning with self-retiring on-policy distillation, improving performance on ALFWorld and WebShop benchmarks.
Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.
Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.
openjev is a model based on Qwen3.5 trained for entailment tasks, enabling applications in reranking, grading, and real-time game playing, with the v2 version adding multi-modal capabilities and improved zero-shot performance.
Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.
Needle 3 is a compact AI foundation model optimized for edge devices like mobiles and wearables, offering tool calling, structured extraction, and text embedding in a single 8-29 MB file.
A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.
ALPINE introduces an ultra-lightweight spatial-relational architecture for few-shot image classification that achieves accuracy gains with fewer parameters, faster convergence, and better robustness compared to baselines like Prototypical Networks and MAML.
This paper introduces GAVEL, a framework that uses graph world models to verify and repair long-horizon LLM planning for robotic tasks, significantly improving success rates and efficiency in simulations.
FRAUDSkill is a structured frozen-weight adaptation framework for audio anti-fraud detection that optimizes external skill programs without modifying the underlying audio-language model, achieving higher accuracy and reduced invalid outputs.
This paper evaluates MiniMax-H3, an omni-modal generative model, by introducing a comprehensive framework to assess its reasoning about the physical world through multimodal inputs. The evaluation reveals that video-based decision reasoning performs best, while audio-based disambiguation reasoning is the weakest.