Articles from HuggingFace
The paper demonstrates a method to extract proprietary AI model architecture details by using child-like framing and multi-turn prompts to exploit alignment cues in frontier AI models.
This paper introduces Lightning Weave, a framework that combines separately trained reasoning capabilities into one efficient model using on-policy distillation, enhancing accuracy and reducing token usage in math and code tasks.
The E2A-Bench benchmark for financial chart reasoning reveals that vision-language models often fail to maintain traceable evidence-to-action reliability, highlighting the need to evaluate the full reasoning chain rather than single hallucination scores.
AlayaVista is a camera-controllable streaming video world model that decouples panoramic scene evolution from perspective video synthesis, supported by the large-scale MUGEN dataset for high-fidelity interactive world modeling.
Yandex has released AliceAI-Foundation-80B-A3B-Base, a foundation language model with a hybrid architecture and MoE layers, featuring 80 billion parameters (3 billion active per token), trained from scratch and excelling in Russian factual knowledge tasks.
SiliconBench evaluates nine Apple Silicon LLM serving engines on speed, memory, and fidelity, finding that explicit memory budgets don't ensure headroom and only a few stacks meet all criteria for concurrency scaling and model coverage.
The paper presents StepAudio3Realtime, an audio-language foundation model for real-time spoken interaction, using a continuous listen-converse-think-act loop with Think-While-Speaking to achieve deep reasoning and low latency, with top-tier performance on benchmarks like MMSU and Full-Duplex Bench.
The paper proposes the Convergent Emergence Hypothesis, stating that few-shot in-context learning emerges with a common cross-modality difficulty profile, and provides empirical support through experiments on six modalities, showing correlated effects in five.
Realtime-Venus is a proactive full-duplex interaction system that integrates continuous perception, conversational control, and native speech generation through two separately trained 9B models, achieving high scores on audio and video benchmarks.
Atria Dawn Preview is a new agentic AI model from Shanghai AI Lab, built on a 744B-parameter MoE GLM-5.2 foundation, designed for research and engineering tasks requiring continuous environmental understanding and tool use.
This article details the YuE2 model repackaged for ComfyUI, enabling audio and music generation workflows. It provides model files based on MERT-v2-FullSong and SheetSage2 for easy integration into ComfyUI projects.
The paper introduces SteerDuplex, a full-duplex speech dialogue model with steerable attributes like tone and persona, and presents SteerBench for evaluating spoken steerability, demonstrating improvements over baselines using reinforcement learning.
RelateAnything is a lightweight, real-time open-vocabulary relation prediction model that accepts arbitrary predicate vocabularies and region sources, trained on a large geometrically verified dataset and evaluated on new cross-dataset benchmarks, showing significant performance gains over comparable methods.
StepAudio 3 Music introduces a large-scale, long-form music generation model with explicit musical planning via ABC-CoT, achieving high scores in audio quality and similarity metrics compared to other systems.
This paper details the development of Sophea, a production bilingual Greek-English automatic speech recognition system, using iterative training, data filtering, and model ensembling to meet quality gates and achieve competitive benchmark results.
Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.
ZGCM-1 is a 7B open foundation model trained from scratch with extreme efficiency, combining internal reasoning and external tool use for math and agentic search tasks, achieving competitive performance with much larger models like Qwen3-235B-A22B and GLM-5.1.
A general-purpose agent directly controls a physical robot by interpreting visuals, writing executable programs, and revising actions based on physical feedback across diverse manipulation tasks, achieving high success rates without task-specific training.
StepAudio 3 Gen is a general-purpose audio generation model that unifies text-to-speech, voice design, sound effects, and music within a single discrete autoregressive framework using residual vector quantization tokens.
SAS introduces a gated sparse attention mechanism that optimizes context ranking end-to-end with language modeling loss, improving performance in reasoning and long-context tasks under tight attention budgets.