Articles from HuggingFace
LimiX-2 is a pretrained foundation model for structured data that uses contextual mechanism networks to achieve #1 on major tabular benchmarks, supporting multiple tasks without task-specific parameter updates.
This paper proposes EventEgoHands++, a framework for event-based 3D hand mesh reconstruction from an egocentric viewpoint, incorporating hand detection and adaptive attention, and introduces a new real-world dataset EEH-R for training and evaluation.
FLAT is a representation pre-training framework that jointly optimizes a shared multimodal encoder with downstream decoders for text-to-image and image-to-text tasks, using flexible-length aligned 1D sequences to enable cross-modal retrieval and generation with state-of-the-art results.
This paper introduces ImpossibleRubrics, an open-source evaluation framework that uses fine-grained, adversarial rubrics to stress-test large language models, aiming to reduce evaluation bias and expose hidden failure modes.
ScienceBuddy introduces a recursive-in-recursive self-improvement paradigm for interactive scientific agents, enabling continual evolution through researcher collaboration and feedback.
This paper surveys the use of foundation models in game AI across roles like playing, modeling, design, and evaluation, highlighting transferability challenges and the need for game-specific validation.
Qwen-Image-2.1 is an open-source unified text-to-image and image editing model with 7B parameters, featuring efficient architecture, transparency support, and versatile editing capabilities.
Bellman Policy Optimization (BPO) is a critic-free reinforcement learning method that reformulates Policy Mirror Descent using the Bellman equation for autoregressive generation with terminal rewards, improving mathematical reasoning in large language models.
Introduces MoME, a context-aware memory mechanism for LLMs that uses a mixture of slots to handle token polysemy, improving over baselines in pretraining experiments.
EvoOntology introduces a self-evolving ontology layer for data agents, encapsulated as an MCP server, to bridge the agent-data gap and improve performance on heterogeneous data tasks as shown in benchmarks.
This paper assesses the generalization of nnU-Net for brain tumor segmentation in the BraTS-GoAT 2026 challenge, reporting performance metrics and analyzing failure cases.
HypoEvolve introduces a framework using genetic algorithms to coordinate multi-agent LLMs for scientific hypothesis discovery, demonstrated through drug repurposing in cancer research with improved performance over baselines.
HarnessVLN is a zero-shot, training-free framework for embodied navigation that unifies perception, retrieval, grounding, navigation, recovery, and termination through a unified tool interface, achieving state-of-the-art results on benchmarks like R2R and RxR.
ModularRSI introduces a modular and generalizable framework for recursive self-improvement in AI agent harnesses, using contrastive learning across tasks to evolve modules independently and enhance performance on unseen tasks.
The paper introduces Gavel, a method that elicits native skill routing from frozen LLMs via linear projections, enabling efficient tool selection without context overload and outperforming existing pipelines on benchmarks.
This paper decomposes transformer representation updates into parallel and perpendicular components to study evolution geometry, linking it to editing robustness, compression diagnosis, and training improvements.
ModaLens is a paired image-swap audit that measures how report availability reduces image sensitivity in medical vision-language models, demonstrated using MedGemma-27B on the MIMIC-CXR dataset.
This paper investigates the losslessness of Orthrus, a hybrid autoregressive-diffusion model for inference acceleration, finding that it requires high numerical precision (FP32) for exact trajectory matching, while BF16 divergence does not impair downstream performance.
LynnReal-Omni is a unified multimodal video diffusion framework that integrates agentic visual controls with high-fidelity generation and real-time acceleration for stable, controllable video creation.
This paper proposes an exploration-guided prompt scaffolding framework that dynamically adjusts training prompts for multimodal reinforcement learning, achieving up to 9.7% relative improvement in performance on benchmarks.