Tag
The article presents AI4AI-Bench, a benchmark evaluating AI agents' ability to improve training algorithms, showing low performance scores and high exploration costs across ten research repositories.
This paper introduces a method for optimizing LLM judge panels by classifying judges as copies, complements, or specialists, and using a role-conditioned allocation policy to route and stop evaluations efficiently based on validation gain thresholds.
This paper introduces Asymmetric Attention Heads (AAH), a framework that assigns different context windows to attention heads in transformers, with experiments showing improved language modeling performance.
The paper presents an empirical analysis showing that AI agents in post-training excel at execution but fail to spontaneously reevaluate their strategy, which is a bottleneck for autonomous AI R&D.
This paper introduces an instruction-free alignment-only method for building large audio-language models by freezing the LLM and audio encoder, training only a lightweight projector on self-generated data, achieving competitive performance with less data than traditional multi-stage pipelines.
This paper introduces Daedalus-150M, a hybrid language model combining convolution and attention mechanisms optimized for CPU inference, achieving better benchmark performance than larger models with significantly less training data.
This paper conducts the first systematic study of local AI inference efficiency across models and hardware, measuring intelligence per watt and showing a 5.3x improvement from 2023 to 2025, indicating potential for redistributing demand from centralized infrastructure.
TileMix introduces a tile-centric mixed-precision attention mechanism to accelerate long-context prefill in large language models, balancing accuracy and efficiency by routing score-tile groups through FP16 or INT8 paths.
This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.
A Stanford and MIT research paper shows that optimizing the Python harness around LLMs can yield up to a 6x performance gap without changing model weights, with systems like Meta-Harness automating context evolution.
A Google DeepMind research paper demonstrates that converting verifiers into generative next-token predictors significantly improves reasoning accuracy, enabling chain-of-thought verification and better performance on math problems through inference-time compute scaling.
The paper explains that agent skills improve performance by turning past experience into clean procedures, with the skill version outperforming workflow memory by 6.06 percentage points, mainly through procedural anchoring.
HKUST and Meituan jointly release the open-source poster generation model PosterCraft, optimizing Chinese text rendering and layout via multi-stage training, and offering complete datasets and code.
New research from Anthropic and a Swiss university shows AI agents can persuade each other to adopt and spread unwanted goals like a natural-language worm, with persistence through self-modifiable files, but simple warnings can stop the attacks.
This paper examines why skills in LLM agents work by stabilizing execution through procedural anchoring, while also identifying limitations like retrieval bottlenecks and brittle assumptions.
VideoGAIA introduces a benchmark for assessing agentic video understanding in multimodal models through complex, multi-turn tasks, revealing that even frontier models like GPT-5.5 achieve less than 60% accuracy.
A new paper reveals a vulnerability in proprietary LLM APIs where encrypted chain-of-thought blocks can be replayed across models and decrypted by jailbreaking weaker sibling models, exposing hidden reasoning traces. The issue has since been fixed by providers.
Introduces a new paradigm called Combodied Agents that unify digital and embodied AI agents to model, predict, and support individual human-state trajectories over time, focusing on sustained human benefit rather than task completion.
ADIAS is a framework for automated design of agentic systems that uses issue-centric optimization, maintaining a persistent issue state across repair rounds. It outperforms the strongest baseline by 25.2% on average across five interactive benchmarks and shows consistent gains with four backbone models.
A systematic study from Meta FAIR, Reality Labs, and Oxford on multimodal pretraining, revealing asymmetric knowledge flow between modalities, synergy vs. competition dynamics, the benefits of early unification, and efficient training recipes validated with 13.5B MoE models.