Tag
This paper introduces TaH2, an adaptive looped transformer that improves test-time scaling by dynamically allocating extra iterations to beneficial tokens, achieving a 53% improvement in accuracy-compute slope over baselines on benchmarks like AIME.
This paper introduces Planned Test-Time Scaling (PTTS), a method that coordinates reasoning branches to enhance performance on challenging tasks, achieving significant gains over repeated sampling in mathematical reasoning benchmarks.
Hill Sampling is a simple test-time scaling method that repeatedly samples edits to the best verified program using frozen LLMs, achieving state-of-the-art results on algorithmic problems like circle packing and Erdős' minimum-overlap problem.
This paper introduces LLM-as-an-Improver, a method that uses verification feedback to generate improved candidate solutions for LLMs, enhancing performance beyond initial candidate pools.
This paper explores the impact of different candidate-generation schedules on the energy consumption and performance of large language models during test-time scaling, demonstrating that larger batch sizes reduce energy use and latency.
Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.
This paper investigates test-time scaling for scientific equation discovery, formulating it as an iterative search process and finding that search width is the dominant allocation parameter for improving performance and efficiency under compute budgets.
The paper introduces Retrieval-Grounded Voting (RGV) to address the limitations of confidence-based voting in multi-turn search agents by using lexical overlap with retrieved documents, achieving up to 5.4% accuracy gains.
Prefix Sliding reduces memory costs during long reasoning by discarding unimportant intermediate tokens, enabling efficient test-time scaling without retraining, achieving up to 3x speedup in existing models.
Gambit improves reasoning model efficiency by using thought-level beam search to dynamically allocate compute to promising reasoning traces under fixed hardware budgets, yielding significant accuracy and throughput gains.
This paper introduces the 'crystallization problem' for evaluating reusable memory in text-to-SQL systems, showing that storing verified corrected queries in a per-database bank improves held-out first-attempt accuracy by 4.34 points on BIRD, capturing 44.4% of the headroom provided by on-demand repair. Controlled interventions identify database-specific content as the main driver.
A new verifier-free breadth-depth refinement framework improves LLM reasoning at test time by sampling multiple rollouts, iteratively refining each via self-critique, and aggregating with majority voting. It consistently outperforms greedy decoding, majority voting, and verifier-based selection across several math benchmarks and open-weight models.
This paper introduces RL^2, an adaptive inference-time steering framework for Vision-Language-Action models that uses offline RL on latent representations to compose action flows, activating steering only when failure is predicted. It achieves up to +17.3% success rate improvements on SIMPLER and PolaRiS benchmarks and demonstrates real-world transfer.
This paper presents MetaEvolve, a framework that uses reinforcement learning to train LLMs in self-evolution meta-skills for iterative refinement, achieving significant improvements on coding benchmarks.
Introduces Surrogate Latent Policy Optimization (SLPO) to apply outcome-reward RL to autoregressive latent reasoners, enabling test-time scaling and variable-horizon policies that improve accuracy on harder instances.
This paper re-evaluates the methodology of automatic harness evolution for LLM agents, highlighting that its gains may stem from additional test-time search rather than improved harness design, and that evaluation on the same benchmark risks overfitting. Experiments show that harness evolution does not consistently outperform simpler test-time scaling methods.
Introduces the Energy-guided Recursive Model (ERM), which uses Hopfield energies to guide selection among recursive reasoning trajectories, achieving state-of-the-art performance on Sudoku, Pencil Puzzle Bench, and Maze tasks.
This paper rethinks how automatic harness evolution for agents should be evaluated, showing that gains may be due to increased compute rather than genuine improvements, and that evolved harnesses transfer poorly to unseen tasks.
This paper investigates test-time scaling techniques for small open vision-language models (≤7B parameters) on the multilingual visual MCQ benchmark EXAMS-V, finding that inference budget and parseability matter more than complex search or verification methods. The best configuration achieves 84.1% on the ImageCLEF 2026 test split, ranking first on the leaderboard.
KV-PRM introduces a process reward model that leverages KV-cache transfer to avoid re-encoding, achieving up to 5000x FLOP reduction while maintaining or improving performance on reasoning benchmarks.