test-time-scaling

Tag

Cards List
#test-time-scaling

Improving Test-Time Scaling with Adaptive Looped Transformers

Hugging Face Daily Papers ↗ · yesterday Cached

This paper introduces TaH2, an adaptive looped transformer that improves test-time scaling by dynamically allocating extra iterations to beneficial tokens, achieving a 53% improvement in accuracy-compute slope over baselines on benchmarks like AIME.

0 favorites 0 likes
#test-time-scaling

Planned Test-Time Scaling with Coordinated Reasoning Paths

arXiv cs.CL ↗ · 5d ago Cached

This paper introduces Planned Test-Time Scaling (PTTS), a method that coordinates reasoning branches to enhance performance on challenging tasks, achieving significant gains over repeated sampling in mathematical reasoning benchmarks.

0 favorites 0 likes
#test-time-scaling

Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training

arXiv cs.LG ↗ · 6d ago Cached

Hill Sampling is a simple test-time scaling method that repeatedly samples edits to the best verified program using frozen LLMs, achieving state-of-the-art results on algorithmic problems like circle packing and Erdős' minimum-overlap problem.

0 favorites 0 likes
#test-time-scaling

LLM-as-an-Improver: Turning Verification into Better Candidates

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper introduces LLM-as-an-Improver, a method that uses verification feedback to generate improved candidate solutions for LLMs, enhancing performance beyond initial candidate pools.

0 favorites 0 likes
#test-time-scaling

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

arXiv cs.LG ↗ · 2026-09-18 Cached

This paper explores the impact of different candidate-generation schedules on the energy consumption and performance of large language models during test-time scaling, demonstrating that larger batch sizes reduce energy use and latency.

0 favorites 0 likes
#test-time-scaling

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Hugging Face Daily Papers ↗ · 2026-09-11 Cached

Dynin-Robotics is an omnimodal unified diffusion model that integrates vision, language, and action for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.

0 favorites 0 likes
#test-time-scaling

Test-Time Scaling for Scientific Equation Discovery

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper investigates test-time scaling for scientific equation discovery, formulating it as an iterative search process and finding that search width is the dominant allocation parameter for improving performance and efficiency under compute budgets.

0 favorites 0 likes
#test-time-scaling

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

arXiv cs.AI ↗ · 2026-08-26 Cached

The paper introduces Retrieval-Grounded Voting (RGV) to address the limitations of confidence-based voting in multi-turn search agents by using lexical overlap with retrieved documents, achieving up to 5.4% accuracy gains.

0 favorites 0 likes
#test-time-scaling

Prefix Sliding for efficient test-time scaling

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

Prefix Sliding reduces memory costs during long reasoning by discarding unimportant intermediate tokens, enabling efficient test-time scaling without retraining, achieving up to 3x speedup in existing models.

0 favorites 0 likes
#test-time-scaling

Thought-Level Beam Search for Reasoning

Hugging Face Daily Papers ↗ · 2026-08-11 Cached

Gambit improves reasoning model efficiency by using thought-level beam search to dynamically allocate compute to promising reasoning traces under fixed hardware budgets, yielding significant accuracy and throughput gains.

0 favorites 0 likes
#test-time-scaling

From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper introduces the 'crystallization problem' for evaluating reusable memory in text-to-SQL systems, showing that storing verified corrected queries in a per-database bank improves held-out first-attempt accuracy by 4.34 points on BIRD, capturing 44.4% of the headroom provided by on-demand repair. Controlled interventions identify database-specific content as the main driver.

0 favorites 0 likes
#test-time-scaling

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv cs.AI ↗ · 2026-08-07 Cached

A new verifier-free breadth-depth refinement framework improves LLM reasoning at test time by sampling multiple rollouts, iteratively refining each via self-critique, and aggregating with majority voting. It consistently outperforms greedy decoding, majority voting, and verifier-based selection across several math benchmarks and open-weight models.

0 favorites 0 likes
#test-time-scaling

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

This paper introduces RL^2, an adaptive inference-time steering framework for Vision-Language-Action models that uses offline RL on latent representations to compose action flows, activating steering only when failure is predicted. It achieves up to +17.3% success rate improvements on SIMPLER and PolaRiS benchmarks and demonstrates real-world transfer.

0 favorites 0 likes
#test-time-scaling

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

arXiv cs.CL ↗ · 2026-07-27 Cached

This paper presents MetaEvolve, a framework that uses reinforcement learning to train LLMs in self-evolution meta-skills for iterative refinement, achieving significant improvements on coding benchmarks.

0 favorites 0 likes
#test-time-scaling

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Hugging Face Daily Papers ↗ · 2026-07-22 Cached

Introduces Surrogate Latent Policy Optimization (SLPO) to apply outcome-reward RL to autoregressive latent reasoners, enabling test-time scaling and variable-horizon policies that improve accuracy on harder instances.

0 favorites 0 likes
#test-time-scaling

Rethinking the Evaluation of Harness Evolution for Agents

arXiv cs.AI ↗ · 2026-07-15 Cached

This paper re-evaluates the methodology of automatic harness evolution for LLM agents, highlighting that its gains may stem from additional test-time search rather than improved harness design, and that evaluation on the same benchmark risks overfitting. Experiments show that harness evolution does not consistently outperform simpler test-time scaling methods.

0 favorites 0 likes
#test-time-scaling

Energy-guided Recursive Model

arXiv cs.LG ↗ · 2026-07-14 Cached

Introduces the Energy-guided Recursive Model (ERM), which uses Hopfield energies to guide selection among recursive reasoning trajectories, achieving state-of-the-art performance on Sudoku, Pencil Puzzle Bench, and Maze tasks.

0 favorites 0 likes
#test-time-scaling

Rethinking the Evaluation of Harness Evolution for Agents

Hugging Face Daily Papers ↗ · 2026-07-14 Cached

This paper rethinks how automatic harness evolution for agents should be evaluated, showing that gains may be due to increased compute rather than genuine improvements, and that evolved harnesses transfer poorly to unseen tasks.

0 favorites 0 likes
#test-time-scaling

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

arXiv cs.CL ↗ · 2026-07-13 Cached

This paper investigates test-time scaling techniques for small open vision-language models (≤7B parameters) on the multilingual visual MCQ benchmark EXAMS-V, finding that inference budget and parseability matter more than complex search or verification methods. The best configuration achieves 84.1% on the ImageCLEF 2026 test split, ranking first on the leaderboard.

0 favorites 0 likes
#test-time-scaling

KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling

arXiv cs.AI ↗ · 2026-07-13 Cached

KV-PRM introduces a process reward model that leverages KV-cache transfer to avoid re-encoding, achieving up to 5000x FLOP reduction while maintaining or improving performance on reasoning benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback