Tag
This paper introduces Meta-Skill, a method letting a frozen Builder model learn reusable principles from Target execution feedback to construct better agent harnesses for unseen tasks, improving performance by 8.95 points on Harness-Bench and NewtonBench. It suggests a path toward system-level self-improvement by learning to build better environments rather than changing model weights.
The paper presents Spectral Feedback, an algorithm that enhances test-time alignment for discrete diffusion models in protein inverse folding by iteratively selecting edit-positions using sparse Fourier representations, resulting in improved performance for reward maximization.
This paper introduces GradCuit, a method for test-time latent reasoning that inserts optimizable latent states at a selected Transformer layer. It achieves 64.5% average accuracy across five backbones and three reasoning benchmarks, outperforming chain-of-thought prompting and showing improved robustness and interpretability.
Introduces SEVRA, a selective verification controller for budget-aware reasoning that decides when to accept a model's initial answer versus spending extra compute on verification, improving accuracy and reducing unnecessary tokens on benchmarks like MATH500 and GSM8K.