Tag
The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method for LLMs that uses comparison oracles to avoid likelihood displacement. It includes theoretical guarantees and experimental improvements over existing methods.
Introduces HERO, an LLM-based program optimizer that overcomes the weakest-link effect by generating and recombining heterogeneous atomic edits, achieving faster convergence and higher scores across algorithmic, game, agentic, and robotic domains.
This paper introduces a model-free deep learning method for solving high-dimensional nonlinear partial differential equations with unknown coefficients, using zeroth-order derivative estimators derived from perturbed Monte Carlo trajectories. The approach avoids automatic differentiation, provides theoretical error bounds, and demonstrates competitive performance in numerical experiments.
This paper reveals that zeroth-order fine-tuning of LLMs is dominated by a single decoding layer, which can be identified by activation outliers, and fine-tuning only that layer matches or exceeds full-model fine-tuning with up to 4.52x speedup.