Tag
This paper proposes GUPO, a Gradient Uncertainty-aware Policy Optimization method that improves post-training of large language models by modeling gradient uncertainties using a Bayesian approach to handle conflicts among group gradients.
This article promotes the book 'Before Machine Learning Volume 3: Probability and Statistics for AI,' which uses the Antwerp Diamond Heist narrative to teach statistical concepts for AI, such as Bayesian reasoning and Markov Chains, with practical code examples.
CalibratedRubric is a task-adaptive framework for building compact, measurable rubric banks for open-ended LLM evaluation, using Bayesian measurability filtering and IRT-based selection to improve human-gold agreement and rank fidelity across financial, healthcare, general, and legal benchmarks.
This paper presents BayesPO, a Bayesian prompt optimization framework using gradient-guided discrete MCMC with parallel tempering, achieving improved accuracy on instruction-induction tasks.
Introduces Bayesian Manifold Curriculum (BMC), an adaptive curriculum learning method for LLMs that leverages the model's latent geometry to allocate training effort across diverse problem types, improving efficiency beyond traditional difficulty-based curricula.