importance-sampling

Tag

Cards List
#importance-sampling

Group Adaptive Clipping Policy Optimization

arXiv cs.LG · 2026-09-02 Cached

The paper proposes Group Adaptive Clustering Policy Optimization (GAPO), a plug-in modification to GRPO methods that adapts the clipping boundary to rollout advantage, improving Pass@1 and Pass@k on math reasoning and coding benchmarks.

0 favorites 0 likes
#importance-sampling

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

arXiv cs.LG · 2026-07-28 Cached

This paper introduces DiffTilt, a distributional framework that exponentially tilts a diffusion model-induced joint distribution over environments and executions to efficiently discover rare safety-critical failures in autonomous and cyber-physical systems, outperforming conditional sampling strategies on ARCH-COMP benchmarks and a new tractor-trailer benchmark.

0 favorites 0 likes
#importance-sampling

Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling

arXiv cs.LG · 2026-07-20 Cached

This paper introduces Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct method for sampling from unnormalized Boltzmann densities by learning stochastic Hamiltonian-style paths and correcting them using non-equilibrium work.

0 favorites 0 likes
#importance-sampling

@xennygrimmato_: if you’re wondering how token-level rejection sampling works in this paper, here’s how they do it: M_t = max_v [ pi_the…

X AI KOLs Timeline · 2026-07-11 Cached

Explains token-level rejection sampling for RLHF/PPO, where importance ratio M_t is the maximum over vocabulary and tokens are accepted with Bernoulli sampling based on w_t / M_t.

0 favorites 0 likes
#importance-sampling

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

arXiv cs.LG · 2026-07-09 Cached

This paper introduces Unbounded Positive Asymmetric Optimization (UP), a universal plug-and-play objective that resolves the exploration-stability dilemma in RL-based LLM training by anchoring the policy with stop-gradient, enabling unclipped gradients for positive advantages while clipping negative ones.

0 favorites 0 likes
← Back to home

Submit Feedback