clipping

Tag

Cards List
#clipping

CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions

arXiv cs.LG · 2026-08-10 Cached

CertBind introduces a multiscale theory for certifiable composition of frozen multimodal encoder connectors, enabling certified retrieval decisions with fallback and abstain mechanisms. Evaluations on a shared route show it can recover native CLIP retrieval performance while preserving no-harm on passing branches.

0 favorites 0 likes
#clipping

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Hugging Face Daily Papers · 2026-07-21 Cached

Introduces Staleness-Adaptive Trust Regions (SAT) to stabilize asynchronous reinforcement learning by adaptively controlling update intervals based on staleness. Evaluated on a decoupled asynchronous RL setup using Qwen3-30B-A3B-Base, achieving improved results.

0 favorites 0 likes
#clipping

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization

Hugging Face Daily Papers · 2026-07-11 Cached

This paper reveals that PPO-Clipping's use of Euclidean metric causes exploration collapse in LLM RL, and proposes Riemannian Isometric Policy Optimization (RIPO) to ensure geometrically consistent policy updates, achieving up to 60% improvement over GRPO on AIME24.

0 favorites 0 likes
#clipping

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

arXiv cs.LG · 2026-07-09 Cached

This paper introduces Unbounded Positive Asymmetric Optimization (UP), a universal plug-and-play objective that resolves the exploration-stability dilemma in RL-based LLM training by anchoring the policy with stop-gradient, enabling unclipped gradients for positive advantages while clipping negative ones.

0 favorites 0 likes
#clipping

@johnschulman2: PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective f…

X AI KOLs Following · 2026-06-18 Cached

This paper reveals that the clipping mechanism in PPO and GRPO biases entropy in RLVR for LLMs: clip-low increases entropy, clip-high decreases it. The authors prove that standard clipping reduces entropy even with random rewards, and show that adjusting clip-low can prevent entropy collapse and promote exploration.

0 favorites 0 likes
#clipping

MuCon: Clipped Muon Updates for LLM Training

arXiv cs.LG · 2026-05-27 Cached

This paper introduces MuCon, a clipped-Muon optimizer for LLM training that applies singular-value clipping instead of full polarization, preserving smaller singular values while clipping only the largest ones. It explores approximations to avoid full SVD, including polar/absolute-value formulas and rational Newton filters, noting numerical challenges near the threshold.

0 favorites 0 likes
#clipping

How clips ate the internet

The Verge · 2026-05-26 Cached

The Vergecast episode explores how social media feeds are dominated by clipped content and algorithmic brute force, and also reviews the new Fitbit Air fitness tracker and discusses smart glasses as a product category.

0 favorites 0 likes
← Back to home

Submit Feedback