math-benchmarks

Tag

Cards List
#math-benchmarks

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

arXiv cs.LG · 23h ago Cached

The paper identifies the Matthew Effect in reinforcement learning for large language models, where easy problems improve more than hard ones, and introduces Never Give Up (NGU), an adaptive sampling method to allocate more compute to hard problems, demonstrating performance gains on math and coding benchmarks.

0 favorites 0 likes
#math-benchmarks

Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

Hugging Face Daily Papers · 3d ago Cached

This paper introduces Lightning Weave, a framework that combines separately trained reasoning capabilities into one efficient model using on-policy distillation, enhancing accuracy and reducing token usage in math and code tasks.

0 favorites 0 likes
#math-benchmarks

Concretized Proposition Prompting Resolves Composition-Knowledge Dichotomy in Large Language Models

arXiv cs.AI · 2026-07-10 Cached

This paper proposes Concretized Proposition Prompting (CPP), a framework that resolves the composition-knowledge dichotomy in LLMs by explicitly concretizing propositions relevant to questions, significantly enhancing reasoning performance especially in medical and math benchmarks.

0 favorites 0 likes
#math-benchmarks

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

arXiv cs.CL · 2026-07-02 Cached

This paper introduces DASH, a method that uses intermediate answer commitments within reasoning traces to assign segment-level credit, reducing overthinking behaviors and improving accuracy on competition-level math benchmarks.

0 favorites 0 likes
#math-benchmarks

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

Hugging Face Daily Papers · 2026-06-17 Cached

This paper proposes Trajectory-Augmented Policy Optimization (TAPO), which constructs micro-reflective correction trajectories using the model's own correct and incorrect rollouts to improve reasoning in large language models, outperforming standard self-distillation methods on math benchmarks.

0 favorites 0 likes
#math-benchmarks

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning

arXiv cs.AI · 2026-05-26 Cached

This paper formalizes reasoning redundancy in LLMs as the fraction of trailing steps that can be truncated without affecting correctness, quantifying 61-93% redundancy across frontier models and proving that redundancy is a structural consequence of length-agnostic outcome rewards.

0 favorites 0 likes
← Back to home

Submit Feedback