mathematical-reasoning

Tag

Cards List
#mathematical-reasoning

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv cs.AI · 2d ago Cached

A new verifier-free breadth-depth refinement framework improves LLM reasoning at test time by sampling multiple rollouts, iteratively refining each via self-critique, and aggregating with majority voting. It consistently outperforms greedy decoding, majority voting, and verifier-based selection across several math benchmarks and open-weight models.

0 favorites 0 likes
#mathematical-reasoning

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

arXiv cs.AI · 2d ago Cached

This paper proposes Hyper-ES, a subspace-based evolution strategy framework for LLM reasoning that obtains descent directions via lightweight gradient-based fine-tuning and then uses CMA-ES to merge layer-wise DARE-TIES coefficients, consistently outperforming GRPO-LoRA while requiring fewer gradient updates.

0 favorites 0 likes
#mathematical-reasoning

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

arXiv cs.AI · 6d ago Cached

Presents AMTFV, an agentic framework that decouples mathematical verification modeling from execution via a Mathematical Tool Flow interface, improving LLM answer verification and revision on five challenging math datasets.

0 favorites 0 likes
#mathematical-reasoning

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

arXiv cs.CL · 2026-07-31 Cached

This paper introduces LEEPS, a latent-guided explore-exploit prompt sampler for efficient reinforcement learning with verifiable rewards (RLVR) in LLMs. It adaptively balances reuse of informative prompts and exploration of uncertain ones, improving reasoning benchmark scores by 2.6-3.7% over baselines while adding only ~2 seconds of overhead per training step.

0 favorites 0 likes
#mathematical-reasoning

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

arXiv cs.AI · 2026-07-31 Cached

This paper investigates internal representational differences between RL and SFT fine-tuned models on mathematical reasoning, finding that RL models exhibit more linearly separable hidden states and hierarchical layer importance. Token allocation variability under repeated sampling suggests training pipeline dependence rather than RL vs SFT alone.

0 favorites 0 likes
#mathematical-reasoning

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Hugging Face Daily Papers · 2026-07-30 Cached

This paper introduces β-OPSD, a generalization of on-policy self-distillation that frames it as a policy optimization family with a controllable KL penalty. The method uses distillation to approximate expensive policy optimization, improving stability and reasoning performance on mathematical benchmarks.

0 favorites 0 likes
#mathematical-reasoning

Pass the Baton: Trajectory-Relayed On-Policy Distillation

arXiv cs.CL · 2026-07-29 Cached

Proposes Relay On-Policy Distillation (Relay-OPD) that addresses prefix failure in on-policy distillation by having the teacher briefly take over at detected handoff triggers to correct reasoning trajectories. Achieves consistent improvements on mathematical reasoning benchmarks and reduces training trajectory length by over 50%.

0 favorites 0 likes
#mathematical-reasoning

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

arXiv cs.AI · 2026-07-28 Cached

This paper investigates how LLMs' answers change under meaning-preserving paraphrases, finding that instance-level behavior is unstable (flip rates >23%) and that single-prompt accuracy masks substantial inconsistency, while a self-paraphrasing strategy can partially recover latent knowledge.

0 favorites 0 likes
#mathematical-reasoning

Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

arXiv cs.AI · 2026-07-24 Cached

This paper investigates representation robustness in LLMs for mathematical problem solving by systematically varying surface representations of equivalent problems, finding substantial sensitivity and showing that code-augmented reasoning does not uniformly eliminate brittleness.

0 favorites 0 likes
#mathematical-reasoning

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization

arXiv cs.AI · 2026-07-21 Cached

PPO-HSC introduces a High-order Sampling Coverage reward to encourage exploration of diverse reasoning patterns in RL fine-tuning of LLMs, improving solution diversity and state-space coverage on math and code tasks.

0 favorites 0 likes
#mathematical-reasoning

After OpenAI’s CDC proof announcement, GPT-5.6 used a similar prompt to close a 30-year gap in convex optimization, verified in Lean

Reddit r/singularity · 2026-07-18

Following OpenAI's CDC proof announcement, GPT-5.6 reportedly solved a 30-year open problem in convex optimization using a similar prompt, with the solution verified in the Lean proof assistant.

0 favorites 0 likes
#mathematical-reasoning

Another 50+ year-old Erdős problem falls to GPT-5.6

Reddit r/singularity · 2026-07-13

GPT-5.6 has solved another 50+ year-old unsolved problem posed by mathematician Paul Erdős, showcasing a major leap in AI's mathematical reasoning capabilities.

0 favorites 0 likes
#mathematical-reasoning

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Hugging Face Daily Papers · 2026-07-13 Cached

AdvancedMathBench is a new benchmark suite for evaluating LLMs on advanced mathematical proof generation and verification. It includes ProverBench for generation and VerifierBench for verification, demonstrating that current models like GPT-5.5-xhigh achieve only modest performance.

0 favorites 0 likes
#mathematical-reasoning

@SebastienBubeck: https://x.com/SebastienBubeck/status/2075596982622835006

X AI KOLs Timeline · 2026-07-10 Cached

GPT-5.6 significantly outperforms published state-of-the-art on a fundamental mathematical problem about gradient flow length, achieving exponential improvements. This marks a major advance in AI's ability to reason about complex mathematical questions.

0 favorites 0 likes
#mathematical-reasoning

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

arXiv cs.CL · 2026-07-10 Cached

This position paper reviews the current state of LLM-driven formal mathematics, identifies key limitations in applying these systems to open-ended research mathematics, and proposes a strategic roadmap for developing AI agents capable of advancing mathematical frontiers.

0 favorites 0 likes
#mathematical-reasoning

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

arXiv cs.AI · 2026-07-09 Cached

Introduces MIRA-Math, a benchmark that isolates and evaluates a model's ability to recognize missing atomic facts, request them precisely in natural language, and integrate the returned information into a correct final answer, using deterministic instance generation and a fixed constrained responder.

0 favorites 0 likes
#mathematical-reasoning

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

arXiv cs.AI · 2026-07-09 Cached

This paper proposes a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, evaluating it on research-level mathematical problems. Results show substantial performance gains across models, with GPT-5.5 achieving the highest solve rate of 75.2%.

0 favorites 0 likes
#mathematical-reasoning

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

arXiv cs.CL · 2026-07-08 Cached

PluraMath extends the PolyMath dataset to 18 underrepresented languages, providing a human-validated benchmark for evaluating multilingual mathematical reasoning in LLMs. The paper reveals a persistent performance gap between high-resource and low-resource languages across 27 models.

0 favorites 0 likes
#mathematical-reasoning

Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory

arXiv cs.AI · 2026-07-08 Cached

Danus is an orchestration system for research-level mathematical reasoning that uses a shared fact graph as global memory to manage parallel proof search by multiple LLM agents, with a stateless verifier for incremental proof construction.

0 favorites 0 likes
#mathematical-reasoning

Reward Granularity in RLVR: Comparing Process and Outcome Reward Structures for Mathematical Reasoning in Small Language Models

arXiv cs.LG · 2026-07-07 Cached

This paper systematically compares process and outcome reward structures for reinforcement learning with verifiable rewards (RLVR) in small language models for mathematical reasoning. The study finds that process-only supervision significantly improves accuracy and reasoning trace fidelity over outcome-only supervision, and analyzes failure modes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback