SHAPE of Chain-of-Thought in Math Reasoning
Summary
SHAPE is a framework that analyzes chain-of-thought reasoning in large language models using semantic spaces and heuristics to diagnose and improve mathematical reasoning through post-training.
View Cached Full Text
Cached at: 09/01/26, 11:48 AM
Paper page - SHAPE of Chain-of-Thought in Math Reasoning
Source: https://huggingface.co/papers/2608.28600
Abstract
SHAPE analyzes chain-of-thought reasoning via semantic spaces and heuristics to diagnose LLM mathematical reasoning and improve post-training.
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce SHAPE, a framework that analyzesChain-of-Thought(CoT) trajectories through two lenses developed in mathematics education: (1)semantic spaces: the model’s evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2)heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use SHAPE to analyze the reasoning patterns of various models. Our findings reveal that the mathematicalheuristicsemployed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a fewsemantic spacesrather than exploring many disparate ones -- a pattern consistent with human behavior. Next, we utilize the SHAPE lens to evaluate whetherpost-trainingtruly enhances mathematical proficiency. We find thatreinforcement learninginducesmode-seekingin heuristic usage. Lastly, we post-train LLMs by promoting diverseheuristicsand demonstrate its effectiveness in improving accuracy. Overall, SHAPE provides a theoretically-grounded diagnostic framework for decoding LLM reasoning and offers a new path towardpost-trainingLLMs for math reasoning. The code for our model is available at https://github.com/holi-lab/SHAPE-of-CoT
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2608\.28600
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.28600 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.28600 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.28600 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
This paper evaluates three approaches (pure chain-of-thought reasoning, single-shot code execution, and iterative code execution) on 1,000 GSM-Symbolic problems using Claude Haiku 4.5, finding that chain-of-thought is the most robust to perturbation, while code execution does not improve reasoning robustness on grade-school math problems.
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts
This research paper from MediaTek and National Taiwan University challenges the assumption that reasoning chains must be dense and sequential, showing that models can extract answers from sparse, shuffled, and noisy reasoning traces. The findings suggest that answer extraction is robust and order-independent, potentially enabling more efficient, parallelized reasoning generation.
Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
The paper proposes a mean-field framework to model chain-of-thought reasoning in LLMs as a guided discovery process on a clue graph, deriving an ODE for the fraction of discovered clues and validating it experimentally.
Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
This paper introduces Semi-CoT, a semi-supervised learning framework for chain-of-thought reasoning that uses unlabeled questions with an entropy-based selection to generate reliable pseudo reasoning chains, showing promising but mixed results on math reasoning benchmarks.
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories
Proposes Step-Aware Reasoning Energy (SARE), a geometric framework using CKA to quantify computational effort in individual chain-of-thought steps of LLMs, revealing non-uniform effort allocation and improved confidence predictions.