Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
Summary
This survey synthesizes recent advancements in mathematical reasoning with large language models, covering benchmarks, architectures, training strategies, and evaluation protocols. It identifies key challenges such as reasoning faithfulness and benchmark biases.
View Cached Full Text
Cached at: 05/20/26, 08:26 AM
# Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
Source: [https://arxiv.org/html/2605.19723](https://arxiv.org/html/2605.19723)
###### Abstract
Mathematical reasoning is essential for problem\-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems\. As Large Language Models \(LLMs\) improve their reasoning capabilities, understanding how well they perform mathematical reasoning has become increasingly important\. This survey synthesizes recent advancements in mathematical reasoning with LLMs through a structured analysis of datasets, architectures, training strategies, and evaluation protocols\. Our systematic review encompasses approximately 120 peer\-reviewed studies and preprints, examining the evolution of this research area and providing a unified analytical framework to understand current progress and limitations\.\. Our study particularly introduces a unified taxonomy of mathematical datasets, distinguishing between pretraining corpora, supervised fine\-tuning resources, and evaluation benchmarks across varying levels of reasoning complexity\. A systematic analysis of reasoning architectures and training strategies, including tool integration, verifier\-guided reasoning, and parameter\-efficient adaptation, is presented to assess their effects on reasoning robustness and generalization\. Moreover, a comparative evaluation of existing metrics highlights the gap between final\-answer accuracy and process\-level reasoning verification\. By synthesizing insights across these areas, our analysis identifies recurring failure modes, such as reasoning faithfulness issues, benchmark biases, and generalization limitations, and outlines key research directions toward improving symbolic grounding, evaluation reliability, and the development of more robust and trustworthy LLM\-based reasoning systems\.
###### keywords:
Large Language Models , Artificial Intelligence , Natural Language , Math Word Problem , Reasoning
††journal:Intelligent Systems with Applications\\affiliation
\[NUST\]organization=School of Electrical Engineering and Computer Science, National University of Science and Technology, Islamabad, Pakistan
\\affiliation
\[West\]organization=School of Computing, Data, and Mathematical Sciences, Western Sydney University, Indonesia
\\affiliation
\[MIUN\]organization=Department of Communication, Quality Management and Information Systems, Mid Sweden University, Östersund Campus, Sweden
## 1Introduction
Mathematics formalizes reasoning about number, structure, space, and change through precise symbolic systems\(Devlin and Gray,[1998](https://arxiv.org/html/2605.19723#bib.bib17); Kline,[1990](https://arxiv.org/html/2605.19723#bib.bib50)\)\. Beyond arithmetic computation, it requires compositional abstraction, multi\-step deduction, variable binding, and manipulation of formal constraints across interconnected concepts\(Huang and Chang,[2023](https://arxiv.org/html/2605.19723#bib.bib42)\)\. Mathematical reasoning therefore differs fundamentally from surface\-level numerical pattern recognition; it demands structured inference, proof construction, and long\-horizon dependency tracking\(Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110)\)\. These characteristics make mathematics a rigorous benchmark for evaluating the reasoning capabilities of LLMs\.
Recent LLMs demonstrate strong performance across diverse natural language tasks through large\-scale pretraining\(Weiet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib104); Matzakoset al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib69)\)\. Prompting strategies such as Chain\-of\-Thought \(CoT\)\(Weiet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib104)\)and Tool\-Integrated Reasoning \(TIR\)\(Li,[2024](https://arxiv.org/html/2605.19723#bib.bib58)\)improve multi\-step problem solving and enable progress in algebraic manipulation, theorem proving, and structured word problems\(Imaniet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib45)\)\. Nevertheless, solving mathematical problems remains a long\-standing challenge in artificial intelligence\(Feigenbaum and Feldman,[1963](https://arxiv.org/html/2605.19723#bib.bib23); Hosseiniet al\.,[2014](https://arxiv.org/html/2605.19723#bib.bib40)\)\. Despite scaling advances, LLMs continue to exhibit brittle generalization, arithmetic inconsistencies \(see Figure[1](https://arxiv.org/html/2605.19723#S1.F1)\), symbolic reasoning failures, and error propagation across long reasoning chains\(Raeet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib81); Lewkowyczet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib55)\)\.
As illustrated in Figure[1](https://arxiv.org/html/2605.19723#S1.F1), the model generates a largely coherent chain\-of\-thought solution that correctly formulates intermediate steps \(e\.g\., derivingx\+10=20x\+10=20and solvingx=10x=10\), yet produces an incorrect final answer\. This discrepancy highlights a critical limitation: while intermediate reasoning steps may appear logically consistent, the final output can still deviate due to arithmetic slips or weak answer verification\. Such behavior reflects the lack of global consistency in LLM reasoning, where local step\-by\-step plausibility does not guarantee correctness of the final result\. Consequently, performance declines substantially on competition\-level and university\-level problems requiring abstraction, formal rigor, and non\-standard solution strategies\.
These limitations arise from multiple factors\. Existing benchmarks often emphasize final\-answer accuracy and may contain distributional biases or pretraining contamination, limiting their ability to measure genuine reasoning generalization\(Mishraet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib73); Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43); Guanet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib31)\)\. Moreover, the next token prediction objective optimizes local coherence rather than globally valid derivations, which can encourage shortcut learning\(Ahnet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib3)\)\. Although meta cognitive prompting and self reflection techniques improve solution quality\(Wang and Zhao,[2023](https://arxiv.org/html/2605.19723#bib.bib103)\), the faithfulness of generated reasoning traces remains uncertain\(Didolkaret al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib18)\)\.
Despite rapid progress, research on mathematical reasoning in LLMs remains fragmented across datasets, architectures, training strategies, and evaluation methodologies\. Existing surveys often focus on individual aspects such as reasoning prompting, benchmark datasets, or evaluation protocols\. In contrast, this survey provides a unified analysis across datasets, model architectures, training strategies, and evaluation frameworks, highlighting how these components jointly influence mathematical reasoning performance\.
Recent surveys have examined related aspects of reasoning in large language models, including prompting strategies, benchmark datasets, and reasoning evaluation protocols\. However, most existing surveys focus on a single dimension of the problem, such as prompting techniques, benchmark comparisons, or reasoning evaluation\. In contrast, this survey provides a unified analysis that jointly examines datasets, architectural adaptations, training strategies, and evaluation methodologies\. By integrating these components within a single framework, the survey highlights how design choices across the entire model development lifecycle collectively influence mathematical reasoning performance\.
Figure 1:Top:Math word problem\.Bottom:Step\-by\-step erroneous solution\.InputQuestion: Dane’s two daughters need to plant a certain number of flowers each to grow a garden\. As the days passed, the flowers grew into 20 more but 10 of them died\. Dane’s daughters harvested the flowers and split them between 5 different baskets, with 4 flowers in each basket\. How many flowers did each daughter plant initially? Answer:
Step\-by\-Step SolutionCalculate the total number of flowers harvested: Dane’s daughters split the flowers into 5 baskets, with 4 flowers in each basket\. The total number of flowers harvested is:5×4=20flowers5\\times 4=20\\ \\text\{flowers\}Account for the flowers that grew and died: The problem states that 20 additional flowers grew, but 10 of them died\. Let the initial number of flowers planted by the two daughters bexx\. After accounting for the growth and death of flowers, the total number of flowers becomes:x\+20−10=x\+10x\+20\-10=x\+10Set up the equation for the total flowers harvested: The total number of flowers harvested is equal to the total number of flowers after growth and death:x\+10=20x\+10=20Solve forxx: Subtract 10 from both sides:x=10x=10Determine how many flowers each daughter planted: The two daughters together plantedx=10x=10flowers\. Since the flowers were planted equally between the two daughters, each daughter planted:102=5flowers\\frac\{10\}\{2\}=5\\ \\text\{flowers\}Final Answer: 6
This survey provides a structured synthesis of mathematical reasoning in LLMs across the full model development lifecycle\. The research landscape can be organized into three major components: datasets, model architectures and training strategies, and evaluation methodologies\. We begin by analyzing benchmark datasets spanning foundational corpora, grade\-school and competition\-level problems\(Heet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib34); Gaoet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib28); Fanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib22); Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43)\), and formal proof corpora\(Zhenget al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib118); Luet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib67); Azerbayevet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib7); Zhanget al\.,[2024b](https://arxiv.org/html/2605.19723#bib.bib114)\)\(Section[4](https://arxiv.org/html/2605.19723#S4)\)\. We then examine architectural adaptations \(Section[5](https://arxiv.org/html/2605.19723#S5)\) and training strategies—including pretraining\(Devlinet al\.,[2019](https://arxiv.org/html/2605.19723#bib.bib16)\), supervised fine\-tuning\(Liet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib59)\), and reinforcement learning\(Stoneet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib94)\)—that aim to improve reasoning robustness and fidelity \(Section[6](https://arxiv.org/html/2605.19723#S6)\)\. Finally, we review evaluation methodologies and metrics, emphasizing the gap between answer\-level correctness and step\-level reasoning verification \(Section[7](https://arxiv.org/html/2605.19723#S7)\)\.
## Research Questions
To structure the analysis of mathematical reasoning in LLMs, this survey is guided by the following research questions:
- •RQ1:What datasets and benchmarks are currently used to evaluate mathematical reasoning in LLMs, particularly for competition\-level and Olympiad\-style problems?
- •RQ2:Which architectural designs and training strategies most effectively improve mathematical reasoning capabilities in LLMs?
- •RQ3:How do LLMs perform multi\-step reasoning, and to what extent are their generated reasoning traces faithful and verifiable?
- •RQ4:What evaluation methodologies and metrics are used to assess mathematical reasoning performance, and what limitations do current evaluation protocols exhibit?
- •RQ5:What key challenges and open research directions remain for advancing robust mathematical reasoning in LLMs?
## Contributions
This survey provides a comprehensive synthesis of research on mathematical reasoning in large language models\. The main contributions are summarized as follows:
- •Unified dataset taxonomy\.We organize existing mathematical datasets according to their functional role in LLM development, distinguishing between pretraining corpora, supervised fine\-tuning datasets, and evaluation benchmarks across different levels of reasoning difficulty\.
- •Systematic analysis of reasoning architectures and training strategies\.We review architectural adaptations, reasoning\-enhancement mechanisms, and training pipelines—including tool integration, verifier\-guided reasoning, and parameter\-efficient fine\-tuning—and analyze how these methods influence reasoning robustness and generalization\.
- •Comparative evaluation of reasoning metrics\.We examine existing evaluation methodologies and highlight the gap between answer\-level correctness and process\-level reasoning verification\.
- •Cross\-sectional synthesis of failure modes and limitations\.By integrating insights across datasets, architectures, training strategies, and evaluation frameworks, we identify recurring failure patterns that limit reliable mathematical reasoning in LLMs\.
- •Future research directions\.We outline key open challenges and promising research directions for improving reasoning faithfulness, symbolic grounding, and evaluation reliability in future LLM systems\.
## Paper Organization
The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2605.19723#S2)presents the structured literature review methodology used to collect and analyze relevant studies\. Section[3](https://arxiv.org/html/2605.19723#S3)introduces the background on mathematical reasoning in LLMs\. Section[4](https://arxiv.org/html/2605.19723#S4)examines benchmark datasets, Section[5](https://arxiv.org/html/2605.19723#S5)discusses model architectures and training adaptations, Section[6](https://arxiv.org/html/2605.19723#S6)reviews evaluation methodologies, Section[7](https://arxiv.org/html/2605.19723#S7)synthesizes open challenges and future research directions, and Section[8](https://arxiv.org/html/2605.19723#S8)concludes the survey\.
For clarity, Figure[2](https://arxiv.org/html/2605.19723#Sx3.F2)summarizes the overall research landscape examined in this survey, highlighting the relationship between datasets, model architectures, and evaluation methodologies\.
Mathematical Reasoning in LLMsDatasetsTraining DatasetsBenchmarksArchitecturesBase LLM ModelsReasoning EnhancementsEvaluationAnswer MetricsProcess Verification
Figure 2:Conceptual landscape of research on mathematical reasoning in large language models\. The field spans dataset design, model architectures, and evaluation methodologies, each addressing different aspects of reasoning capability\.
## 2Literature Collection Methodology
Given the rapid publication cycle and preprint\-driven dissemination in AI, strict exhaustiveness is impractical\. We therefore adopt a structured and reproducible scoping review methodology with systematic filtering to maximize thematic coverage while maintaining methodological transparency\.
Table 1:Eligibility and inclusion criteria applied during structured literature screening\.### 2\.1Search Strategy and Keyword Formulation
The search protocol is aligned with the five research questions guiding this survey\. A multi\-tier Boolean query matrix captures intersections among LLMs, mathematical reasoning, benchmark datasets, architectural optimization strategies, reasoning faithfulness, and evaluation protocols\.
Ten overlapping keyword configurations span multiple semantic projections of the research space, ranging from broad queries \(e\.g\.,*LLMs AND mathematical reasoning AND benchmarks*\) to targeted methodological intersections \(e\.g\.,*formal verification*,*tool integration*, or*reasoning faithfulness*\)\. This redundancy mitigates query sensitivity and ranking bias by ensuring that studies weakly surfaced in one query are retrieved through alternative formulations\.
All query strings, execution dates, and retrieval counts are logged to ensure traceability\. Table[2](https://arxiv.org/html/2605.19723#S2.T2)summarizes the search matrix\.
Table 2:Structured search strategy organized by thematic objective and keyword blocks\.
### 2\.2Phase 1: Identification and Structured Extraction
Searches are conducted across Google Scholar, arXiv, ScienceDirect, ACM Digital Library, and IEEE Xplore\. Executing the ten query configurations yields an initial retrieval of 271,100 records, reflecting search breadth rather than topical precision\.
Due to export restrictions and API limits—particularly in Google Scholar and ACM DL—we apply a relevance\-ranked extraction window\. For each query, the top 200–300 ranked results \(approximately the first 20–30 pages\) are collected\. Manual inspection confirms significant topical dilution beyond this threshold\.
This relevance\-window strategy aggregates 12,679 accessible records for downstream screening\. Table[3](https://arxiv.org/html/2605.19723#S2.T3)summarizes database\-wise retrieval distributions\.
Table 3:Database\-level retrieval, screening, and deduplication statistics for the structured review process\.
### 2\.3Phase 2: Deduplication and Relevance Screening
Database\-level deduplication removes 3,720 overlapping records, leaving 8,959 unique entries\. Consolidation across query intersections further reduces the corpus to 7,551 distinct articles\.
Given this scale, automated scripts extract titles and abstracts and compute keyword\-density scores aligned with the survey’s research questions\. These scores are used only to prioritize screening order; all inclusion decisions are performed through manual review\.
Title screening removes clearly out\-of\-scope work \(e\.g\., generic NLP tasks or simple arithmetic computation\), reducing the pool to 1,406 candidate papers\. Abstract screening then applies a stricter conceptual filter requiring explicit engagement with LLM\-based mathematical reasoning \(e\.g\., multi\-step reasoning, complex word problems, or formal proofs\)\. This stage produces the candidate set for full\-text evaluation \(Table[4](https://arxiv.org/html/2605.19723#S2.T4)\)\.
Table 4:Retrieval and multi\-stage screening statistics across boolean keyword query sets\.Articles are excluded if they:
- •do not focus on LLM\-based methodologies,
- •address purely linguistic tasks without quantitative reasoning,
- •lack substantive engagement with mathematical reasoning, benchmarks, or reasoning\-specific modeling\.
### 2\.4Phase 3: Full\-Text Review and Final Inclusion
Remaining candidates undergo full\-text evaluation under stricter inclusion criteria emphasizing:
- •explicit focus on mathematical reasoning tasks,
- •empirical evaluation on recognized benchmarks,
- •reasoning\-specific modeling or training strategies,
- •analysis of reasoning faithfulness or verification\.
This stage yields a final corpus of 120 core papers forming the empirical foundation of this survey\. Figure[3](https://arxiv.org/html/2605.19723#S2.F3)presents the PRISMA\-style flow diagram summarizing the progression from identification to inclusion\.
Figure 3:PRISMA flow diagram of the systematic literature review selection processIdentificationScreeningIncludedRecords identified from:Databases \(n = 271,100\)Records removed before screening:Unretrieved / API Limits \(n = 258,421\)Intra\-database duplicates \(n = 3,720\)Cross\-database duplicates \(n = 1,408\)Records screened\(n = 7,551\)Records excluded:Title screening \(n = 6,145\)Reports assessed for eligibility\(n = 1,406\)Reports excluded:Abstract screening \(n = 1,286\)Included Articles\(n = 120\)
### 2\.5Mitigation of Selection Bias and Reproducibility
To mitigate ranking bias and extraction limitations, three safeguards are implemented\.
#### 2\.5\.1Strategic Query Redundancy
Overlapping Boolean query structures ensure that studies weakly ranked in one search configuration are recoverable through alternative semantic formulations\.
#### 2\.5\.2Bi\-directional Snowballing
Backward and forward citation tracking during full\-text review identifies seminal contributions and niche subfields potentially missed by search engine ranking thresholds\.
#### 2\.5\.3Execution Logging
All searches are conducted within a fixed temporal window\. Query strings, ranking thresholds, filtering scripts, and screening decisions are archived to ensure methodological transparency and approximate reproducibility\.
Overall, this methodology prioritizes structured thematic coverage and analytical rigor over strict exhaustiveness, which remains infeasible in rapidly evolving AI research ecosystems\.
## 3Background
This section introduces the conceptual foundations of mathematical reasoning and its interaction with LLMs\. It outlines how LLMs approximate reasoning and highlights the fundamental challenges involved in evaluating their mathematical reasoning capabilities\.
### 3\.1Mathematical Reasoning: Definitions and Scope
Mathematical reasoning formalizes inference over symbolic structures, numerical quantities, and abstract relations\. Unlike arithmetic computation, which follows fixed procedural rules, mathematical reasoning requires compositional abstraction, variable binding, constraint propagation, multi\-step deduction, and often proof construction\(Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110)\)\. These characteristics make mathematics a stringent test of systematic generalization and logical consistency\.
Within the context of LLM evaluation, mathematical reasoning spans multiple levels: \(i\) arithmetic and numerical manipulation, \(ii\) algebraic and equation\-based transformation, \(iii\) word\-problem reasoning that integrates semantic parsing with quantitative inference, and \(iv\) formal theorem proving over axiomatic systems\. Although these tasks differ in surface representation, they all depend on systematic generalization and logically consistent intermediate reasoning steps\.
### 3\.2Reasoning Mechanisms in LLMs
LLMs implement autoregressive next token prediction trained on large\-scale textual corpora\(Weiet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib104)\)\. At scale, they acquire strong linguistic competence and partial numerical reasoning capabilities\(Matzakoset al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib69)\)\. However, the maximum\-likelihood training objective optimizes local token prediction rather than globally valid logical derivations, challenging rigorous theorem proving and complex problem\-solving\(Xinet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib107); Forootani,[2025](https://arxiv.org/html/2605.19723#bib.bib25)\)\.
This creates a structural mismatch with mathematics, which requires discrete symbolic manipulation and verifiable inference\. In hybrid reasoning tasks, models must jointly perform semantic parsing and symbolic transformation by identifying operators, binding operands, mapping language to formal structures, and executing multi\-step derivations\(Lample and Charton,[2020](https://arxiv.org/html/2605.19723#bib.bib54); Kukreja and Sakshi,[2022](https://arxiv.org/html/2605.19723#bib.bib53); Liuet al\.,[2025b](https://arxiv.org/html/2605.19723#bib.bib65); Fu,[2025](https://arxiv.org/html/2605.19723#bib.bib26)\)\. Purely parametric models often entangle these processes, leading to operator misuse, arithmetic inconsistencies, and cascading reasoning errors in complex problems\(Liet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib56)\)\.
Moreover, next token prediction encourages pattern imitation rather than principled derivation, lacking true cognitive adaptability and iterative refinement\(Duanet al\.,[2020](https://arxiv.org/html/2605.19723#bib.bib21); Didolkaret al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib18); Tianet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib97)\)\. This contributes to memorization effects and discontinuous performance across problem types and difficulty levels\(Ahnet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib3); Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43)\)\.
### 3\.3Prompting and Tool\-Augmented Reasoning
Prompting strategies partially mitigate these limitations\. CoT prompting elicits intermediate reasoning steps and improves performance on multi\-step problems\(Weiet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib104)\)\. Self\-consistency decoding\(Wanget al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib100)\)and structured prompting variants further enhance robustness\(Aaliet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib1)\)\. TIR augments language models with external symbolic engines such as calculators or theorem provers, enabling more reliable algebraic manipulation and theorem\-level reasoning\(Li,[2024](https://arxiv.org/html/2605.19723#bib.bib58); Imaniet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib45)\)\.
Despite these advances, mathematical problem solving remains a long\-standing challenge in artificial intelligence\(Feigenbaum and Feldman,[1963](https://arxiv.org/html/2605.19723#bib.bib23); Hosseiniet al\.,[2014](https://arxiv.org/html/2605.19723#bib.bib40)\)\. Even state\-of\-the\-art models struggle on competition\-level and university\-level benchmarks that require abstraction, deep formula knowledge, and non\-standard solution strategies\(Lewkowyczet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib55); Chernyshevet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib12); Liuet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib63)\)\. Improvements from prompting often reflect improved elicitation of reasoning rather than fundamentally stronger internal reasoning capabilities\.
The current landscape of mathematical reasoning in LLMs is shaped not by a single reasoning mechanism, but by a family of prompting, search, tool\-use, and verification strategies that differ in both capability and reliability\. Table[5](https://arxiv.org/html/2605.19723#S3.T5)summarizes the major reasoning enhancement approaches and compares them in terms of their core idea, practical strengths, and key limitations\.
Table 5:Comparison of major reasoning enhancement strategies for mathematical problem solving in LLMs\.As the comparison indicates, improvements in mathematical performance often arise from better elicitation, sampling, decomposition, or verification rather than from fundamentally solved reasoning\. This distinction is crucial, because a method may improve final\-answer accuracy while still leaving unresolved questions about faithfulness, internal computation, and robustness\. These issues motivate the evaluation challenges discussed next\.
### 3\.4Foundations of Evaluation and Faithfulness
Evaluating mathematical reasoning introduces additional challenges\. Many benchmarks emphasize final answer accuracy without verifying the correctness of intermediate reasoning steps\(Mishraet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib73)\)\. This evaluation paradigm makes it difficult to distinguish systematic reasoning from pattern matching over dataset regularities\.
Meta\-cognitive and role\-based prompting techniques attempt to improve self\-reflection and error correction\(Didolkaret al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib18); Han and Wang,[2024](https://arxiv.org/html/2605.19723#bib.bib33); Wang and Zhao,[2023](https://arxiv.org/html/2605.19723#bib.bib103)\)\. However, generated reasoning traces may not faithfully represent the model’s internal inference process\(Didolkaret al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib18)\)\. The gap between plausible explanations and verifiable derivations therefore remains significant\. Furthermore, feedback\-driven refinement in LLMs has not yet matched the structured guidance provided by human instruction\(Antonet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib4)\)\.
Overall, LLM\-based mathematical reasoning emerges from statistical sequence modeling rather than explicit symbolic computation\. While prompting, search strategies, and tool integration can partially mitigate these limitations, they do not fully resolve the tension between probabilistic language modeling and the symbolic precision required for mathematics\. These challenges motivate the systematic analysis of datasets, architectures, training strategies, and evaluation protocols presented in the subsequent sections\.
Mathematical reasoning tasks addressed by LLMs can be broadly categorized according to the cognitive operations required to reach a solution\. Prior work commonly identifies five major categories: arithmetic reasoning, equation and algebraic reasoning, word problem reasoning, symbolic or formal reasoning, and algorithmic reasoning\. A structured overview of these categories, along with their descriptions and representative benchmarks, is presented in Table[6](https://arxiv.org/html/2605.19723#S3.T6)\. While these categories often overlap in practice, they provide a useful conceptual framework for analyzing datasets, model architectures, and evaluation strategies discussed in the following sections\.
Table 6:Taxonomy of mathematical reasoning tasks studied in LLM research\.
## 4Mathematical Reasoning Datasets
The progress of mathematical reasoning in LLMs depends critically on the availability and quality of curated mathematical datasets\. Dataset design determines what models learn, how they generalize, and which forms of reasoning they can reliably perform\. Effective mathematical datasets span diverse problem types, cover multiple levels of cognitive complexity, and ideally provide intermediate reasoning traces that expose latent solution structure\(Gonget al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib30); Karra and Lasfar,[2023](https://arxiv.org/html/2605.19723#bib.bib49)\)\.
Table 7:Taxonomy of mathematical datasets for training LLMs\.Mathematical corpora also vary along several orthogonal dimensions, including cognitive depth, supervision format, symbolic structure, and evaluation protocol\. Consequently, mathematical benchmarks cannot be treated as homogeneous natural language resources; they must be analyzed as structured datasets that support different stages of learning and evaluation\.
This section organizes existing work along three dimensions\. First, it categorizes datasets by cognitive complexity and functional role in training or evaluation\. Second, it reviews semantic and structural representations that map natural language problems to formal mathematical objects\. Third, it examines symbolic retrieval and tokenization mechanisms that influence how models encode and manipulate mathematical expressions\. The section concludes with a critical discussion of current benchmark limitations\.
Table 8:Evaluation benchmarks for assessing mathematical reasoning capabilities of LLMs\.Tables[7](https://arxiv.org/html/2605.19723#S4.T7)and[8](https://arxiv.org/html/2605.19723#S4.T8)organize datasets by training and evaluation roles, a benchmark\-level comparison is needed to clarify how widely used corpora differ in reasoning type, difficulty, scoring format, and evaluation value\. Table[9](https://arxiv.org/html/2605.19723#S4.T9)provides this comparative view and highlights that current benchmarks do not measure a single unified notion of mathematical competence; instead, they probe different layers of reasoning ranging from arithmetic decomposition to formal symbolic verification\.
Table 9:Comparative overview of major benchmarks for mathematical reasoning in LLMs\.This comparison shows that benchmark choice strongly shapes the conclusions drawn about model capability\. Benchmarks such as GSM8K emphasize multi\-step numerical reasoning\(Cobbeet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib14)\), whereas MATH\(Hendryckset al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib35)\)and Olympiad\-style corpora probe deeper abstraction and symbolic manipulation\(Toshniwalet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib98); Pasteret al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib78)\)\. Formal proof datasets\(Zhanget al\.,[2024b](https://arxiv.org/html/2605.19723#bib.bib114)\), in contrast, prioritize verifiability and logical rigor\. These differences reinforce the need to examine not only dataset content but also the structural representations through which mathematical problems are encoded and solved\.
### 4\.1Semantic and Structural Representations
As reasoning complexity increases, natural language descriptions alone become insufficient for representing mathematical structure\. Mathematical problems typically involve operator hierarchies, symbolic dependencies, variable binding, and latent quantitative relations that plain text representations often obscure\(Shalytet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib88)\)\. Consequently, many mathematical datasets incorporate semantic or structural representations that map problem statements to formal mathematical objects\.
Semantic parsing methods convert textual descriptions into executable expressions or formal logical structures, enabling equation construction and symbolic reasoning\(Roy and Roth,[2015](https://arxiv.org/html/2605.19723#bib.bib85)\)\. Unit Dependency Graphs explicitly represent relationships among quantities, units, and operations, helping systems infer equation structure and resolve numerical references\(Roy and Roth,[2017](https://arxiv.org/html/2605.19723#bib.bib86); Daveet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib75)\)\. Domain\-adapted encoders such as MathBERT further support this process by learning symbol\-aware representations from math\-rich corpora containing LaTeX expressions and structured notation\(Penget al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib79)\)\. Contextual embedding and in\-context learning techniques\(Devlinet al\.,[2019](https://arxiv.org/html/2605.19723#bib.bib16); Brownet al\.,[2020](https://arxiv.org/html/2605.19723#bib.bib9)\)complement these methods by allowing models to infer latent schemas from demonstrations and better interpret variables, operators, and mathematical relations\.
These approaches collectively move mathematical datasets beyond surface\-level text representations toward structured semantic supervision, which is essential for tasks involving equation mapping, operator prediction, symbolic grounding, and formal derivation\.
Beyond semantic representations, mathematical datasets must also address how symbolic expressions are encoded, indexed, and retrieved during reasoning\.
### 4\.2Symbolic Representation and Retrieval
Once semantic structure is identified, reasoning systems must retrieve, encode, and manipulate formal symbolic objects\. This requirement becomes particularly important in theorem proving, proof synthesis, and autoformalization tasks, where reasoning depends on operations over structured expressions rather than free\-form text generation\(Stathopoulos,[2022](https://arxiv.org/html/2605.19723#bib.bib93)\)\.
Formal representation standards such as MathML address this challenge by separating visual layout from semantic structure\. Presentation MathML captures the two\-dimensional rendering of expressions, while Content MathML encodes operator–argument relationships and functional composition\(Ausbrooks and others,[2003](https://arxiv.org/html/2605.19723#bib.bib5); Zanibbiet al\.,[2011](https://arxiv.org/html/2605.19723#bib.bib111)\)\. Building on these representations, Symbol Layout Trees and Operator Trees model spatial and logical expression structure, enabling indexing, normalization, and structural matching of mathematical expressions\.
Hybrid retrieval frameworks such as MCAT combine textual and symbolic views to support structured math retrieval\(Yokoet al\.,[2014](https://arxiv.org/html/2605.19723#bib.bib109); Kristiantoet al\.,[2016](https://arxiv.org/html/2605.19723#bib.bib52)\)\. Dense retrieval approaches such as Tangent\-CFT extend this paradigm by learning similarity functions over symbolic paths and expression structures\(Mansouriet al\.,[2019](https://arxiv.org/html/2605.19723#bib.bib68)\)\. These methods collectively shift reasoning systems from token\-level processing toward structure\-aware retrieval and manipulation of mathematical expressions\.
### 4\.3Tokenization Mechanisms
Despite advances in semantic parsing and symbolic retrieval, numerical tokenization remains a persistent bottleneck for LLM\-based mathematical reasoning\. Tokenizers designed for natural language frequently fragment numbers in ways that disrupt place\-value structure and weaken arithmetic reasoning\.
Standard Byte Pair Encoding \(BPE\) remains widely used but often splits numbers inconsistently and fails to preserve numerical regularities\(Zouharet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib120)\)\. Alternative schemes explicitly encode positional structure\. Left\-to\-Right tokenization processes digits from the most significant position, whereas Right\-to\-Left tokenization aligns tokenization with carry\-based arithmetic operations\(Singh and Strouse,[2024](https://arxiv.org/html/2605.19723#bib.bib91)\)\. Digit\-level tokenization preserves exact place\-value information but increases sequence length\(Chowdheryet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib13)\)\. Continuous numerical encodings such as xVal bypass discrete tokenization and embed magnitude directly in continuous space\(Golkaret al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib29)\)\.
These findings highlight that mathematical reasoning performance depends not only on model scale and training data but also on the representational assumptions introduced during tokenization\.
### 4\.4Benchmark Limitations and Research Gaps
Despite rapid progress, current mathematical datasets still leave substantial gaps in evaluating genuine reasoning ability\.
First, benchmark validity remains vulnerable to training–test contamination\. Large web\-derived corpora such as OpenWebMath improve scale but also increase the risk of noisy extraction, weak provenance, and overlap with downstream evaluation sets\(Pasteret al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib78)\)\. Consequently, reported improvements may sometimes reflect memorization or benchmark exposure rather than true reasoning generalization\.
Second, many training pipelines rely heavily on synthetic reasoning traces\. Synthetic CoT supervision enables scalable instruction tuning but often reproduces the biases and shortcuts of the teacher model generating the traces\(Guanet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib31)\)\. This dependency reduces diversity in reasoning strategies and may limit downstream improvements\(Tanget al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib95); Huanget al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib44)\)\.
Third, current benchmarks inadequately evaluate out\-of\-distribution generalization\. Many datasets rely on stable problem templates, encouraging pattern recognition rather than transferable abstraction\(Lindseyet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib61)\)\. Empirical evidence shows that models frequently solve familiar formats while failing on structurally novel variants requiring the same underlying principle\(Ahnet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib3)\)\. These observations suggest that many benchmarks still measure pattern recognition rather than robust mathematical reasoning\.
Fourth, evaluation remains fragmented across reasoning modes\. Arithmetic reasoning\(Koncel\-Kedziorskiet al\.,[2016](https://arxiv.org/html/2605.19723#bib.bib51); Miaoet al\.,[2020](https://arxiv.org/html/2605.19723#bib.bib70); Cobbeet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib14); Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110)\), competition mathematics\(Hendryckset al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib35); Heet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib34); Gaoet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib28); Fanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib22)\), theorem proving\(Zhenget al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib118); Luet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib67); Azerbayevet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib7)\),higher\-education level problems\(Alavi Naeiniet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib6); Chernyshevet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib12)\)and adversarial robustness are often studied separately, making it difficult to determine whether models exhibit unified mathematical competence or narrow task\-specific abilities\.
Finally, dataset design remains weakly aligned with representational constraints such as symbolic structure and numerical tokenization\. Recent studies demonstrate that improvements in tokenization and structured representation can significantly affect mathematical accuracy\(Singh and Strouse,[2024](https://arxiv.org/html/2605.19723#bib.bib91)\)\. This finding highlights the need for benchmarks that evaluate not only final answers but also structural consistency, symbolic validity, and representation\-sensitive reasoning\.
Overall, future benchmarks must reduce contamination risks, incorporate diverse human\-authored reasoning traces, test principled generalization, and integrate symbolic, semantic, and numerical evaluation within a unified framework\. Taken together, existing benchmarks capture different and only partially overlapping aspects of mathematical reasoning\. As a result, improvements on one benchmark family do not necessarily indicate broader mathematical competence\. This fragmentation highlights the need for evaluation frameworks capable of measuring reasoning across levels of abstraction and representation\.
## 5Architectures and Training Strategies for Mathematical Reasoning
Modern LLMs rely primarily on Transformer\-based architectures trained on massive text corpora containing billions to trillions of tokens\. Decoder\-only Transformer models dominate current LLM designs because they support efficient autoregressive generation and scale effectively across distributed computing infrastructure\. These architectures integrate multi\-head self\-atten\- tion, feed\-forward networks, and layer normalization to capture long\-range dependencies and contextual relationships in text\.
Transformers scale efficiently through distributed training strategies such as data parallelism, tensor parallelism, and pipeline parallelism, enabling the use of large GPU or TPU clusters for training\(Smithet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib92)\)\. Once the architecture is defined, models undergo a multi\-stage training pipeline that includes large\-scale pretraining followed by several post\-training alignment stages\. These stages collectively determine the reasoning, generation, and task\-adaptation capabilities of modern LLMs\.
### 5\.1Reasoning Oriented Extensions of LLM Architectures
While standard transformer architectures provide the foundation for most modern LLMs, mathematical reasoning performance often depends on additional reasoning\-oriented mechanisms rather than architecture alone\. Several extensions have therefore emerged to improve structured reasoning\.
One line of work augments language models with explicit reasoning traces, such as chain\-of\-thought prompting and stepwise scratchpad generation, allowing intermediate computations to be externalized during inference\. Another direction integrates external tools such as symbolic solvers, calculators, or programming environments, enabling models to offload precise computation to deterministic systems\. A third approach introduces verifier or critic models that evaluate candidate reasoning paths and guide the model toward more reliable solutions\.
These approaches suggest that improvements in mathematical reasoning often arise from hybrid reasoning pipelines that combine neural generation with symbolic verification, structured search, or external computation rather than from architectural scaling alone\.
### 5\.2Pretraining
pretraining forms the computational foundation of LLM development\. During this stage, models learn general linguistic and semantic representations by training on extremely large corpora that often contain trillions of tokens\. Training typically requires distributed infrastructure consisting of hundreds or thousands of GPUs or TPUs\. Efficient training relies on optimization strategies such as data parallelism, tensor parallelism, pipeline parallelism, and memory\-efficient techniques including Zero Redundancy Optimization \(ZeRO\)\.
Successful pretraining requires careful balancing of model capacity, training data quality, and computational resources\. Research on scaling laws shows that model performance improves predictably as model parameters, dataset size, and training compute increase\(Kaplanet al\.,[2020](https://arxiv.org/html/2605.19723#bib.bib48); Hernandezet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib38); Zhanget al\.,[2025a](https://arxiv.org/html/2605.19723#bib.bib116); Weiet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib105)\)\. These scaling relationships guide the design of modern LLM training pipelines and explain the rapid capability improvements observed in recent models\(Hoffmannet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib39)\)\.
### 5\.3Post\-Training and Alignment
Although pretraining provides broad language competence, additional training stages are required to adapt models for downstream tasks and align them with human instructions\. Post\-training typically includes supervised fine\-tuning, instruction tuning, and reinforcement learning based on feedback\.
#### 5\.3\.1Supervised Fine\-Tuning
Supervised fine\-tuning adapts a pre\-trained model to specific tasks using labeled input–output pairs\(Ziegleret al\.,[2020](https://arxiv.org/html/2605.19723#bib.bib119)\)\. Task\-specific datasets, such as question–answer pairs or reasoning examples, guide the model toward producing structured and task\-relevant responses\(Wuet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib106)\)\. Compared with pretraining corpora, these datasets are significantly smaller but provide high\-quality supervision for specialized capabilities\.
#### 5\.3\.2Instruction Tuning
Instruction tuning improves the ability of language models to follow natural language instructions\. In this approach, models are trained on diverse instruction–response pairs that represent human task descriptions\(Minet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib72)\)\. This training enables models to generalize to new tasks through zero\-shot and few\-shot prompting\(Xianet al\.,[2017](https://arxiv.org/html/2605.19723#bib.bib108); Wanget al\.,[2020](https://arxiv.org/html/2605.19723#bib.bib102)\), which is critical for handling novel mathematical word problems with diverse linguistic formulations\(Pourpanahet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib80)\)\.
Many instruction\-tuning datasets incorporate CoT reasoning traces, where intermediate steps explicitly describe the reasoning process\. These structured responses significantly improve performance on multi\-step reasoning tasks\(Parnami and Lee,[2022](https://arxiv.org/html/2605.19723#bib.bib77); Chen,[2023](https://arxiv.org/html/2605.19723#bib.bib11)\)\.
#### 5\.3\.3Reinforcement Learning
Reinforcement learning further aligns model behavior with desired outputs by optimizing responses according to reward signals\. A typical reinforcement learning framework includes a reward model that evaluates response quality and a policy optimization stage that updates the language model accordingly\.
Reward models can be rule\-based or model\-based\(Liuet al\.,[2025a](https://arxiv.org/html/2605.19723#bib.bib62)\)\. Rule\-based reward systems verify outputs through deterministic constraints such as numerical correctness or structured formatting, providing high reliability for tasks like mathematical reasoning\(Muet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib74)\)\. Model\-based reward models estimate response quality based on learned preference signals\(Ouyanget al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib76)\), though they are more prone to reward hacking, necessitating specific mitigation strategies\(Miaoet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib71); Fuet al\.,[2026](https://arxiv.org/html/2605.19723#bib.bib27)\)\. Policy optimization algorithms such as Proximal Policy Optimization \(PPO\) and Trust Region Policy Optimization \(TRPO\) enable stable reinforcement learning for LLMs\(Schulmanet al\.,[2015](https://arxiv.org/html/2605.19723#bib.bib87); Achiamet al\.,[2017](https://arxiv.org/html/2605.19723#bib.bib2); Zhanget al\.,[2024a](https://arxiv.org/html/2605.19723#bib.bib112)\)\.
In mathematical reasoning, model\-based evaluation targets different levels of reasoning\. Outcome Reward Models \(ORMs\) check the final answer’s logical validity, guiding search and test\-time scaling\(Thatikondaet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib96)\)\. Process Reward Models \(PRMs\) evaluate each intermediate step to ensure correctness and prevent flawed reasoning\(Zhanget al\.,[2025b](https://arxiv.org/html/2605.19723#bib.bib117)\)\. Recent Reward Reasoning Models \(RRMs\) go further by generating reasoning steps to judge complex proofs more accurately\(Guoet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib32)\)\.
For policy optimization, foundational algorithms such as Proximal Policy Optimization \(PPO\) and Trust Region Policy Optimization \(TRPO\) provide stable reinforcement learning\(Schulmanet al\.,[2015](https://arxiv.org/html/2605.19723#bib.bib87); Achiamet al\.,[2017](https://arxiv.org/html/2605.19723#bib.bib2)\)\. Tailored algorithms like Group Relative Policy Optimization \(GRPO\) efficiently optimize reasoning without memory\-intensive critic models\(Shaoet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib89); Zhanget al\.,[2024a](https://arxiv.org/html/2605.19723#bib.bib112)\)\.
#### 5\.3\.4Parameter\-Efficient Fine\-Tuning
Adapting extremely large models through full fine\-tuning is computationally expensive\. Parameter\-efficient fine\-tuning \(PEFT\) methods address this limitation by updating only a small subset of model parameters while freezing the majority of pretrained weights\. Among these methods, Low\-Rank Adaptation \(LoRA\) and its variants have become widely adopted because they significantly reduce memory and compute requirements while preserving model performance\.
LoRA introduces trainable low\-rank matrices into Transformer layers while keeping the original weights fixed\. Several variants extend this approach through quantization, adaptive rank allocation, sparsity mechanisms, or expert routing\. Table[10](https://arxiv.org/html/2605.19723#S5.T10)summarizes representative LoRA\-based techniques\.
Table 10:Representative variants of LoRA for parameter\-efficient fine\-tuning\.Although the preceding discussion outlines the general LLM training pipeline, mathematical reasoning places additional demands on both model adaptation and training design\. In practice, researchers have introduced a range of math\-specific modifications spanning corpus design, reasoning supervision, external tool integration, and verification\-oriented training\. Table[11](https://arxiv.org/html/2605.19723#S5.T11)summarizes these adaptations and clarifies the distinct role each plays in improving mathematical reasoning performance\.
These adaptations show that progress in mathematical reasoning rarely depends on architecture alone\. Instead, performance emerges from the interaction between mathematical data, supervision format, symbolic support, and verification mechanisms\. At the same time, many of these strategies introduce their own bottlenecks, including contamination risk, synthetic bias, verifier dependence, and limited transfer to structurally novel problems\. These tensions are discussed in the following subsection\.
Table 11:Math\-specific architectural and training adaptations for improving reasoning in LLMs\.
### 5\.4Training Trends and Challenges
Despite significant progress, several challenges remain in training LLMs for complex reasoning tasks\. One major limitation arises from the mismatch between training objectives and reasoning behavior\. Language models optimize next token prediction rather than logical correctness, which can produce fluent yet incorrect reasoning chains\(Lindseyet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib61)\)\.
Another challenge involves parametric knowledge bias\. During pretraining, models store large amounts of knowledge in their parameters\. While this internal knowledge improves downstream performance, it can also lead to hallucinations when the model relies on memorized patterns rather than input context\(Jiet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib46); Longpreet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib66)\)\.
Large\-scale training also raises concerns about memorization and data leakage\. Models may reproduce segments of training data, which introduces risks related to copyright violations, privacy exposure, and sensitive data leakage\(Feldman,[2020](https://arxiv.org/html/2605.19723#bib.bib24); Shiet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib90); Carliniet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib10); Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43)\)\. Addressing these issues requires improved training objectives, better dataset curation practices, and evaluation methods that emphasize reasoning faithfulness and data safety\.
## 6Evaluation of Mathematical Reasoning
Evaluating the performance of LLMs on mathematical reasoning tasks remains challenging\. Unlike human evaluation, model outputs often consist of fluent natural language explanations that appear coherent while containing logical or computational errors\(Ahnet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib3)\)\. Mathematical reasoning tasks involve structured symbolic manipulation and multi\-step derivations, making them fundamentally different from standard text generation tasks\. Consequently, evaluating only the final answer is often insufficient\. Reliable evaluation must assess both the correctness of the final solution and the validity of intermediate reasoning steps to avoid false positives and false negatives\(Wanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib101)\)\.
Researchers therefore employ several metrics to measure model performance on mathematical reasoning tasks\. Traditional metrics such as accuracy and exact match compare the generated answer with the reference solution\. Additional measures including F1\-score, Macro F1\-score, True Positive Rate \(TPR\), and True Negative Rate \(TNR\) are also used to measure alignment between generated outputs and ground truth\(Bostrom and Durrett,[2020](https://arxiv.org/html/2605.19723#bib.bib8); Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43)\)\. However, these metrics primarily evaluate outcome correctness and provide limited insight into the reasoning process that produced the answer\.
Mathematical reasoning can be assessed at multiple levels, from final\-answer correctness to stepwise derivational validity, no single metric provides a complete picture of model performance\. Table[12](https://arxiv.org/html/2605.19723#S6.T12)compares the most common evaluation approaches and highlights the trade\-off between simplicity, diagnostic value, and verification rigor\.
Table 12:Evaluation metrics and verification approaches for mathematical reasoning in LLMs\.The comparison makes clear that widely used metrics such as accuracy and exact match remain attractive because they are simple and reproducible, yet they fail to capture whether a model reasoned correctly\. More diagnostic approaches exist, but they are harder to scale and often require structured outputs, executable traces, or formal verification environments\. This limitation explains why answer\-level evaluation remains dominant despite its conceptual weakness\.
Accuracy based evaluation is particularly problematic for mathematical reasoning\. A model may produce the correct answer through memorization or shortcut pattern recognition rather than genuine reasoning, while correct reasoning chains may be penalized due to minor numerical deviations such as rounding errors\. As a result, accuracy metrics fail to distinguish robust reasoning ability from superficial pattern matching\.
To address this limitation, recent work evaluates reasoning processes directly\. Metrics such as skill success rate, secondary skill success rate, completion rate, and calculation success rate measure the model’s ability to execute specific reasoning skills required to solve mathematical problems\(Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110); Didolkaret al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib18)\)\. For formal mathematical reasoning tasks, syntactic and semantic correctness can also be verified using automated theorem proving tools such as the Isabelle proof assistant\(Azerbayevet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib7)\)\.
Recent research further explores process\-level evaluation approaches\. Automated process supervision techniques analyze intermediate reasoning steps rather than only final outputs\. Methods based on Monte Carlo Tree Search enable efficient error localization, while hybrid verification approaches combine symbolic solvers with executable reasoning traces to improve reliability\(Ren,[2025](https://arxiv.org/html/2605.19723#bib.bib83)\)\. However, the diversity and complexity of reasoning patterns produced by LLMs make automated verification difficult\(Herbert,[2019](https://arxiv.org/html/2605.19723#bib.bib36),[2021](https://arxiv.org/html/2605.19723#bib.bib37)\)\. Consequently, verifying intermediate reasoning steps remains computationally expensive and often requires human oversight\.
### 6\.1Limitations and Gaps
Despite the availability of multiple evaluation metrics, several limitations remain\. Accuracy remains the most widely used metric for mathematical reasoning benchmarks\(Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43); Singh and Strouse,[2024](https://arxiv.org/html/2605.19723#bib.bib91); Guanet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib31); Rhomrasiet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib84)\)\. However, accuracy evaluates only the final answer and does not capture the correctness of intermediate reasoning steps\. Models may therefore achieve high scores through memorization or heuristic shortcuts rather than principled reasoning\.
Evaluation consistency also presents challenges\. LLMs may generate different reasoning paths or error localization patterns for the same problem across multiple runs, complicating reproducible evaluation\(Ren,[2025](https://arxiv.org/html/2605.19723#bib.bib83)\)\. Furthermore, discrepancies frequently arise between explicit reasoning traces generated by the model and the internal mechanisms that produce the final answer\. Although interpretability methods such as attribution graphs provide partial insights into model behavior, they do not fully reveal the reasoning processes used by the model\(Lindseyet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib61)\)\.
Intermediate evaluation metrics such as relative error \(RE\) introduce additional limitations because they can be dominated by outliers or disproportionately large deviations\(Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110)\)\. These challenges highlight the need for more robust evaluation frameworks that jointly assess answer correctness, reasoning validity, and model reliability in mathematical problem solving\.
Overall, current evaluation protocols still measure mathematical reasoning primarily through final\-answer correctness\. While convenient, this approach provides limited insight into the reasoning processes that produce correct answers\. A model may arrive at the correct solution through memorization, heuristic shortcuts, or reasoning paths that are logically inconsistent\. Future evaluation frameworks must therefore move beyond answer\-level metrics and incorporate scalable process\-level verification to better assess genuine reasoning capability\.
## 7Discussion and Future Directions
This section synthesizes insights across datasets, architectures, training strategies, and evaluation protocols to analyze how their limitations jointly influence the progress of LLM\-based mathematical reasoning\. Figure[4](https://arxiv.org/html/2605.19723#S7.F4)summarizes how these limitations interact across the reasoning pipeline\.
LINGUISTICPROCESSINGZONEInputs: Word Problems, LanguageStrongLinguistic Fluency\(Input Flow\)Limitations: Dataset contamination \(Sec\. 4\) Synthetic CoT bias Tokenization limitations Weak numerical encoding Symbol ambiguity Pattern imitation vs\. reasoning \(Sec\. 5\) Limited symbolic manipulation Unfaithful CoT Parametric knowledge bias/ Memorization Accuracy\-only metrics \(Sec\. 6\) No verification/ Weak fault localizationSYMBOLICLOGICZONEOutputs:Correct Answers,Step\-by\-Step ProofsReliableMathematical Reasoning\(Target Output\)HIGH LINGUISTIC FLOWCHOKED REASONING FLOWTHE TRANSFORMATION GAPReasoning Breakdown
Figure 4:Challenge pipeline summarizing the interconnected limitations affecting mathematical reasoning in LLMs\. The “transformation gap” represents the failure to translate linguistic fluency into reliable symbolic logic due to compounding technical limitations\.### 7\.1Cross\-Section Analysis
Challenges in LLM\-based mathematical reasoning are deeply interconnected rather than isolated\. Limitations across datasets \(Section[4](https://arxiv.org/html/2605.19723#S4)\), model architectures, training strategies \(Section[5](https://arxiv.org/html/2605.19723#S5)\), and evaluation frameworks \(Section[6](https://arxiv.org/html/2605.19723#S6)\) collectively form a reinforcing cycle that restricts progress\.
A central observation emerging from the preceding sections is that current limitations in mathematical reasoning arise from interactions across multiple components of the LLM pipeline rather than from a single bottleneck\. Dataset design, representation choices, training supervision, and evaluation protocols jointly shape the reasoning behavior exhibited by modern models\(Lianget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib60)\)\. Improvements in one component, such as larger datasets\(Cobbeet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib14); Hendryckset al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib35)\)or more advanced prompting strategies, often expose weaknesses in others\(Jiet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib46); Longpreet al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib66)\), including symbolic representation\(Feldman,[2020](https://arxiv.org/html/2605.19723#bib.bib24)\), reasoning faithfulness, and evaluation reliability\.
Training strategies further reinforce shallow reasoning patterns\. Instruction tuning\(Minet al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib72); Parnami and Lee,[2022](https://arxiv.org/html/2605.19723#bib.bib77)\)and reinforcement learning from human feedback \(RLHF\)\(Ouyanget al\.,[2022](https://arxiv.org/html/2605.19723#bib.bib76)\)significantly improve instruction following and response fluency\. However, they may also produce*unfaithful CoT reasoning*\(Lindseyet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib61)\), where generated explanations appear plausible but do not correspond to the model’s internal decision process\. Consequently, models often perform well on in\-distribution tasks while failing on structurally novel problems requiring deeper conceptual reasoning\(Ahnet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib3); Huang and Chang,[2023](https://arxiv.org/html/2605.19723#bib.bib42)\)\.
Evaluation methodologies further obscure these limitations\. The dominant metric, accuracy\(Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43); Rhomrasiet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib84); Singh and Strouse,[2024](https://arxiv.org/html/2605.19723#bib.bib91)\), evaluates only final answers and does not verify reasoning correctness\(Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110)\)\. Models may therefore receive credit for correct answers derived through incorrect reasoning, while correct reasoning processes may be penalized due to small numerical deviations\(Wanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib101)\)\. In addition, current evaluation frameworks lack robust mechanisms for reasoning verification and fault localization\(Ren,[2025](https://arxiv.org/html/2605.19723#bib.bib83)\)\.
Finally, fundamental challenges remain in the representation and processing of mathematical structures\. LLMs frequently struggle with mathematical symbols, structured expressions, and numerical formats\(Rajaramanet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib82); Shalytet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib88)\)\. Although improvements in tokenization and structured representations can mitigate certain arithmetic errors\(Singh and Strouse,[2024](https://arxiv.org/html/2605.19723#bib.bib91)\), persistent failures in numerical reasoning indicate that simply scaling model size cannot resolve these representational limitations\.
The cross\-sectional analysis reveals that current limitations are not merely abstract methodological concerns; they manifest as recurring and recognizable failure patterns during mathematical problem solving\. Table[13](https://arxiv.org/html/2605.19723#S7.T13)summarizes the most common failure modes observed in LLM\-based mathematical reasoning and links them to likely causes and broader implications for reliability and evaluation\.
Table 13:Common failure modes in LLM\-based mathematical reasoning and their broader implications\.These failure modes reinforce a central conclusion of this survey: mathematical reasoning errors in LLMs are rarely isolated defects\. Rather, they emerge from coupled weaknesses in representation, supervision, inference control, and evaluation design\. Accordingly, the most promising future research directions are those that address these failures jointly rather than optimizing isolated components in separation\.
### 7\.2Open Problems and Future Directions
The analysis above highlights several important research directions for advancing mathematical reasoning in LLMs\.
#### 7\.2\.1Beyond Synthetic Reasoning Datasets
Heavy reliance on synthetic CoT datasets\(Guanet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib31)\)and diminishing returns from dataset refinement\(Huanget al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib44); Tanget al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib95)\)indicate the need for higher\-quality training resources\. Future research should prioritize human\-curated datasets that capture diverse reasoning strategies and explicitly test out\-of\-distribution generalization\(Huanget al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib43)\)\. Such datasets should emphasize complex multi\-step reasoning processes that cannot be easily memorized\.
#### 7\.2\.2Faithful Reasoning Mechanisms
Unfaithful reasoning remains a central limitation of current LLMs\(Lindseyet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib61)\)\. Future work should explore training paradigms and architectural modifications that encourage reasoning processes aligned with the model’s internal computations\. Potential approaches include improved reward modeling for reasoning tasks\(Muet al\.,[2024](https://arxiv.org/html/2605.19723#bib.bib74)\), structured reasoning supervision, and architectures designed for greater interpretability\.
#### 7\.2\.3Process\-Level Evaluation
Evaluation must extend beyond final\-answer accuracy\(Yuanet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib110)\)\. Developing scalable methods for assessing reasoning processes remains a major open challenge\. Promising directions include automated process supervision\(Ren,[2025](https://arxiv.org/html/2605.19723#bib.bib83)\), hybrid evaluation frameworks combining symbolic verifiers with neural models\(Azerbayevet al\.,[2023](https://arxiv.org/html/2605.19723#bib.bib7)\), and reasoning\-aware evaluation metrics capable of fault localization\.
#### 7\.2\.4Numerical and Symbolic Representation
Improving how LLMs represent and manipulate numerical and symbolic information remains a fundamental challenge\. Current tokenization schemes often fail to preserve mathematical structure\(Rajaramanet al\.,[2025](https://arxiv.org/html/2605.19723#bib.bib82)\)\. Future work should explore specialized representations and architectures that treat numbers and symbols as structured entities rather than simple text tokens\. Structure\-aware encoders and symbolic representations\(Penget al\.,[2021](https://arxiv.org/html/2605.19723#bib.bib79); Zanibbiet al\.,[2011](https://arxiv.org/html/2605.19723#bib.bib111)\)may provide more reliable foundations for mathematical reasoning\.
## 8Conclusions
Mathematical reasoning represents one of the most demanding benchmarks for evaluating LLMs because it requires structured symbolic manipulation, multi\-step deduction, and logically consistent inference across long reasoning chains\. This survey synthesizes current research on mathematical reasoning in LLMs by examining benchmark datasets, architectural adaptations, training strategies, and evaluation protocols across the full model development lifecycle\. The analysis shows that progress depends not only on scaling models or expanding datasets but also on improving structured representations, reasoning supervision, and evaluation methods that verify intermediate reasoning steps rather than final answers alone\. Advancing this area therefore requires tighter integration between language modeling, symbolic reasoning mechanisms, and robust evaluation frameworks capable of assessing reasoning faithfulness and generalization across diverse mathematical tasks\. Future progress in mathematical reasoning will likely depend on closer integration between neural language models and symbolic reasoning systems, supported by benchmarks and evaluation frameworks that measure reasoning processes rather than final answers alone\.
## Appendix AComprehensive Reference Coverage and Categorization
This appendix provides a structured overview of the references used in this survey\. Given the large number of cited works \(approximately 120\), we present a categorized subset of representative and influential references for each thematic area\. Given the breadth of included works spanning foundational theories, datasets, architectures, and evaluation methodologies, we organize the references into thematic categories for clarity and reproducibility\.
### A\.1Foundational Works in Language Models
These works establish the theoretical and architectural foundations of modern large language models, including transformer\-based scaling and pretraining paradigms\.
- •Brownet al\.\([2020](https://arxiv.org/html/2605.19723#bib.bib9)\)– Few\-shot learning with GPT models
- •Devlinet al\.\([2019](https://arxiv.org/html/2605.19723#bib.bib16)\)– Pretraining bidirectional transformers
- •Chowdheryet al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib13)\)– Scaling language models with Pathways
- •Kaplanet al\.\([2020](https://arxiv.org/html/2605.19723#bib.bib48)\); Hernandezet al\.\([2021](https://arxiv.org/html/2605.19723#bib.bib38)\); Raeet al\.\([2022](https://arxiv.org/html/2605.19723#bib.bib81)\)– Scaling laws for LLMs
- •Hoffmannet al\.\([2022](https://arxiv.org/html/2605.19723#bib.bib39)\)– Compute\-optimal training strategies
### A\.2Mathematical Reasoning Benchmarks and Datasets
These references introduce datasets and benchmarks used to evaluate mathematical reasoning capabilities of LLMs\.
- •Hendryckset al\.\([2021](https://arxiv.org/html/2605.19723#bib.bib35)\)– MATH dataset
- •Azerbayevet al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib7)\)– ProofNet formal reasoning dataset
- •Mishraet al\.\([2022](https://arxiv.org/html/2605.19723#bib.bib73)\)– Unified reasoning benchmark
- •Heet al\.\([2024](https://arxiv.org/html/2605.19723#bib.bib34)\); Gaoet al\.\([2024](https://arxiv.org/html/2605.19723#bib.bib28)\)– Olympiad\-level benchmarks
- •Chernyshevet al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib12)\)– University\-level evaluation \(U\-MATH\)
- •Fanget al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib22)\)– Large\-scale evaluation dataset
- •Huanget al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib43)\)– Robustness\-focused perturbation dataset
### A\.3Reasoning Techniques and Prompting Strategies
This category includes works that propose improvements to reasoning in LLMs through prompting and inference\-time strategies\.
- •Wanget al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib100)\)– Self\-consistency decoding
- •Imaniet al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib45)\)– Prompt engineering for math reasoning
- •Han and Wang \([2024](https://arxiv.org/html/2605.19723#bib.bib33)\)– Role\-play prompting analysis
- •Aaliet al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib1)\)– Structured prompting for evaluation
### A\.4Training and Optimization Methods
These works focus on training paradigms, reinforcement learning, and optimization techniques used to enhance LLM reasoning\.
- •Ouyanget al\.\([2022](https://arxiv.org/html/2605.19723#bib.bib76)\)– RLHF training
- •Achiamet al\.\([2017](https://arxiv.org/html/2605.19723#bib.bib2)\); Schulmanet al\.\([2015](https://arxiv.org/html/2605.19723#bib.bib87)\)– Policy optimization methods
- •Fuet al\.\([2026](https://arxiv.org/html/2605.19723#bib.bib27)\); Miaoet al\.\([2024](https://arxiv.org/html/2605.19723#bib.bib71)\); Guoet al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib32)\)– Reward modeling
- •Muet al\.\([2024](https://arxiv.org/html/2605.19723#bib.bib74)\)– Rule\-based reward strategies
### A\.5Parameter\-Efficient Fine\-Tuning \(PEFT\)
These references explore efficient adaptation techniques for large\-scale models\.
- •Huet al\.\([2021](https://arxiv.org/html/2605.19723#bib.bib41)\)– LoRA
- •Dettmerset al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib15)\)– QLoRA
- •Valipouret al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib99)\); Liuet al\.\([2024](https://arxiv.org/html/2605.19723#bib.bib64)\); Dinget al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib19)\)– Variants of low\-rank adaptation
- •Lianget al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib60)\)– Survey on PEFT methods
### A\.6Symbolic and Tool\-Augmented Reasoning
These works integrate symbolic reasoning or external tools with LLMs\.
- •Luet al\.\([2021](https://arxiv.org/html/2605.19723#bib.bib67)\)– Geometry reasoning with symbolic systems
- •Li \([2024](https://arxiv.org/html/2605.19723#bib.bib58)\)– Tool\-integrated reasoning
- •Lample and Charton \([2020](https://arxiv.org/html/2605.19723#bib.bib54)\)– Neural symbolic mathematics
### A\.7Evaluation, Faithfulness, and Limitations
This category includes works addressing hallucination, reasoning reliability, and evaluation challenges\.
- •Jiet al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib46)\)– Hallucination in LLMs
- •Wanget al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib101)\)– False positives in reasoning
- •Herbert \([2019](https://arxiv.org/html/2605.19723#bib.bib36),[2021](https://arxiv.org/html/2605.19723#bib.bib37)\)– Challenges in assessing reasoning
- •Carliniet al\.\([2022](https://arxiv.org/html/2605.19723#bib.bib10)\); Shiet al\.\([2024](https://arxiv.org/html/2605.19723#bib.bib90)\)– Memorization and data leakage
### A\.8Data Quality and Representation Challenges
These works highlight issues in data quality, tokenization, and numerical representation\.
- •Gonget al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib30)\); Karra and Lasfar \([2023](https://arxiv.org/html/2605.19723#bib.bib49)\)– Dataset quality
- •Bostrom and Durrett \([2020](https://arxiv.org/html/2605.19723#bib.bib8)\); Singh and Strouse \([2024](https://arxiv.org/html/2605.19723#bib.bib91)\); Rajaramanet al\.\([2025](https://arxiv.org/html/2605.19723#bib.bib82)\)– Tokenization limitations
- •Golkaret al\.\([2023](https://arxiv.org/html/2605.19723#bib.bib29)\)– Numerical encoding methods
### A\.9Historical and Theoretical Foundations of Mathematical Reasoning
These references provide broader context for mathematical reasoning and cognition\.
- •Feigenbaum and Feldman \([1963](https://arxiv.org/html/2605.19723#bib.bib23)\)– Early AI reasoning
- •Kline \([1990](https://arxiv.org/html/2605.19723#bib.bib50)\)– History of mathematics
- •Devlin and Gray \([1998](https://arxiv.org/html/2605.19723#bib.bib17)\)– Language of mathematics
### A\.10Additional Surveys and Meta\-Analyses
These works complement the present survey and provide broader perspectives\.
- •Forootani \([2025](https://arxiv.org/html/2605.19723#bib.bib25)\)– Mathematical reasoning survey
- •Huang and Chang \([2023](https://arxiv.org/html/2605.19723#bib.bib42)\)– Reasoning in LLMs
- •Liuet al\.\([2025b](https://arxiv.org/html/2605.19723#bib.bib65)\)– Mathematical language models survey
## Appendix BReproducibility Statement
All references included in this survey were collected through a structured and iterative process combining keyword\-based search, citation chaining, and manual filtering\. Efforts were made to ensure diversity across datasets, methods, and evaluation strategies\.
Despite these efforts, the rapidly evolving nature of LLM research implies that some recent works may not be included\.
## Appendix CLimitations of Reference Coverage
- •Rapid publication cycles in arXiv may lead to missing recent contributions
- •Some references may overlap across multiple categories
- •Categorization involves subjective judgment
Future updates to this survey may incorporate automated literature mining techniques for improved coverage\.
## References
- A\. Aali, M\. A\. Mohsin, V\. Bikia, A\. Singhvi, R\. Gaus, S\. Bedi, H\. Cui, M\. Fuentes,et al\.\(2025\)Structured prompting enables more robust evaluation of language models\.External Links:2511\.20836,[Link](https://arxiv.org/abs/2511.20836)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I3.i4.p1.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p1.1)\.
- J\. Achiam, D\. Held, A\. Tamar, and P\. Abbeel \(2017\)Constrained policy optimization\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 22–31\.External Links:[Link](https://proceedings.mlr.press/v70/achiam17a.html)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I4.i2.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p4.1)\.
- J\. Ahn, R\. Verma, R\. Lou, D\. Liu, R\. Zhang, and W\. Yin \(2024\)Large language models for mathematical reasoning: progresses and challenges\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop,St\. Julian’s, Malta,pp\. 225–237\.External Links:[Link](https://aclanthology.org/2024.eacl-srw.17/),[Document](https://dx.doi.org/10.18653/v1/2024.eacl-srw.17)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p4.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p3.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p4.1),[§6](https://arxiv.org/html/2605.19723#S6.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p3.1)\.
- S\. Alavi Naeini, R\. Saqur, M\. Saeidi, J\. Giorgi, and B\. Taati \(2023\)Large language models are fixated by red herrings: exploring creative problem solving and einstellung effect using the only connect wall dataset\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 5631–5652\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/11e3e0f1b29dcd31bd0952bfc1357f68-Paper-Datasets_and_Benchmarks.pdf)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- J\. Anton, G\. Cosentino, K\. Sharma, M\. Gelsomini, M\. Mok, M\. Giannakos, and D\. Abrahamson \(2025\)The human condition: modal and interactive advantages of teacher over ai feedback on children’s mathematical performance\.InProceedings of the 24th Interaction Design and Children,pp\. 183–203\.External Links:ISBN 9798400714733,[Link](https://doi.org/10.1145/3713043.3728863)Cited by:[§3\.4](https://arxiv.org/html/2605.19723#S3.SS4.p2.1)\.
- R\. Ausbrookset al\.\(2003\)Mathematical markup language \(mathml\) version 2\.0\.W3C Recommendation\.External Links:[Link](https://www.w3.org/TR/MathML2/)Cited by:[§4\.2](https://arxiv.org/html/2605.19723#S4.SS2.p2.1)\.
- Z\. Azerbayev, B\. Piotrowski, H\. Schoelkopf, E\. W\. Ayers, D\. Radev, and J\. Avigad \(2023\)ProofNet: autoformalizing and formally proving undergraduate\-level mathematics\.External Links:2302\.12433,[Link](https://arxiv.org/abs/2302.12433)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I2.i2.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1),[§6](https://arxiv.org/html/2605.19723#S6.p6.1),[§7\.2\.3](https://arxiv.org/html/2605.19723#S7.SS2.SSS3.p1.1)\.
- K\. Bostrom and G\. Durrett \(2020\)Byte pair encoding is suboptimal for language model pretraining\.InFindings of the Association for Computational Linguistics: EMNLP 2020,Online,pp\. 4617–4624\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.414/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.414)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I8.i2.p1.1),[§6](https://arxiv.org/html/2605.19723#S6.p2.1)\.
- T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss,et al\.\(2020\)Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I1.i1.p1.1),[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p2.1)\.
- N\. Carlini, M\. Jagielski, C\. Zhang, N\. Papernot, A\. Terzis, and F\. Tramer \(2022\)The privacy onion effect: memorization is relative\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 13263–13276\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/564b5f8289ba846ebc498417e834c253-Paper-Conference.pdf)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I7.i4.p1.1),[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p3.1)\.
- W\. Chen \(2023\)Large language models are few\(1\)\-shot table reasoners\.InFindings of the Association for Computational Linguistics: EACL 2023,A\. Vlachos and I\. Augenstein \(Eds\.\),Dubrovnik, Croatia,pp\. 1120–1130\.External Links:[Link](https://aclanthology.org/2023.findings-eacl.83/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-eacl.83)Cited by:[§5\.3\.2](https://arxiv.org/html/2605.19723#S5.SS3.SSS2.p2.1)\.
- K\. Chernyshev, V\. Polshkov, V\. Stepanov, A\. Myasnikov, E\. Artemova, A\. Miasnikov, and S\. Tilga \(2025\)U\-MATH: a university\-level benchmark for evaluating mathematical skills in large language models\.InProceedings of the Fourth Workshop on Generation, Evaluation and Metrics \(GEM²\),Vienna, Austria and virtual meeting,pp\. 974–1001\.External Links:[Link](https://aclanthology.org/2025.gem-1.77/),ISBN 979\-8\-89176\-261\-9Cited by:[5th item](https://arxiv.org/html/2605.19723#A1.I2.i5.p1.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p2.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- A\. Chowdhery, S\. Narang, J\. Devlin, M\. Bosma, G\. Mishra, A\. Roberts, P\. Barham, H\. W\. Chung, C\. Sutton, S\. Gehrmann, P\. Schuh,et al\.\(2023\)PaLM: scaling language modeling with pathways\.Journal of Machine Learning Research24\(240\),pp\. 1–113\.External Links:[Link](http://jmlr.org/papers/v24/22-1144.html)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I1.i3.p1.1),[§4\.3](https://arxiv.org/html/2605.19723#S4.SS3.p2.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman \(2021\)Training verifiers to solve math word problems\.External Links:2110\.14168,[Link](https://arxiv.org/abs/2110.14168)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1),[§4](https://arxiv.org/html/2605.19723#S4.p5.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p2.1)\.
- N\. Dave, D\. Kifer, C\. L\. Giles, and A\. Mali \(2024\)Investigating symbolic capabilities of large language models\.InProceedings of the First International Workshop on Logical Foundations of Neuro\-Symbolic AI \(LNSAI 2024\),G\. Cima, M\. Console, V\. Gutièrrez\-Basulto, P\. Hitzler, and M\. Lenzerini \(Eds\.\),CEUR Workshop Proceedings, Vol\.3819\.External Links:[Link](https://ceur-ws.org/Vol-3819/paper2.pdf)Cited by:[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p2.1)\.
- T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. Zettlemoyer \(2023\)QLoRA: efficient finetuning of quantized llms\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 10088–10115\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I5.i2.p1.1),[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.3.2.1.1.1)\.
- J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova \(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),Minneapolis, Minnesota,pp\. 4171–4186\.External Links:[Link](https://aclanthology.org/N19-1423/),[Document](https://dx.doi.org/10.18653/v1/N19-1423)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I1.i2.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p2.1)\.
- K\. Devlin and J\. Gray \(1998\)The language of mathematics: making the invisible visible\.Nature396\(6710\),pp\. 428–428\.External Links:[Link](https://www.math.chalmers.se/~ulfp/Review/langmath.pdf)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I9.i3.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p1.1)\.
- A\. Didolkar, A\. Goyal, N\. R\. Ke, S\. Guo, M\. Valko, T\. Lillicrap, D\. Rezende, Y\. Bengio, M\. Mozer, and S\. Arora \(2024\)Metacognitive capabilities of llms: an exploration in mathematical problem solving\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 19783–19812\.External Links:[Document](https://dx.doi.org/10.52202/079017-0623),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/2318d75a06437eaa257737a5cf3ab83c-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p4.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p3.1),[§3\.4](https://arxiv.org/html/2605.19723#S3.SS4.p2.1),[§6](https://arxiv.org/html/2605.19723#S6.p6.1)\.
- N\. Ding, X\. Lv, Q\. Wang, Y\. Chen, B\. Zhou, Z\. Liu, and M\. Sun \(2023\)Sparse low\-rank adaptation of pre\-trained language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 4133–4145\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.252/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.252)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I5.i3.p1.1),[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.6.5.1.1.1)\.
- S\. Dou, E\. Zhou, Y\. Liu, S\. Gao, W\. Shen, L\. Xiong, Y\. Zhou, X\. Wang, Z\. Xi, X\. Fan,et al\.\(2024\)LoRAMoE: alleviating world knowledge forgetting in large language models via MoE\-style plugin\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 1932–1945\.External Links:[Link](https://aclanthology.org/2024.acl-long.106/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.106)Cited by:[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.7.6.1.1.1)\.
- N\. Duan, D\. Tang, and M\. Zhou \(2020\)Machine reasoning: technology, dilemma and future\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts,Online,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-tutorials.1),[Link](https://aclanthology.org/2020.emnlp-tutorials.1/)Cited by:[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p3.1)\.
- M\. Fang, X\. Wan, F\. Lu, F\. Xing, and K\. Zou \(2025\)MathOdyssey: Benchmarking Mathematical Problem\-Solving Skills in Large Language Models Using Odyssey Math Data\.12\(1\),pp\. 1392\.External Links:ISSN 2052\-4463,[Document](https://dx.doi.org/10.1038/s41597-025-05283-3),[Link](https://doi.org/10.1038/s41597-025-05283-3)Cited by:[6th item](https://arxiv.org/html/2605.19723#A1.I2.i6.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- E\. A\. Feigenbaum and J\. Feldman \(Eds\.\) \(1963\)Computers and thought\.McGraw\-Hill,New York\.External Links:[Link](https://archive.org/details/computersthought00feig)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I9.i1.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p2.1)\.
- V\. Feldman \(2020\)Does learning require memorization? a short tale about a long tail\.InProceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing,STOC 2020,New York, NY, USA,pp\. 954–959\.External Links:ISBN 9781450369794,[Link](https://doi.org/10.1145/3357713.3384290),[Document](https://dx.doi.org/10.1145/3357713.3384290)Cited by:[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p3.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p2.1)\.
- A\. Forootani \(2025\)A survey on mathematical reasoning and optimization with large language models\.External Links:2503\.17726,[Link](https://arxiv.org/abs/2503.17726)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I10.i1.p1.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p1.1)\.
- J\. Fu, X\. Zhao, C\. Yao, H\. Wang, Q\. Han, and Y\. Xiao \(2026\)Reward shaping to mitigate reward hacking in rlhf\.External Links:2502\.18770,[Link](https://arxiv.org/abs/2502.18770)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I4.i3.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1)\.
- Y\. Fu \(2025\)Improving complex reasoning in large language models\.PhD thesis,The University of EdinburghSchool of Informatics,Edinburgh, United Kingdom\.External Links:[Document](https://dx.doi.org/10.7488/era/6083),[Link](https://doi.org/10.7488/era/6083)Cited by:[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p2.1)\.
- B\. Gao, F\. Song, Z\. Yang, Z\. Cai, Y\. Miao, Q\. Dong, L\. Li, C\. Ma, L\. Chen, R\. Xu, Z\. Tang, B\. Wang,et al\.\(2024\)Omni\-math: a universal olympiad level mathematic benchmark for large language models\.External Links:2410\.07985,[Link](https://arxiv.org/abs/2410.07985)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I2.i4.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- S\. Golkar, M\. Pettee, M\. Eickenberg, A\. Bietti, M\. Cranmer, G\. Krawezik, F\. Lanusse, M\. McCabe, R\. Ohana, L\. Parker,et al\.\(2023\)XVal: a continuous number encoding for large language models\.InNeurIPS 2023 AI for Science Workshop,External Links:[Link](https://openreview.net/forum?id=KHDMZtoF4i)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I8.i3.p1.1),[§4\.3](https://arxiv.org/html/2605.19723#S4.SS3.p2.1)\.
- Y\. Gong, G\. Liu, Y\. Xue, R\. Li, and L\. Meng \(2023\)A survey on dataset quality in machine learning\.Information and Software Technology162,pp\. 107268\.External Links:ISSN 0950\-5849,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.infsof.2023.107268),[Link](https://www.sciencedirect.com/science/article/pii/S0950584923001222)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I8.i1.p1.1),[§4](https://arxiv.org/html/2605.19723#S4.p1.1)\.
- X\. Guan, L\. L\. Zhang, Y\. Liu, N\. Shang, Y\. Sun, Y\. Zhu, F\. Yang, and M\. Yang \(2025\)RStar\-math: small llms can master math reasoning with self\-evolved deep thinking\.External Links:2501\.04519,[Link](https://arxiv.org/abs/2501.04519)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p4.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p3.1),[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p1.1),[§7\.2\.1](https://arxiv.org/html/2605.19723#S7.SS2.SSS1.p1.1)\.
- J\. Guo, Z\. Chi, L\. Dong, Q\. Dong, X\. Wu, S\. Huang, and F\. Wei \(2025\)Reward reasoning models\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=V8Kbz7l2cr)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I4.i3.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p3.1)\.
- Z\. Han and Z\. Wang \(2024\)Rethinking the role\-play prompting in mathematical reasoning tasks\.InProceedings of the 1st Workshop on Efficiency, Security, and Generalization of Multimedia Foundation Models,ESGMFM ’24,New York, NY, USA,pp\. 13–17\.External Links:ISBN 9798400711916,[Link](https://doi.org/10.1145/3688864.3689149),[Document](https://dx.doi.org/10.1145/3688864.3689149)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I3.i3.p1.1),[§3\.4](https://arxiv.org/html/2605.19723#S3.SS4.p2.1)\.
- C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang,et al\.\(2024\)OlympiadBench: a challenging benchmark for promoting AGI with olympiad\-level bilingual multimodal scientific problems\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 3828–3850\.External Links:[Link](https://aclanthology.org/2024.acl-long.211/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.211)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I2.i4.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021\)Measuring mathematical problem solving with the math dataset\.External Links:2103\.03874,[Link](https://arxiv.org/abs/2103.03874)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I2.i1.p1.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1),[§4](https://arxiv.org/html/2605.19723#S4.p5.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p2.1)\.
- S\. Herbert \(2019\)Challenges in assessing mathematical reasoning\.\.Mathematics Education Research Group of Australasia\.Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I7.i3.p1.1),[§6](https://arxiv.org/html/2605.19723#S6.p7.1)\.
- S\. Herbert \(2021\)Overcoming challenges in assessing mathematical reasoning\.Australian Journal of Teacher Education46\(8\),pp\. 17–30\.External Links:[Document](https://dx.doi.org/10.14221/ajte.2021v46n8.2),[Link](https://doi.org/10.14221/ajte.2021v46n8.2)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I7.i3.p1.1),[§6](https://arxiv.org/html/2605.19723#S6.p7.1)\.
- D\. Hernandez, J\. Kaplan, T\. Henighan, and S\. McCandlish \(2021\)Scaling laws for transfer\.External Links:2102\.01293,[Link](https://arxiv.org/abs/2102.01293)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I1.i4.p1.1),[§5\.2](https://arxiv.org/html/2605.19723#S5.SS2.p2.1)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, Hendricks,et al\.\(2022\)An empirical analysis of compute\-optimal large language model training\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 30016–30030\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/c1e2faff6f588870935f114ebe04a3e5-Paper-Conference.pdf)Cited by:[5th item](https://arxiv.org/html/2605.19723#A1.I1.i5.p1.1),[§5\.2](https://arxiv.org/html/2605.19723#S5.SS2.p2.1)\.
- M\. J\. Hosseini, H\. Hajishirzi, O\. Etzioni, and N\. Kushman \(2014\)Learning to solve arithmetic word problems with verb categorization\.InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Doha, Qatar,pp\. 523–533\.External Links:[Link](https://aclanthology.org/D14-1058/),[Document](https://dx.doi.org/10.3115/v1/D14-1058)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p2.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2021\)LoRA: low\-rank adaptation of large language models\.External Links:2106\.09685,[Link](https://arxiv.org/abs/2106.09685)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I5.i1.p1.1),[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.2.1.1.1.1)\.
- J\. Huang and K\. C\. Chang \(2023\)Towards reasoning in large language models: a survey\.InFindings of the Association for Computational Linguistics: ACL 2023,A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 1049–1065\.External Links:[Link](https://aclanthology.org/2023.findings-acl.67/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.67)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I10.i2.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p3.1)\.
- K\. Huang, J\. Guo, Z\. Li, X\. Ji, J\. Ge, W\. Li, Y\. Guo, T\. Cai, H\. Yuan, R\. Wang, Y\. Wu, M\. Yin, S\. Tang, Y\. Huang, C\. Jin, X\. Chen, C\. Zhang, and M\. Wang \(2025\)MATH\-perturb: benchmarking llms’ math reasoning abilities against hard perturbations\.External Links:2502\.06453,[Link](https://arxiv.org/abs/2502.06453)Cited by:[7th item](https://arxiv.org/html/2605.19723#A1.I2.i7.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p4.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p3.1),[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p3.1),[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p1.1),[§6](https://arxiv.org/html/2605.19723#S6.p2.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p4.1),[§7\.2\.1](https://arxiv.org/html/2605.19723#S7.SS2.SSS1.p1.1)\.
- Z\. Huang, H\. Zou, X\. Li, Y\. Liu, Y\. Zheng, E\. Chern, S\. Xia, Y\. Qin, W\. Yuan, and P\. Liu \(2024\)O1 replication journey – part 2: surpassing o1\-preview through simple distillation, big progress or bitter lesson?\.External Links:2411\.16489,[Link](https://arxiv.org/abs/2411.16489)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p3.1),[§7\.2\.1](https://arxiv.org/html/2605.19723#S7.SS2.SSS1.p1.1)\.
- S\. Imani, L\. Du, and H\. Shrivastava \(2023\)MathPrompter: mathematical reasoning using large language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 5: Industry Track\),Toronto, Canada,pp\. 37–42\.External Links:[Link](https://aclanthology.org/2023.acl-industry.4/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-industry.4)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I3.i2.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p1.1)\.
- Z\. Ji, N\. Lee, R\. Frieske, T\. Yu, D\. Su, Y\. Xu, E\. Ishii, Y\. J\. Bang, A\. Madotto, and P\. Fung \(2023\)Survey of hallucination in natural language generation\.ACM Comput\. Surv\.55\(12\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3571730),[Document](https://dx.doi.org/10.1145/3571730)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I7.i1.p1.1),[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p2.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p2.1)\.
- R\. Jiang, L\. Liu, and C\. Chen \(2025\)MoPE: mixture of prompt experts for parameter\-efficient and scalable multimodal fusion\.External Links:2403\.10568,[Link](https://arxiv.org/abs/2403.10568)Cited by:[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.7.6.1.1.1)\.
- J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei \(2020\)Scaling laws for neural language models\.External Links:2001\.08361,[Link](https://arxiv.org/abs/2001.08361)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I1.i4.p1.1),[§5\.2](https://arxiv.org/html/2605.19723#S5.SS2.p2.1)\.
- R\. Karra and A\. Lasfar \(2023\)Impact of data quality on question answering system performances\.Intelligent Automation & Soft Computing35\(1\),pp\. 335–349\.External Links:[Document](https://dx.doi.org/10.32604/iasc.2023.026695),[Link](https://www.techscience.com/iasc/v35n1/48146),ISSN 1079\-8587Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I8.i1.p1.1),[§4](https://arxiv.org/html/2605.19723#S4.p1.1)\.
- M\. Kline \(1990\)Mathematical thought from ancient to modern times\.Vol\.3,Oxford University Press,New York\.Note:Originally published in 1972; Oxford University Press paperback editionExternal Links:ISBN 0\-19\-506137\-3,LCCN 89\-25520,[Link](https://www.hlevkin.com/hlevkin/90MathPhysBioBooks/mathHistory/KLINE_MathematicalThoughtFromAncientToModernTimes3.pdf)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I9.i2.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p1.1)\.
- R\. Koncel\-Kedziorski, S\. Roy, A\. Amini, N\. Kushman, and H\. Hajishirzi \(2016\)MAWPS: a math word problem repository\.InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,San Diego, California,pp\. 1152–1157\.External Links:[Link](https://aclanthology.org/N16-1136/),[Document](https://dx.doi.org/10.18653/v1/N16-1136)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- G\. Y\. Kristianto, G\. Topic, and A\. Aizawa \(2016\)MCAT math retrieval system for ntcir\-12 mathir task\.InProceedings of the 12th NTCIR Conference on Evaluation of Information Access Technologies,External Links:[Link](https://research.nii.ac.jp/ntcir/workshop/OnlineProceedings12/pdf/ntcir/MathIR/04-NTCIR12-MathIR-KristiantoGY.pdf)Cited by:[§4\.2](https://arxiv.org/html/2605.19723#S4.SS2.p3.1)\.
- V\. Kukreja and Sakshi \(2022\)Machine learning models for mathematical symbol recognition: a stem to stern literature analysis\.Multimedia Tools Appl\.81\(20\),pp\. 28651–28687\.External Links:ISSN 1380\-7501,[Link](https://doi.org/10.1007/s11042-022-12644-2),[Document](https://dx.doi.org/10.1007/s11042-022-12644-2)Cited by:[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p2.1)\.
- G\. Lample and F\. Charton \(2020\)Deep learning for symbolic mathematics\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=S1eZYeHFDS)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I6.i3.p1.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p2.1)\.
- A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil,et al\.\(2022\)Solving quantitative reasoning problems with language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 3843–3857\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/18abbeef8cfe9203fdf9053c9c4fe191-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p2.1)\.
- C\. Li, X\. Fei, and X\. Yang \(2025\)Optimizing numerical reasoning with pre\-trained language models: an enhanced algorithm\.InProceedings of the 2025 6th International Conference on Computer Information and Big Data Applications,CIBDA ’25,New York, NY, USA,pp\. 285–295\.External Links:ISBN 9798400713163,[Link](https://doi.org/10.1145/3746709.3746759),[Document](https://dx.doi.org/10.1145/3746709.3746759)Cited by:[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p2.1)\.
- S\. Li \(2024\)Enhancing mathematical problem solving in large language models through tool\-integrated reasoning and python code execution\.In2024 5th International Conference on Big Data & Artificial Intelligence & Software Engineering \(ICBASE\),Vol\.,pp\. 165–168\.External Links:[Document](https://dx.doi.org/10.1109/ICBASE63199.2024.10762312)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I6.i2.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p1.1)\.
- Z\. Li, X\. Li, Y\. Liu, H\. Xie, J\. Li, F\. Wang, Q\. Li, and X\. Zhong \(2023\)Label supervised llama finetuning\.External Links:2310\.01208,[Link](https://arxiv.org/abs/2310.01208)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p7.1)\.
- C\. X\. Liang, Z\. Bi, T\. Wang, M\. Liu, X\. Song, Y\. Zhang, J\. Song, Q\. Niu, B\. Peng, K\. Chen,et al\.\(2025\)Low\-rank adaptation for scalable large language models: a comprehensive survey\.Authorea Preprints\.Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I5.i4.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p2.1)\.
- J\. Lindsey, W\. Gurnee, E\. Ameisen, B\. Chen, A\. Pearce, N\. L\. Turner, C\. Citro, D\. Abrahams, S\. Carter, B\. Hosmer,et al\.\(2025\)On the biology of a large language model \(2025\)\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2025/attribution-graphs/biology.html)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p4.1),[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p1.1),[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p2.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p3.1),[§7\.2\.2](https://arxiv.org/html/2605.19723#S7.SS2.SSS2.p1.1)\.
- A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.\(2025a\)Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.External Links:2412\.19437,[Link](https://arxiv.org/abs/2412.19437)Cited by:[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1)\.
- J\. Liu, Z\. Huang, Z\. Ma, Q\. Liu, E\. Chen, T\. Su, and H\. Liu \(2023\)Guiding mathematical reasoning via mastering commonsense formula knowledge\.InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,KDD ’23,New York, NY, USA,pp\. 1477–1488\.External Links:ISBN 9798400701030,[Link](https://doi.org/10.1145/3580305.3599375),[Document](https://dx.doi.org/10.1145/3580305.3599375)Cited by:[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p2.1)\.
- S\. Liu, C\. Wang, H\. Yin, P\. Molchanov, Y\. F\. Wang, K\. Cheng, and M\. Chen \(2024\)DoRA: weight\-decomposed low\-rank adaptation\.InInternational Conference on Machine Learning,External Links:[Link](https://api.semanticscholar.org/CorpusID:267657886)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I5.i3.p1.1),[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.5.4.1.1.1)\.
- W\. Liu, H\. Hu, J\. Zhou, Y\. Ding, J\. Li, J\. Zeng, M\. He, Q\. Chen, B\. Jiang, A\. Zhou, and L\. He \(2025b\)Mathematical language models: a survey\.ACM Comput\. Surv\.58\(6\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3773985),[Document](https://dx.doi.org/10.1145/3773985)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I10.i3.p1.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p2.1)\.
- S\. Longpre, K\. Perisetla, A\. Chen, N\. Ramesh, C\. DuBois, and S\. Singh \(2021\)Entity\-based knowledge conflicts in question answering\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,Online and Punta Cana, Dominican Republic,pp\. 7052–7063\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.565/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.565)Cited by:[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p2.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p2.1)\.
- P\. Lu, R\. Gong, S\. Jiang, L\. Qiu, S\. Huang, X\. Liang, and S\. Zhu \(2021\)Inter\-GPS: interpretable geometry problem solving with formal language and symbolic reasoning\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),Online,pp\. 6774–6786\.External Links:[Link](https://aclanthology.org/2021.acl-long.528/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.528)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I6.i1.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- B\. Mansouri, S\. Rohatgi, D\. W\. Oard, J\. Wu, C\. L\. Giles, and R\. Zanibbi \(2019\)Tangent\-cft: an embedding model for mathematical formulas\.InProceedings of the 2019 ACM SIGIR International Conference on the Theory of Information Retrieval \(ICTIR ’19\),Santa Clara, CA, USA,pp\. 11–18\.External Links:[Document](https://dx.doi.org/10.1145/3341981.3344235),ISBN 978\-1\-4503\-6881\-0,[Link](https://doi.org/10.1145/3341981.3344235)Cited by:[§4\.2](https://arxiv.org/html/2605.19723#S4.SS2.p3.1)\.
- N\. Matzakos, S\. Doukakis, and M\. Moundridou \(2023\)Learning mathematics with large language models: a comparative study with computer algebra systems and other tools\.International Journal of Emerging Technologies in Learning \(iJET\)18\(20\),pp\. 51–71\.External Links:[Document](https://dx.doi.org/10.3991/ijet.v18i20.42979),[Link](https://online-journals.org/index.php/i-jet/article/view/42979)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p1.1)\.
- S\. Miao, C\. Liang, and K\. Su \(2020\)A diverse corpus for evaluating and developing English math word problem solvers\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,Online,pp\. 975–984\.External Links:[Link](https://aclanthology.org/2020.acl-main.92/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.92)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- Y\. Miao, S\. Zhang, L\. Ding, R\. Bao, L\. Zhang, and D\. Tao \(2024\)InfoRM: mitigating reward hacking in rlhf via information\-theoretic reward modeling\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 134387–134429\.External Links:[Document](https://dx.doi.org/10.52202/079017-4270),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/f25d75fc760aec0a6174f9f5d9da59b8-Paper-Conference.pdf)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I4.i3.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1)\.
- S\. Min, M\. Lewis, L\. Zettlemoyer, and H\. Hajishirzi \(2022\)MetaICL: learning to learn in context\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,Seattle, United States,pp\. 2791–2809\.External Links:[Link](https://aclanthology.org/2022.naacl-main.201/),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.201)Cited by:[§5\.3\.2](https://arxiv.org/html/2605.19723#S5.SS3.SSS2.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p3.1)\.
- S\. Mishra, M\. Finlayson, P\. Lu, L\. Tang, S\. Welleck, C\. Baral, T\. Rajpurohit, O\. Tafjord, A\. Sabharwal, P\. Clark, and A\. Kalyan \(2022\)LILA: a unified benchmark for mathematical reasoning\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Abu Dhabi, United Arab Emirates,pp\. 5807–5832\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.392/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.392)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I2.i3.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p4.1),[§3\.4](https://arxiv.org/html/2605.19723#S3.SS4.p1.1)\.
- T\. Mu, A\. Helyar, J\. Heidecke, J\. Achiam, A\. Vallone, I\. Kivlichan, M\. Lin, A\. Beutel, J\. Schulman, and L\. Weng \(2024\)Rule based rewards for language model safety\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 108877–108901\.External Links:[Document](https://dx.doi.org/10.52202/079017-3457),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/c4e380fb74dec9da9c7212e834657aa9-Paper-Conference.pdf)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I4.i4.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1),[§7\.2\.2](https://arxiv.org/html/2605.19723#S7.SS2.SSS2.p1.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, Slama,et al\.\(2022\)Training language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 27730–27744\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I4.i1.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p3.1)\.
- A\. Parnami and M\. Lee \(2022\)Learning from few examples: a summary of approaches to few\-shot learning\.External Links:2203\.04291,[Link](https://arxiv.org/abs/2203.04291)Cited by:[§5\.3\.2](https://arxiv.org/html/2605.19723#S5.SS3.SSS2.p2.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p3.1)\.
- K\. Paster, M\. D\. Santos, Z\. Azerbayev, and J\. Ba \(2023\)OpenWebMath: an open dataset of high\-quality mathematical web text\.External Links:2310\.06786,[Link](https://arxiv.org/abs/2310.06786)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p2.1),[§4](https://arxiv.org/html/2605.19723#S4.p5.1)\.
- S\. Peng, K\. Yuan, L\. Gao, and Z\. Tang \(2021\)MathBERT: a pre\-trained model for mathematical formula understanding\.External Links:2105\.00377,[Link](https://arxiv.org/abs/2105.00377)Cited by:[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p2.1),[§7\.2\.4](https://arxiv.org/html/2605.19723#S7.SS2.SSS4.p1.1)\.
- F\. Pourpanah, M\. Abdar, Y\. Luo, X\. Zhou, R\. Wang, C\. P\. Lim, X\. Wang, and Q\. M\. J\. Wu \(2023\)A review of generalized zero\-shot learning methods\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(4\),pp\. 4051–4070\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2022.3191696)Cited by:[§5\.3\.2](https://arxiv.org/html/2605.19723#S5.SS3.SSS2.p1.1)\.
- J\. W\. Rae, S\. Borgeaud, T\. Cai, K\. Millican, J\. Hoffmann, F\. Song, J\. Aslanides, S\. Henderson, R\. Ring,et al\.\(2022\)Scaling language models: methods, analysis & insights from training gopher\.External Links:2112\.11446,[Link](https://arxiv.org/abs/2112.11446)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I1.i4.p1.1),[§1](https://arxiv.org/html/2605.19723#S1.p2.1)\.
- N\. Rajaraman, J\. Jiao, and K\. Ramchandran \(2025\)Toward a theory of tokenization in llms\.External Links:2404\.08335,[Link](https://arxiv.org/abs/2404.08335)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I8.i2.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p5.1),[§7\.2\.4](https://arxiv.org/html/2605.19723#S7.SS2.SSS4.p1.1)\.
- R\. Ren \(2025\)The multi\-agent fault localization system based on monte carlo tree search approach\.External Links:2507\.22800,[Link](https://arxiv.org/abs/2507.22800)Cited by:[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p2.1),[§6](https://arxiv.org/html/2605.19723#S6.p7.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p4.1),[§7\.2\.3](https://arxiv.org/html/2605.19723#S7.SS2.SSS3.p1.1)\.
- L\. Rhomrasi, Y\. Ahsini, A\. Igualde, R\. Vinuesa, S\. Hoyas, J\. García\-Sabater, M\. Fullana i Alfonso, and A\. Conejero \(2025\)LLM performance on mathematical reasoning in catalan language\.Results in Engineering25,pp\. 104366\.External Links:[Document](https://dx.doi.org/10.1016/j.rineng.2025.104366)Cited by:[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p4.1)\.
- S\. Roy and D\. Roth \(2015\)Solving general arithmetic word problems\.InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing,L\. Màrquez, C\. Callison\-Burch, and J\. Su \(Eds\.\),Lisbon, Portugal,pp\. 1743–1752\.External Links:[Link](https://aclanthology.org/D15-1202/),[Document](https://dx.doi.org/10.18653/v1/D15-1202)Cited by:[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p2.1)\.
- S\. Roy and D\. Roth \(2017\)Unit dependency graph and its application to arithmetic word problem solving\.InProceedings of the Thirty\-First AAAI Conference on Artificial Intelligence,AAAI’17,pp\. 3082–3088\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v31i1.10959),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/10959/10818)Cited by:[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p2.1)\.
- J\. Schulman, S\. Levine, P\. Abbeel, M\. Jordan, and P\. Moritz \(2015\)Trust region policy optimization\.InProceedings of the 32nd International Conference on Machine Learning,F\. Bach and D\. Blei \(Eds\.\),Proceedings of Machine Learning Research, Vol\.37,Lille, France,pp\. 1889–1897\.External Links:[Link](https://proceedings.mlr.press/v37/schulman15.html)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I4.i2.p1.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p4.1)\.
- M\. Shalyt, R\. Elimelech, and I\. Kaminer \(2025\)Do LLMs understand calculus? evaluating symbolic math generalization with ASyMOB\.InAI4X 2025 International Conference,External Links:[Link](https://openreview.net/forum?id=kdatAr0mcd)Cited by:[§4\.1](https://arxiv.org/html/2605.19723#S4.SS1.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p5.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. Guo \(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.External Links:2402\.03300,[Link](https://arxiv.org/abs/2402.03300)Cited by:[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p4.1)\.
- W\. Shi, A\. Ajith, M\. Xia, Y\. Huang, D\. Liu, T\. Blevins, D\. Chen, and L\. Zettlemoyer \(2024\)Detecting pretraining data from large language models\.External Links:2310\.16789,[Link](https://arxiv.org/abs/2310.16789)Cited by:[4th item](https://arxiv.org/html/2605.19723#A1.I7.i4.p1.1),[§5\.4](https://arxiv.org/html/2605.19723#S5.SS4.p3.1)\.
- A\. K\. Singh and D\. Strouse \(2024\)Tokenization counts: the impact of tokenization on arithmetic in frontier llms\.External Links:2402\.14903,[Link](https://arxiv.org/abs/2402.14903)Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I8.i2.p1.1),[§4\.3](https://arxiv.org/html/2605.19723#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p6.1),[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p4.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p5.1)\.
- S\. Smith, M\. Patwary, B\. Norick, P\. LeGresley, S\. Rajbhandari, J\. Casper, Z\. Liu, S\. Prabhumoye,et al\.\(2022\)Using deepspeed and megatron to train megatron\-turing nlg 530b, a large\-scale generative language model\.External Links:2201\.11990,[Link](https://arxiv.org/abs/2201.11990)Cited by:[§5](https://arxiv.org/html/2605.19723#S5.p2.1)\.
- Y\. Stathopoulos \(2022\)Retrieval of research\-level mathematics via joint modelling of text and types\.Ph\.D\. Thesis,Apollo \- University of Cambridge Repository\.External Links:[Link](https://www.repository.cam.ac.uk/handle/1810/347577),[Document](https://dx.doi.org/10.17863/CAM.94992)Cited by:[§4\.2](https://arxiv.org/html/2605.19723#S4.SS2.p1.1)\.
- G\. B\. Stone, D\. A\. Talbert, and W\. Eberle \(2022\)A survey of scalable reinforcement learning\.International Journal of Intelligent Computing Research \(IJICR\)13\(1\)\.Note:Copyright © 2022, Infonomics SocietyExternal Links:[Document](https://dx.doi.org/10.20533/ijicr.2042.4655.2022.0136),[Link](https://infonomics-society.org/wp-content/uploads/A-Survey-of-Scalable-Reinforcement-Learning.pdf)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p7.1)\.
- Z\. Tang, X\. Zhang, B\. Wang, and F\. Wei \(2024\)MathScale: scaling instruction tuning for mathematical reasoning\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 47885–47900\.External Links:[Link](https://proceedings.mlr.press/v235/tang24k.html)Cited by:[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p3.1),[§7\.2\.1](https://arxiv.org/html/2605.19723#S7.SS2.SSS1.p1.1)\.
- R\. K\. Thatikonda, W\. Buntine, and E\. Shareghi \(2025\)Logical reasoning with outcome reward models for test\-time scaling\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 26102–26112\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1326/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1326),ISBN 979\-8\-89176\-332\-6Cited by:[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p3.1)\.
- Y\. Tian, B\. Peng, L\. Song, L\. Jin, D\. Yu, L\. Han, H\. Mi, and D\. Yu \(2024\)Toward self\-improvement of llms via imagination, searching, and criticizing\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 52723–52748\.External Links:[Document](https://dx.doi.org/10.52202/079017-1670),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/5e5853f35164e434015716a8c2a66543-Paper-Conference.pdf)Cited by:[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p3.1)\.
- S\. Toshniwal, I\. Moshkov, S\. Narenthiran, D\. Gitman, F\. Jia, and I\. Gitman \(2024\)OpenMathInstruct\-1: a 1\.8 million math instruction tuning dataset\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§4](https://arxiv.org/html/2605.19723#S4.p5.1)\.
- M\. Valipour, M\. Rezagholizadeh, I\. Kobyzev, and A\. Ghodsi \(2023\)DyLoRA: parameter\-efficient tuning of pre\-trained models using dynamic search\-free low\-rank adaptation\.InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics,Dubrovnik, Croatia,pp\. 3274–3287\.External Links:[Link](https://aclanthology.org/2023.eacl-main.239/),[Document](https://dx.doi.org/10.18653/v1/2023.eacl-main.239)Cited by:[3rd item](https://arxiv.org/html/2605.19723#A1.I5.i3.p1.1),[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.6.5.1.1.1)\.
- X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. Zhou \(2023\)Self\-consistency improves chain of thought reasoning in language models\.External Links:2203\.11171,[Link](https://arxiv.org/abs/2203.11171)Cited by:[1st item](https://arxiv.org/html/2605.19723#A1.I3.i1.p1.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p1.1)\.
- Y\. Wang, Q\. Yao, J\. T\. Kwok, and L\. M\. Ni \(2020\)Generalizing from a few examples: a survey on few\-shot learning\.ACM Comput\. Surv\.53\(3\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3386252),[Document](https://dx.doi.org/10.1145/3386252)Cited by:[§5\.3\.2](https://arxiv.org/html/2605.19723#S5.SS3.SSS2.p1.1)\.
- Y\. Wang, N\. Yang, L\. Wang, F\. Wei, and F\. Feng \(2025\)Examining false positives under inference scaling for mathematical reasoning\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 12501–12520\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.632/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.632),ISBN 979\-8\-89176\-332\-6Cited by:[2nd item](https://arxiv.org/html/2605.19723#A1.I7.i2.p1.1),[§6](https://arxiv.org/html/2605.19723#S6.p1.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p4.1)\.
- Y\. Wang and Y\. Zhao \(2023\)Metacognitive prompting improves understanding in large language models\.InNorth American Chapter of the Association for Computational Linguistics,External Links:[Link](https://api.semanticscholar.org/CorpusID:260775822)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p4.1),[§3\.4](https://arxiv.org/html/2605.19723#S3.SS4.p2.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, b\. ichter, F\. Xia, E\. Chi, Q\. V\. Le, and D\. Zhou \(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 24824–24837\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p2.1),[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p1.1),[§3\.3](https://arxiv.org/html/2605.19723#S3.SS3.p1.1)\.
- T\. Wei, L\. Zhao, L\. Zhang, B\. Zhu, L\. Wang, H\. Yang, B\. Li, C\. Cheng, W\. Lü, R\. Hu, C\. Li, L\. Yang, X\. Luo, X\. Wu, L\. Liu, W\. Cheng, P\. Cheng, J\. Zhang, X\. Zhang, L\. Lin, X\. Wang, Y\. Ma, C\. Dong, Y\. Sun, Y\. Chen, Y\. Peng, X\. Liang, S\. Yan, H\. Fang, and Y\. Zhou \(2023\)Skywork: a more open bilingual foundation model\.External Links:2310\.19341,[Link](https://arxiv.org/abs/2310.19341)Cited by:[§5\.2](https://arxiv.org/html/2605.19723#S5.SS2.p2.1)\.
- X\. Wu, M\. Chen, W\. Li, R\. Wang, L\. Lu, J\. Liu, K\. Hwang, Y\. Hao, Y\. Pan, Q\. Meng, K\. Huang, L\. Hu, M\. Guizani, N\. Chao, G\. Fortino, F\. Lin, Y\. Tian, D\. Niyato, and F\. Wang \(2025\)LLM fine\-tuning: concepts, opportunities, and challenges\.Big Data and Cognitive Computing9\(4\)\.External Links:[Link](https://www.mdpi.com/2504-2289/9/4/87),ISSN 2504\-2289,[Document](https://dx.doi.org/10.3390/bdcc9040087)Cited by:[§5\.3\.1](https://arxiv.org/html/2605.19723#S5.SS3.SSS1.p1.1)\.
- Y\. Xian, B\. Schiele, and Z\. Akata \(2017\)Zero\-Shot Learning — The Good, the Bad and the Ugly\.In2017 IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),Vol\.,Los Alamitos, CA, USA,pp\. 3077–3086\.External Links:ISSN 1063\-6919,[Document](https://dx.doi.org/10.1109/CVPR.2017.328),[Link](https://doi.ieeecomputersociety.org/10.1109/CVPR.2017.328)Cited by:[§5\.3\.2](https://arxiv.org/html/2605.19723#S5.SS3.SSS2.p1.1)\.
- H\. Xin, D\. Guo, Z\. Shao, Z\.Z\. Ren, Q\. Zhu, B\. Liu, C\. Ruan, W\. Li, and X\. Liang \(2024\)Advancing theorem proving in LLMs through large\-scale synthetic data\.InThe 4th Workshop on Mathematical Reasoning and AI at NeurIPS’24,External Links:[Link](https://openreview.net/forum?id=TPtXLihkny)Cited by:[§3\.2](https://arxiv.org/html/2605.19723#S3.SS2.p1.1)\.
- K\. G\. Yoko, H\. Florence, A\. Akiko,et al\.\(2014\)The mcat math retrieval system for ntcir\-11 math track\.\.InProceedings of the 11th NTCIR Conference,External Links:[Link](https://research.nii.ac.jp/ntcir/workshop/OnlineProceedings11/pdf/NTCIR/Math-2/06-NTCIR11-MATH-KristiantoGY.pdf)Cited by:[§4\.2](https://arxiv.org/html/2605.19723#S4.SS2.p3.1)\.
- Z\. Yuan, H\. Yuan, C\. Tan, W\. Wang, and S\. Huang \(2023\)How well do large language models perform in arithmetic tasks?\.External Links:2304\.02015,[Link](https://arxiv.org/abs/2304.02015)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p1.1),[§3\.1](https://arxiv.org/html/2605.19723#S3.SS1.p1.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1),[§6\.1](https://arxiv.org/html/2605.19723#S6.SS1.p3.1),[§6](https://arxiv.org/html/2605.19723#S6.p6.1),[§7\.1](https://arxiv.org/html/2605.19723#S7.SS1.p4.1),[§7\.2\.3](https://arxiv.org/html/2605.19723#S7.SS2.SSS3.p1.1)\.
- R\. Zanibbi, D\. Blostein, R\. Zanibbi, and D\. Blostein \(2011\)Recognition and retrieval of mathematical expressions\.International Journal on Document Analysis and Recognition \(IJDAR\)15,pp\. 331–357\.External Links:[Document](https://dx.doi.org/10.1007/s10032-011-0174-4)Cited by:[§4\.2](https://arxiv.org/html/2605.19723#S4.SS2.p2.1),[§7\.2\.4](https://arxiv.org/html/2605.19723#S7.SS2.SSS4.p1.1)\.
- L\. Zhang, L\. Li, W\. Wei, H\. Song, Y\. Yang, and J\. Liang \(2024a\)Scalable constrained policy optimization for safe multi\-agent reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 138698–138730\.External Links:[Document](https://dx.doi.org/10.52202/079017-4400),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/fa76985f05e0a25c66528308dda33de0-Paper-Conference.pdf)Cited by:[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p2.1),[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p4.1)\.
- M\. Zhang, Z\. Li, F\. Yin, L\. Lin, and C\. Liu \(2024b\)Fuse, reason and verify: geometry problem solving with parsed clauses from diagram\.External Links:2407\.07327,[Link](https://arxiv.org/abs/2407.07327)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4](https://arxiv.org/html/2605.19723#S4.p5.1)\.
- Q\. Zhang, M\. Chen, A\. Bukharin, P\. He, Y\. Cheng, W\. Chen, and T\. Zhao \(2023\)Adaptive budget allocation for parameter\-efficient fine\-tuning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=lq62uWRJjiY)Cited by:[Table 10](https://arxiv.org/html/2605.19723#S5.T10.1.1.4.3.1.1.1)\.
- Y\. Zhang, J\. Yang, Y\. Yuan, and A\. C\. Yao \(2025a\)Cumulative reasoning with large language models\.Transactions on Machine Learning Research\.External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=grW15p4eq2)Cited by:[§5\.2](https://arxiv.org/html/2605.19723#S5.SS2.p2.1)\.
- Z\. Zhang, C\. Zheng, Y\. Wu, B\. Zhang, R\. Lin, B\. Yu, D\. Liu, J\. Zhou, and J\. Lin \(2025b\)The lessons of developing process reward models in mathematical reasoning\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 10495–10516\.External Links:[Link](https://aclanthology.org/2025.findings-acl.547/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.547),ISBN 979\-8\-89176\-256\-5Cited by:[§5\.3\.3](https://arxiv.org/html/2605.19723#S5.SS3.SSS3.p3.1)\.
- K\. Zheng, J\. M\. Han, and S\. Polu \(2022\)MiniF2F: a cross\-system benchmark for formal olympiad\-level mathematics\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=9ZPegFuFTFv)Cited by:[§1](https://arxiv.org/html/2605.19723#S1.p7.1),[§4\.4](https://arxiv.org/html/2605.19723#S4.SS4.p5.1)\.
- D\. M\. Ziegler, N\. Stiennon, J\. Wu, T\. B\. Brown, A\. Radford, D\. Amodei, P\. Christiano, and G\. Irving \(2020\)Fine\-tuning language models from human preferences\.External Links:1909\.08593,[Link](https://arxiv.org/abs/1909.08593)Cited by:[§5\.3\.1](https://arxiv.org/html/2605.19723#S5.SS3.SSS1.p1.1)\.
- V\. Zouhar, C\. Meister, J\. Gastaldi, L\. Du, T\. Vieira, M\. Sachan, and R\. Cotterell \(2023\)A formal perspective on byte\-pair encoding\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 598–614\.External Links:[Link](https://aclanthology.org/2023.findings-acl.38/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.38)Cited by:[§4\.3](https://arxiv.org/html/2605.19723#S4.SS3.p2.1)\.Similar Articles
Can We Understand How Large Language Models Reason?
This article explores the ongoing efforts and challenges in understanding how large language models reason, focusing on interpretability research.
Enhanced and Efficient Reasoning in Large Learning Models
This paper proposes a method for improving reasoning in large language models by recoding data to explicitly represent relationships, enabling efficient principled reasoning with polynomial-time learnability for relational rules, which addresses hallucinations and supports sound reasoning across multiple calls.
Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
This paper investigates representation robustness in LLMs for mathematical problem solving by systematically varying surface representations of equivalent problems, finding substantial sensitivity and showing that code-augmented reasoning does not uniformly eliminate brittleness.
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
This paper introduces MPAR-Bench, a bilingual benchmark for evaluating multi-point associative reasoning in large language models, along with a perturbation suite and coarse-to-fine evaluation protocol. Results show that deeper reasoning does not automatically confer robust reasoning breadth.
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
This paper investigates how large language models perform arithmetic operations by analyzing internal mechanisms through early decoding, revealing that proficient models exhibit a clear division of labor between attention and MLP modules in reasoning tasks.