MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

arXiv cs.CL Papers

Summary

The paper introduces MIRAGE, an inference-time framework for enhancing LLM reasoning by dynamically switching perspectives using reinforcement learning guidance, outperforming existing prompting methods on various benchmarks.

arXiv:2609.21554v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility - our ability to dynamically switch mental perspectives - we propose MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration), a novel inference-time creative thinking framework. MIRAGE includes a Selector that prioritizes effective conceptual perspectives (e.g., algebraic, probabilistic) and a Reasoner that sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives. Tested on GSM8K, MATH500, MMLU-Pro, and Game-of-24 benchmarks, MIRAGE consistently outperforms methods like Chain-of-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications.
Original Article
View Cached Full Text

Cached at: 09/21/26, 09:08 AM

# MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance
Source: [https://arxiv.org/html/2609.21554](https://arxiv.org/html/2609.21554)
Arash LagzianSrinivas AnumasaAffiliation:National University of SingaporeCorrespondence to:[srinu\_pd@nus\.edu\.sg](mailto:[email protected])Dianbo LiuAffiliation:National University of SingaporeCorrespondence to:[dianbo@nus\.edu\.sg](mailto:[email protected])

###### Abstract

Recent advances in Large Language Models \(LLMs\) have revolutionized artificial intelligence and how human interact with AIs\. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks\. Inspired by human cognitive flexibility—our ability to dynamically switch mental perspectives—we proposeMIRAGE\(Multi\-perspectiveInference\-timeReasoning viaAgent\-GuidedExploration\), a novel inference\-time creative thinking framework\. MIRAGE includes aSelectorthat prioritizes effective conceptual perspectives \(e\.g\., algebraic, probabilistic\) and aReasonerthat sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives\. Tested on GSM8K, MATH500, MMLU\-Pro, and Game\-of\-24 benchmarks, MIRAGE consistently outperforms methods like Chain\-of\-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications\.

###### Keywords:

Machine Learning, ICML

## 1Introduction

Recent progress in Large Language Models \(LLMs\) has greatly improved their ability to handle open\-ended language tasks, but they still struggle with complex reasoning in math, science, and logic\. Even small changes in wording or how the model generates its response can cause it to go from a correct answer to a completely wrong one\. This kind of fragility makes it hard to trust LLMs in important applications like tutoring systems, engineering assistants, or financial tools\([Cobbe et al\., 2021b](https://arxiv.org/html/2609.21554#bib.bib9);[Wei et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib1)\)\. To close this gap, researchers have explored ever richer prompting and inference\-time techniques: Chain\-of\-Thought \(CoT\) prompting steers models through explicit reasoning steps\([Wei et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib1)\), zero\-shot CoT and scratchpads remove the need for demonstrations\([Kojima et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib3);[Nye et al\., 2021](https://arxiv.org/html/2609.21554#bib.bib28)\), and Least\-to\-Most prompting decomposes complex problems into simpler sub\-questions\([Zhou et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib7)\)\. Decoding strategies such as Self\-Consistency\([Wang et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib8)\), ensemble methods like Reflexion and DIPPER sample multiple reasoning paths and vote on an answer\([Shinn et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib10);[Lau et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib2)\), while tool\-augmented frameworks call external calculators or verifiers to patch errors\([Schick et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib13)\)\.

Why do existing fixes fall short?*Fragile prompting*—CoT variants depend on carefully crafted exemplars; minor edits can derail generation\([Kojima et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib3)\)\.*Costly ensembling*—sampling three to five reasoning paths per query boosts accuracy but inflates latency and API cost by up to5×5\\times\([Wang et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib8);[Lau et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib2)\)\.*Rigid single\-perspective reasoning*—all methods process the problem through one fixed representation; if that perspective misaligned with the task’s structure, there is no fallback\. Tool\-augmented or fine\-tuned systems further add infrastructure overhead and sacrifice model\-agnostic portability\.

Decades of research in cognitive science and neuroscience reveal that humans seldom resolve such challenges using a single representational approach\. Instead, we fluidly*re\-encode*problems—transforming algebraic tasks into geometric ones, or probability tasks into frequency\-based representations—until one representation clearly facilitates insight\. This principle of cognitive flexibility is supported by studies on strategy switching\([Siegler, 1996](https://arxiv.org/html/2609.21554#bib.bib18)\), representational shifts underlying insights\([Knoblich et al\., 1999](https://arxiv.org/html/2609.21554#bib.bib23)\), and evidence for specialized parallel neural circuits\([Dehaene, 2009](https://arxiv.org/html/2609.21554#bib.bib22);[Deen and Freiwald, 2021](https://arxiv.org/html/2609.21554#bib.bib21)\)\.

Addressing the reasoning challenge using inspiration from cognitive science, we introduceMIRAGE, an creative multi\-perspective inference\-time reasoning framework inspired by the human cognitive strategy of problem re\-representation to discover insightful solutions\. For mathematic reasoning tasks, MIRAGE employs a*Selector*that ranks twenty conceptual reasoning perspectives—including algebraic, probabilistic, and game\-theoretic frameworks—based on their historical effectiveness on similar problems\. Subsequently, a*Reasoner*iteratively attempts these perspectives in ranked order, halting either upon reaching a high\-confidence solution or after forming a small ensemble\. Remarkably, MIRAGE requires no parameter updates to the base LLM and averages fewer than two forward passes on three of four benchmarks, achieving superior accuracy–cost trade\-offs \(Figure[4](https://arxiv.org/html/2609.21554#S5.F4)\)\.

Our contributions in this study are as follows:

- •We introduce the first inference\-time framework that*learns*to choose among conceptual reasoning perspectives \(Selector\) and solves within them \(Reasoner\) without modifying LLM weights\.
- •Across GSM8K, MATH500, MMLU\-Pro, and Game\-of\-24, MIRAGE lifts accuracy by up to\+24\.7 ppover CoT while using at most2×2\\timesthe cost of a single call and up to5×5\\timesless cost than DIPPER \(n=5n\{=\}5\)\.
- •Extensive experiments on five base models—DeepSeek\-v3, ChatGPT\-4o, Claude 3\.7\-Sonnet, Gemini\-Flash 2\.0, and Qwen2\.5\-7B—confirm consistent gains and detailed accuracy\-vs\-cost analysis

![Refer to caption](https://arxiv.org/html/2609.21554v1/images/paper_figures3.png)Figure 1:Overview of our multi\-perspective reasoning framework\.\(a\)Standard reasoning uses a single forward pass through the Reasoner model\.\(b\)Our method proceeds in four stages: \(1\) the Selector ranks conceptual perspectives, \(2\) the Reasoner model solves each perspective, \(3\) if confidence exceeds the threshold, the answer is returned, and \(4\) otherwise, answers are aggregated\. These stages are illustrated with numbered circles\.
## 2Related Works

Reasoning with large language models has advanced along three intertwined threads that culminate in the ideas we pursue inMIRAGE\. First, prompt–engineering techniques expose latent chain\-of\-thought abilities\. Chain\-of\-Thought \(CoT\) prompting\([Wei et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib1)\)and its zero\-shot variant “Let’s think step by step”\([Kojima et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib3)\)showed that supplying intermediate steps can dramatically lift arithmetic and logical accuracy; later extensions introduced*scratchpads*to reveal hidden computations\([Nye et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib6)\), Least\-to\-Most decomposition for hierarchical problem solving\([Zhou et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib7)\), and self\-training with generated rationales in STaR\([Zelikman et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib32)\)\. Second, search\-based and self\-verification methods improve reliability by exploring multiple reasoning paths: Self\-Consistency aggregates sampled solutions\([Wang et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib8)\), verifier models prune incorrect chains\([Cobbe et al\., 2021b](https://arxiv.org/html/2609.21554#bib.bib9)\), and reflection loops iteratively repair errors\([Shinn et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib10)\)\. Third, modular agent frameworks route questions to external tools or expert policies\. Tree\-of\-Thoughts performs deliberative tree search\([Yao et al\., 2023a](https://arxiv.org/html/2609.21554#bib.bib11)\), ReAct interleaves reasoning with tool calls\([Yao et al\., 2023b](https://arxiv.org/html/2609.21554#bib.bib33)\), PAL executes generated code to obtain ground\-truth signals\([Gao et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib12)\), Toolformer learns when to query APIs\([Schick et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib13)\), HuggingGPT orchestrates specialist models\([Shen et al\., 2023](https://arxiv.org/html/2609.21554#bib.bib14)\), and HDFlow adaptively chooses between fast and slow solvers\([Yao et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib15)\)\. Very recent work pushes routing one step further: DIPPER emphasis on the importance of diversity in input prompts and ensemble models\([Lau et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib2)\), and Atomic Reasoner try to extract atomic facts and reason based on them\([Zhang et al\., 2025](https://arxiv.org/html/2609.21554#bib.bib31)\)\.

MIRAGEbuilds on these insights but occupies a distinct niche\. Instead of sampling many chains or invoking external APIs, we maintain a*human\-interpretable library of twenty conceptual perspectives*\(algebraic, probabilistic, network\-flow,*etc\.*\) and train a lightweight selector—once, on solved examples—to choose the most promising perspective at inference time\. This yields a single\-call, multi\-perspective solution that matches or surpasses the accuracy–cost Pareto front of tree search, verifier, and modular\-tool baselines while requiring no additional model fine\-tuning\.

## 3Motivation from Brain and Cognitive Science

Our approach is fundamentally motivated by well\-established cognitive and neuroscientific principles highlighting the importance of cognitive flexibility and multi\-perspective reasoning in effective human problem\-solving\. Cognitive science research consistently demonstrates that humans excel in complex problem\-solving scenarios by dynamically shifting among multiple mental frameworks or representational strategies, adapting flexibly as task demands evolve\([Spiro et al\., 1988](https://arxiv.org/html/2609.21554#bib.bib16);[Miyake et al\., 2000](https://arxiv.org/html/2609.21554#bib.bib17);[Siegler, 1996](https://arxiv.org/html/2609.21554#bib.bib18)\)\. Translating these insights into AI, we posit that enhancing the reasoning capabilities of Large Language Models \(LLMs\) similarly requires an inference strategy capable of fluidly switching among diverse conceptual representations or reasoning perspectives\.

Cognitive Flexibility Theory, as articulated by Spiro et al\.\([Spiro et al\., 1988](https://arxiv.org/html/2609.21554#bib.bib16)\), underscores the necessity of constructing knowledge through multiple, overlapping representations to achieve mastery in complex and ill\-structured domains\. This theory asserts that learners who regularly restructure knowledge across different conceptual perspectives not only deepen their understanding but also enhance their ability to transfer insights effectively to novel contexts\. Parallel evidence from developmental psychology, notably theOverlapping Waves Theoryintroduced by Siegler\([Siegler, 1996](https://arxiv.org/html/2609.21554#bib.bib18)\), further supports this principle\. Siegler demonstrated that human learners naturally employ and switch among multiple strategies to solve problems, progressively refining strategy selection through experience\. Such flexible use of diverse approaches facilitates robust and generalized problem\-solving capabilities\.

Empirical and neuroscientific evidence converges on the same lesson: switching representations boosts performance\. In classrooms, students who compare multiple algebraic methods achieve deeper procedural and conceptual mastery than peers taught a single approach\([Rittle\-Johnson and Star, 2007](https://arxiv.org/html/2609.21554#bib.bib19)\); likewise, Bayesian problems become far easier when reframed from probabilities to natural frequencies\([Gigerenzer and Hoffrage, 1995](https://arxiv.org/html/2609.21554#bib.bib20)\)\. fMRI studies echo this flexibility, revealing parallel circuits dedicated to social versus spatial reasoning\([Deen and Freiwald, 2021](https://arxiv.org/html/2609.21554#bib.bib21)\)and a triple\-code network for numerical quantity, visual, and verbal processing\([Dehaene, 2009](https://arxiv.org/html/2609.21554#bib.bib22)\), underscoring the brain’s propensity to recruit whichever representation best fits the task\.

Research on insight and social cognition paints a similar picture\. Breakthroughs in “aha\!” problems often hinge on abandoning an unproductive framing and relaxing prior constraints\([Knoblich et al\., 1999](https://arxiv.org/html/2609.21554#bib.bib23)\), while exposure to multiple cultural contexts broadens mental representations and boosts creative problem\-solving\([Maddux and Galinsky, 2009](https://arxiv.org/html/2609.21554#bib.bib24)\)\. Taken together, these strands suggest that intelligent systems should likewise pivot between representations to overcome impasses and generalize\.MIRAGEoperationalises this principle by coupling a Selector that chooses among twenty conceptual perspectives with a Reasoner that solves the problem inside each chosen view, aiming to confer human\-like flexibility on large language models\.

## 4Proposed Method

In this section, we present our inference\-time framework designed to improve the reasoning abilities of LLMs\. The core insight is that complex reasoning problems can be solved more effectively when dynamically transformed into multiple different conceptual perspectives specialized for different reasoning paradigms\.

### 4\.1Overall Framework

Given a reasoning task with input promptqq, our goal is to select an appropriate sequence of conceptual reasoning perspectives and iteratively solve within them until a confident solution is found\. Let𝒮=\{s1,s2,…,s20\}\\mathcal\{S\}=\\\{s\_\{1\},s\_\{2\},\\dots,s\_\{20\}\\\}denote the set of predefined conceptual perspectives\. These 20 perspectives were carefully curated by domain experts through an in\-depth analysis of problem\-solving strategies commonly observed in mathematics, science, engineering, and logic\. The selection process drew from techniques emphasized in educational curricula and employed by experts across disciplines\. We prioritized perspectives that are both cognitively distinctive and broadly applicable\. A complete description of all 20 perspectives, including detailed justifications for their inclusion and representative problem examples, is provided in Appendix[A](https://arxiv.org/html/2609.21554#A1)\.

Our approach consists of two phases: \(1\) a*Training Phase*, where the*Selector*is trained to choose relevant reasoning perspectives for each input problem while the Reasoner remains fixed, and \(2\) an*Inference Phase*, where the trained Selector dynamically guides the Reasoner through selected perspectives\.

### 4\.2Training Phase: Selector Reinforcement Learning

##### Setup\.

We train the selector with REINFORCE on5700MMLU\([Hendrycks et al\., 2021a](https://arxiv.org/html/2609.21554#bib.bib4)\)questions \(one hundred samples per subject, shuffled\)\. The selector is aQwen2\.5\-7B\-Instruct\([Team, 2024c](https://arxiv.org/html/2609.21554#bib.bib29)\)classifier fine\-tuned with LoRA \(r=16r\{=\}16,α=32\\alpha\{=\}32\)\. A frozenQwen2\.5\-14B\-Instruct\([Team, 2024b](https://arxiv.org/html/2609.21554#bib.bib30)\)reasoner, queried once per selected perspective, generates the step\-by\-step solution\. Training uses mini\-batches of eight questions and Adam \(η=×10−6\\eta=5\\\!\\times\\\!10^\{\-6\}\) on a pair of A100 GPUs\.

##### Learning dynamics\.

We train the selector using the REINFORCE algorithm with a fixed Reasoner \(Algorithm[2](https://arxiv.org/html/2609.21554#alg2)\)\. At each step, the selector outputs probabilities over the 20 conceptual perspectives, and a binary mask is sampled to decide which perspectives are activated\. The Reasoner attempts the problem under each selected perspective, and a majority vote produces the final answer\. The reward signal encourages both correctness and sparsity:

r=𝕀\[a^=a∗\]−λ⋅k\|𝒮\|r=\\mathbb\{I\}\[\\hat\{a\}=a^\{\*\}\]\-\\lambda\\cdot\\frac\{k\}\{\|\\mathcal\{S\}\|\}\(1\)
wherekkis the number of selected perspectives andλ=0\.05\\lambda=0\.05is the penalty coefficient\.

![Refer to caption](https://arxiv.org/html/2609.21554v1/images/training_metrics_avg_k_with_trend.png)Figure 2:Training curve of the average number of selected perspectiveskkduring REINFORCE training\. The dashed blue line shows the 100‐step EMA ofkk; the solid orange line is the best‐fit linear trend \(slope≈−2\.02×10−4\\approx\-2\.02\\times 10^\{\-4\}\), highlighting a slight but consistent downward drift\. Under penaltyλ=0\.05\\lambda=0\.05, the selector stabilizes at around eight perspectives per query\.Fig\.[2](https://arxiv.org/html/2609.21554#S4.F2)shows how the average number of perspectiveskkselected evolves over training\. The blue dashed line tracks the 100\-step exponential moving average \(EMA\) ofkk, while the orange line shows the best\-fit linear trend\. Despite fluctuations due to stochastic sampling, the selector exhibits a consistent downward drift inkk, stabilizing at around 8 perspectives per query\. This emergent sparsity demonstrates the selector’s ability to learn compact yet effective subspaces of reasoning, achieving a 2\.5× reduction in reasoning cost compared to querying all 20 perspectives\.

##### Baselines\.

Table[1](https://arxiv.org/html/2609.21554#S4.T1)contrasts our policy against two baselines on a held\-out MMLU slice\. Our policy matches the Full\-20 method accuracy while cutting inference cost by 60%\. It also outperforms a Random\-8 selector by\+10\.9pp, validating that the model learns*which*conceptual perspectives matter\.

Table 1:Held\-out comparison of selector policies\.![Refer to caption](https://arxiv.org/html/2609.21554v1/images/space_selection_distribution_improved.png)Figure 3:Final perspective\-Selection Distribution\.Normalized usage rates of each conceptual perspective after REINFORCE training \(λ=0\.05\\lambda=0\.05\)\. The height of each bar is the fraction of queries the selector directed to that perspective; numerical labels show percentage frequency\. Note the clear peak at Algebraic and Probabilistic perspective, followed by a long\-tail distribution across the remaining 18 perspective\.
##### Perspective preferences\.

As shown in Fig\.[3](https://arxiv.org/html/2609.21554#S4.F3), the trained selector develops a strong preference for a subset of highly predictive reasoning perspectives\. The top five—*Algebraic*\(11\.8%\),*Probabilistic*\(11\.6%\),*Network Flow*\(9\.8%\),*Dimensional Analysis*\(8\.7%\), and*Stochastic Process*\(7\.8%\)—account for nearly half of all selections\. This skewed distribution highlights the effectiveness of the learned policy in focusing computation on a few highly informative perspectives while avoiding low\-utility ones\. The long tail across the remaining perspectives suggests retained flexibility, allowing the system to fall back on niche reasoning modes when needed\.

### 4\.3Inference Phase: Dynamic perspective Selection and Sequential Reasoning

Algorithm 1Inference\-Time Multi\-Perspective Reasoning0:Query

qq, perspective set

𝒮\\mathcal\{S\}, selectorSelector, reasonerReasoner, confidence threshold

τ\\tau, max attempts

kk
1:\(0\)Receive query

qq
2:\(1\)

Sranked←Selector​\(q\)S\_\{\\text\{ranked\}\}\\leftarrow\\textsc\{Selector\}\(q\)\{Rank perspectives byp⁡\(si∣q\)p\(s\_\{i\}\\mid q\)\}

3:Initialize history

H←∅H\\leftarrow\\emptyset
4:for

t=1t=1TO

kkdo

5:

s\(t\)←Sranked​\[t\]s\_\{\(t\)\}\\leftarrow S\_\{\\text\{ranked\}\}\[t\]\{Select top\-ranked perspective\}

6:

q\(t\)←Transform​\(q,s\(t\),H\)q\_\{\(t\)\}\\leftarrow\\textsc\{Transform\}\(q,\\,s\_\{\(t\)\},\\,H\)\{Adapt query to current perspective\}

7:\(2\)

\(a\(t\),c\(t\)\)←Reasoner​\(q\(t\),s\(t\)\)\(a\_\{\(t\)\},c\_\{\(t\)\}\)\\leftarrow\\textsc\{Reasoner\}\(q\_\{\(t\)\},s\_\{\(t\)\}\)\{Answer and confidence\}

8:if

c\(t\)≥τc\_\{\(t\)\}\\geq\\tauthen

9:\(3\)RETURN

a\(t\)a\_\{\(t\)\}\{Return confident answer\}

10:endif

11:

H←H∪\{\(s\(t\),a\(t\),c\(t\)\)\}H\\leftarrow H\\cup\\\{\(s\_\{\(t\)\},a\_\{\(t\)\},c\_\{\(t\)\}\)\\\}\{Update reasoning history\}

12:endfor

13:\(4\)RETURN

Aggregate​\(\{\(a\(1\),c\(1\)\),…,\(a\(k\),c\(k\)\)\}\)\\textsc\{Aggregate\}\(\\\{\(a\_\{\(1\)\},c\_\{\(1\)\}\),\\dots,\(a\_\{\(k\)\},c\_\{\(k\)\}\)\\\}\)\{Fallback: aggregate answers\}

## 5Experiments and results

We evaluate the effectiveness of our proposed Selector\-Reasoner framework across four widely\-used benchmarks: GSM8K, MATH500, MMLU\-Pro, and the Game\-of\-24 task\. We compare our method against standard inference strategies, including simple prompting, Chain\-of\-Thought \(CoT\)\([Wei et al\., 2022](https://arxiv.org/html/2609.21554#bib.bib1)\), and DIPPER\([Lau et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib2)\)with 3 and 5 diverse prompts per problem\. Our evaluation considers various powerful baseline LLMs, including DeepSeek\-v3\([Team, 2024a](https://arxiv.org/html/2609.21554#bib.bib34)\), ChatGPT\-4o\([OpenAI, 2024](https://arxiv.org/html/2609.21554#bib.bib35)\), Claude 3\.7\-sonnet\([Anthropic, 2025](https://arxiv.org/html/2609.21554#bib.bib36)\), Gemini 2\.0 Flash 001\([Google, 2025](https://arxiv.org/html/2609.21554#bib.bib37)\), and Qwen2\.5\-7B\([Team, 2024c](https://arxiv.org/html/2609.21554#bib.bib29)\)\.

### 5\.1Results on GSM8K

Table[2](https://arxiv.org/html/2609.21554#S5.T2)illustrates performance improvements on the GSM8K dataset\([Cobbe et al\., 2021a](https://arxiv.org/html/2609.21554#bib.bib25)\)\. Gemini 2\.0 Flash achieves a remarkable accuracy of 95\.53%, surpassing the best CoT and DIPPER results by a large margin, while querying on average only 1\.02 perspectives per problem\. Similarly, Claude 3\.7\-sonnet and ChatGPT\-4o achieve high accuracy rates of 96\.21% and 92\.04%, respectively, with very low average queried perspectives \(approximately one per query\), highlighting not only superior accuracy but also remarkable inference efficiency\.

Table 2:Results on GSM8K\. MIRAGE outperforms others with minimal queried perspectives\.
### 5\.2Results on MATH500

Table[3](https://arxiv.org/html/2609.21554#S5.T3)presents results on the challenging MATH500 dataset\([Hendrycks et al\., 2021b](https://arxiv.org/html/2609.21554#bib.bib26)\), a benchmark known for its complexity and depth in mathematical reasoning\. Our method significantly outperforms all baselines across all models, achieving accuracy improvements\. Notably, Gemini 2\.0 Flash achieves an accuracy of 84\.40%\. results, demonstrating substantial capability in solving advanced mathematical problems with a relatively low query overhead \(average of 1\.45 perspectives per problem\)\.

Table 3:Results on MATH500\. MIRAGE shows strong gains over baseline methods\.
### 5\.3Results on MMLU\-Pro

In Table[4](https://arxiv.org/html/2609.21554#S5.T4), we evaluate performance on the MMLU\-Pro benchmark\([Wang et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib5)\), a diverse and challenging test of general reasoning ability across multiple scientific and logical domains\. For our evaluation, we used thetest setand selected four subjects—math, physics, chemistry, and engineering—taking the first 100 questions from each domain\. Our method consistently achieves superior accuracy across all base models\. Specifically, Gemini 2\.0 Flash achieves 83\.75% accuracy, a notable gain over the baseline 55\.78%, with an efficient average perspective usage of only 2\.14 per problem\.

Table 4:Results on MMLU\-Pro\. MIRAGE demonstrates robustness in general reasoning tasks\.
### 5\.4Results on Game\-of\-24

Table[5](https://arxiv.org/html/2609.21554#S5.T5)shows the results on the Game\-of\-24\([nlile, 2025](https://arxiv.org/html/2609.21554#bib.bib27)\)reasoning task\. Again, our Selector\-Reasoner approach markedly surpasses baseline performances\. Gemini 2\.0 Flash achieves nearly perfect accuracy \(99\.20%\) with extremely low computational overhead \(average 1\.13 perspectives per problem\), highlighting our method’s generalizability and efficiency even in highly structured logical reasoning scenarios\.

Table 5:Results on Game\-of\-24\. MIRAGE shows exceptional gains and efficiency\.Collectively, these empirical results demonstrate the broad efficacy, efficiency, and adaptability of our cognitive\-inspired Selector\-Reasoner approach in significantly enhancing the reasoning capabilities of modern LLMs across diverse and challenging tasks\.

##### Statistical Rigor\.

Our reported results are based on single\-run executions per experiment setting, and we do not include variance estimates such as error bars or confidence intervals\. While this is a limitation, our conclusions are grounded in a broad and systematic evaluation that enhances robustness\. Specifically, we conduct evaluations on four diverse benchmarks \(GSM8K, MATH500, MMLU\-Pro, Game\-of\-24\), across five base models, and under multiple prompting and reasoning paradigms—including direct prompting, Chain\-of\-Thought, and DIPPER \(n=3n\{=\}3andn=5n\{=\}5\)\. Furthermore, this wide empirical coverage strengthens the reliability and generalizability of our findings\.

### 5\.5Accuracy vs\. Computational Cost Analysis

![Refer to caption](https://arxiv.org/html/2609.21554v1/images/accuracy_vs_cost_combined.png)Figure 4:Accuracy vs\. computational cost across four benchmarks \(GSM8K, MATH500, MMLU\-Pro, and Game\-of\-24\)\. Each point corresponds to a \(base model, prompting method\) pair\. Marker shape denotes the model, while color encodes the prompting strategy\. Our method consistently achieves high accuracy with minimal computational cost \(Top\-Left is better\)\.To evaluate the overall performance of our method across diverse reasoning tasks, we visualize accuracy against computational cost \(measured in terms of average inference calls per sample\) for four representative benchmarks:GSM8K,MATH500,MMLU\-Pro, andGame\-of\-24\(Figure[4](https://arxiv.org/html/2609.21554#S5.F4)\)\. Each method is shown using a distinct color, and each base model uses a unique marker shape for clarity\.

Across all datasets, our method consistently achieves a superior balance between accuracy and cost, outperforming standard prompting baselines like Chain\-of\-Thought \(CoT\) and Monte Carlo Sampling \(MCS, or DIPPER\)\. While DIPPER withn=5n=5queries achieves competitive accuracy, it incurs up to 5×\\timesthe inference cost\. In contrast, our method attains equal or better accuracy with only 1–2 queries on average\.

- •GSM8K:Our method yields the highest accuracy across all models \(up to 96\.2%\) while maintaining a modest cost \(1\.1–2\.3×\\times\)\.
- •MATH500:Our method achieves top\-tier accuracy on Claude and Gemini \(up to 96\.2%\) with nearly half the computational cost of DIPPER \(n=5n=5\)\.
- •MMLU\-Pro:Even under general\-purpose reasoning, our method improves performance over CoT by 20–30% absolute accuracy, using only 2–4 queried perspectives\.
- •Game\-of\-24:We observe the most dramatic gains here, with accuracy reaching 99\.2% at significantly lower cost than multi\-sample baselines\.

These results demonstrate that our method is not only accurate but also highly efficient\. It generalizes well across models and domains, making it suitable for real\-world applications where latency and budget constraints are critical\.

##### Compute Resources\.

We use an NVIDIA A100 40GB GPU for all experiments\. Training the Selector model \(Qwen2\.5\-7B\) requires fine\-tuning on solved examples from MMLU using outputs from the Qwen2\.5\-14B model as a Reasoner\. Both models are publicly available and open access\. The Selector training process took approximately 10 GPU\-hours\.

During inference, the computational cost is dominated by the Reasoner model, which is queried conditionally based on the Selector output\. The Selector itself is lightweight: it generates only a small number of tokens representing perspective names \(e\.g\., “Algebraic”, “Probabilistic”\) and can be executed with minimal overhead, comparable to a single forward pass of standard prompting\.

## 6Conclusion and discussion

We introducedMIRAGE, an inference\-time framework that learns to route each problem to the conceptual reasoning perspective—algebraic, probabilistic, game\-theoretic, and more—most likely to yield a correct solution\. Unlike prior approaches that rely on a single prompt, costly multi\-sample decoding, or extensive finetuning, MIRAGE combines a lightweight*Selector*with a perspective\-aware*Reasoner*, requiring on average fewer than two LLM calls for three of four benchmarks\. Comprehensive experiments on GSM8K, MATH500, MMLU\-Pro, and Game\-of\-24, spanning five base models, show that MIRAGE delivers up to\+24\.7 ppabsolute accuracy over Chain\-of\-Thought while using as little as15\\tfrac\{1\}\{5\}the computational budget of DIPPER \(n=5n\{=\}5\)\. These gains confirm that dynamically shifting representational frames—a hallmark of human cognitive flexibility—can be operationalized in modern LLMs for both effectiveness and efficiency\.

##### Limitations and future work\.

Although our twenty predefined perspectives cover a broad spectrum of mathematical and logical reasoning, they remain discrete and manually crafted\. Scaling to open\-domain tasks will require*\(i\)*automatic discovery or synthesis of new perspectives,*\(ii\)*richer confidence estimation for early stopping, and*\(iii\)*tighter integration with symbolic tools or external knowledge bases\.Statistical uncertainty\.All reported numbers stem from single\-run executions due to computational budget constraints; consequently, we do not provide error bars or confidence intervals\. While we partially offset this by evaluating on four diverse benchmarks, five backbone models, and multiple strong baselines, future work will perform multi\-seed experiments to quantify variance and strengthen statistical rigor\. Moreover, selector training currently assumes access to solved examples; semi\-supervised or reinforcement learning in the wild is an important next step\.

By demonstrating that multi\-perspective selection can match or surpass state\-of\-the\-art accuracy at a fraction of the cost,MIRAGEopens a practical path toward deployable, resource\-aware reasoning systems—underscoring the value of cognitive\-science principles for guiding future LLM research\.

## 7Ablation Study

To investigate the contributions of individual components of our Selector\-Reasoner framework, we perform a series of ablation experiments, systematically removing or modifying key components\. We specifically evaluate the importance of \(1\) Per\-perspective Accuracy Analysis, \(2\) the dynamic selection of conceptual reasoning perspectives, \(3\) the aggregation step of multi\-perspective outputs, and \(4\) the total number and types of conceptual perspectives included\.

##### Per\-perspective Accuracy Analysis\.

We conducted an in\-depth per\-perspective performance analysis on the GSM8K dataset using the Qwen2\.5\-7B model as reasoner to investigate the effectiveness of individual conceptual reasoning perspectives \(see Figure[5](https://arxiv.org/html/2609.21554#S7.F5)\)\. Individual perspectives exhibit notable variations in accuracy, with the highest\-performing perspectives including Info\-Theoric \(74\.1%\), and Probabilistic \(73\.6%\)\. Despite these strong individual performances, none of the single perspectives alone achieve accuracy comparable to aggregating predictions across all perspectives \(88\.4%\)\. Further comparisons against established baseline inference methods clearly demonstrate the efficacy of our approach\. The simple prompting method yields a baseline accuracy of 67\.0%, significantly lower than the single\-perspective results\. More advanced prompting techniques such as Chain\-of\-Thought \(CoT\) and DIPPER with 3 and 5 ensembles achieve moderate improvements \(79\.8% and 83\.1%, respectively\)\. However, our Selector\-driven aggregation method surpasses all these techniques, achieving the highest accuracy \(89\.0%\) while utilizing a minimal average of only 1\.97 queried perspectives per problem\. This result highlights not only superior performance but also computational efficiency, validating the necessity and effectiveness of our dynamic selection and multi\-perspective aggregation strategy\.

![Refer to caption](https://arxiv.org/html/2609.21554v1/images/final_accuracy_comparison_plot.png)Figure 5:Unified accuracy comparison on the GSM8K dataset using the Qwen2\.5\-7B model\. Individual reasoning perspectives \(light blue\) vary significantly in accuracy\. Aggregation across all perspectives \(dark blue\) significantly surpasses any single perspective\. Our Selector\-driven approach \(orange\) outperforms all baseline methods \(patterned bars\), achieving the highest accuracy with efficient use of computational resources \(k<2k<2\)\.The remaining experiments—\(2\) dynamic versus random/fixed selection, \(3\) aggregation variants, and \(4\) sensitivity to the number of queried perspectives—are provided in Appendix[E](https://arxiv.org/html/2609.21554#A5)

##### Broader Impacts\.

MIRAGE offers practical benefits by improving LLM reasoning efficiency, especially for STEM tasks, while reducing inference cost by up to 5×\\timescompared to ensemble\-style methods\. This efficiency supports deployment in educational or resource\-constrained settings\. However, it also introduces potential risks such as misuse for deceptive reasoning or automated homework\-solving\. We expose rationale steps and confidence scores to aid transparency and plan to release all code and hyperparameters\. Additional societal risks, mitigations, and environmental considerations are detailed in Appendix[D](https://arxiv.org/html/2609.21554#A4)\.

## References

- Anthropic \(2025\)AnthropicClaude 3\.7 sonnet system card\.Note:[https://www\.anthropic\.com/claude\-3\-7\-sonnet\-system\-card](https://www.anthropic.com/claude-3-7-sonnet-system-card)Accessed: 2025\-05\-16Cited by:[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- Battagliaet al\.\(2013\)P\. W\. Battagliaet al\.Simulation as an engine of physical scene understanding\.PNAS110\(45\),pp\. 18327–18332\.Cited by:[§A\.1\.6](https://arxiv.org/html/2609.21554#A1.SS1.SSS6.Px1.p1.1)\.
- Battagliaet al\.\(2018\)P\. W\. Battagliaet al\.Relational inductive biases, deep learning, and graph networks\.arXiv preprint arXiv:1806\.01261\.Cited by:[§A\.1\.3](https://arxiv.org/html/2609.21554#A1.SS1.SSS3.Px1.p1.1)\.
- Buckingham \(1914\)E\. BuckinghamOn physically similar systems; illustrations of the use of dimensional equations\.Physical Review4\(4\),pp\. 345\.Cited by:[§A\.1\.6](https://arxiv.org/html/2609.21554#A1.SS1.SSS6.Px2.p1.1)\.
- Cobbeet al\.\(2021a\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Note:GSM8K: Grade School Math 8K datasetCited by:[§5\.1](https://arxiv.org/html/2609.21554#S5.SS1.p1.1)\.
- Cobbeet al\.\(2021b\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, and C\. HesseTraining verifiers to solve math word problems\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Deen and Freiwald \(2021\)B\. Deen and W\. A\. FreiwaldParallel systems for social and spatial reasoning in the brain\.bioRxiv\.External Links:[Link](https://www.biorxiv.org/content/10.1101/2021.09.04.458988v1)Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p3.1),[§3](https://arxiv.org/html/2609.21554#S3.p3.1)\.
- Dehaene \(2009\)S\. DehaeneOrigins of mathematical intuitions: the case of arithmetic\.Annals of the New York Academy of Sciences1156\(1\),pp\. 232–259\.Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p3.1),[§3](https://arxiv.org/html/2609.21554#S3.p3.1)\.
- Ford and Fulkerson \(1956\)L\. R\. Ford and D\. R\. FulkersonMaximal flow through a network\.Canadian Journal of Mathematics8,pp\. 399–404\.Cited by:[§A\.1\.3](https://arxiv.org/html/2609.21554#A1.SS1.SSS3.Px2.p1.1)\.
- Gaoet al\.\(2022\)L\. Gao, A\. Madaan, S\. Zhou, U\. Alon, P\. Liu, Y\. Yang, and G\. NeubigPAL: program\-aided language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Gigerenzer and Hoffrage \(1995\)G\. Gigerenzer and U\. HoffrageHow to improve bayesian reasoning without instruction: frequency formats\.Psychological Review102\(4\),pp\. 684–704\.Cited by:[§3](https://arxiv.org/html/2609.21554#S3.p3.1)\.
- Google \(2025\)GoogleGemini 2\.0 flash\.Note:[https://cloud\.google\.com/vertex\-ai/generative\-ai/docs/models/gemini/2\-0\-flash](https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-0-flash)Accessed: 2025\-05\-16Cited by:[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- Hendryckset al\.\(2021a\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2009.03300)Cited by:[§4\.2](https://arxiv.org/html/2609.21554#S4.SS2.SSS0.Px1.p1.1)\.
- Hendryckset al\.\(2021b\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the math dataset\.NeurIPS\.Note:MATH: 12,500 competition\-level math problemsCited by:[§5\.2](https://arxiv.org/html/2609.21554#S5.SS2.p1.1)\.
- Johnson\-Laird and Byrne \(1991\)P\. N\. Johnson\-Laird and R\. M\. J\. ByrneDeduction\.Psychology Press\.Cited by:[§A\.1\.1](https://arxiv.org/html/2609.21554#A1.SS1.SSS1.Px1.p1.1),[§A\.1\.1](https://arxiv.org/html/2609.21554#A1.SS1.SSS1.Px2.p1.1)\.
- Kirsh and Maglio \(1994\)D\. Kirsh and P\. MaglioOn distinguishing epistemic from pragmatic action\.Cognitive Science18\(4\),pp\. 513–549\.Cited by:[§A\.1\.2](https://arxiv.org/html/2609.21554#A1.SS1.SSS2.Px1.p1.1)\.
- Knoblichet al\.\(1999\)G\. Knoblich, S\. Ohlsson, H\. Haider, and D\. RheniusConstraint relaxation and chunk decomposition in insight problem solving\.Journal of Experimental Psychology: Learning, Memory, and Cognition25\(6\),pp\. 1534–1555\.Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p3.1),[§3](https://arxiv.org/html/2609.21554#S3.p4.1)\.
- Kojimaet al\.\(2022\)T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. IwasawaLarge language models are zero\-shot reasoners\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§1](https://arxiv.org/html/2609.21554#S1.p2.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Kuhn \(1956\)H\. W\. KuhnVariants of the hungarian method for assignment problems\.Naval Research Logistics Quarterly3\(4\),pp\. 253–258\.Cited by:[§A\.1\.3](https://arxiv.org/html/2609.21554#A1.SS1.SSS3.Px2.p1.1)\.
- Lample and Charton \(2020\)G\. Lample and F\. ChartonDeep learning for symbolic mathematics\.arXiv preprint arXiv:2006\.16283\.Cited by:[§A\.1\.7](https://arxiv.org/html/2609.21554#A1.SS1.SSS7.Px2.p1.1)\.
- Larkin and Simon \(1987\)J\. H\. Larkin and H\. A\. SimonWhy a diagram is \(sometimes\) worth ten thousand words\.Cognitive Science11\(1\),pp\. 65–100\.Cited by:[§A\.1\.1](https://arxiv.org/html/2609.21554#A1.SS1.SSS1.Px2.p1.1),[§A\.1\.2](https://arxiv.org/html/2609.21554#A1.SS1.SSS2.Px1.p1.1)\.
- Lauet al\.\(2024\)G\. K\. R\. Lau, W\. Hu, D\. Liu, J\. Chen, S\. Ng, and B\. K\. H\. LowDipper: diversity in prompts for producing large language model ensembles in reasoning tasks\.Vol\.abs/2412\.15238\.Cited by:[Appendix C](https://arxiv.org/html/2609.21554#A3.p3.pic1.2.1.1.1),[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§1](https://arxiv.org/html/2609.21554#S1.p2.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1),[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- Liet al\.\(2020\)Z\. Liet al\.Fourier neural operator for parametric partial differential equations\.arXiv preprint arXiv:2010\.08895\.Cited by:[§A\.1\.5](https://arxiv.org/html/2609.21554#A1.SS1.SSS5.Px1.p1.1)\.
- Maddux and Galinsky \(2009\)W\. W\. Maddux and A\. D\. GalinskyCultural borders and mental barriers: the relationship between living abroad and creativity\.Journal of Personality and Social Psychology96\(5\),pp\. 1047–1061\.Cited by:[§3](https://arxiv.org/html/2609.21554#S3.p4.1)\.
- Mikolovet al\.\(2013\)T\. Mikolovet al\.Efficient estimation of word representations in vector space\.arXiv preprint arXiv:1301\.3781\.Cited by:[§A\.1\.5](https://arxiv.org/html/2609.21554#A1.SS1.SSS5.Px2.p1.1)\.
- Miyakeet al\.\(2000\)A\. Miyake, N\. P\. Friedman, M\. J\. Emerson, A\. H\. Witzki, A\. Howerter, and T\. D\. WagerThe unity and diversity of executive functions and their contributions to complex “frontal lobe” tasks: a latent variable analysis\.Cognitive Psychology41\(1\),pp\. 49–100\.Cited by:[§3](https://arxiv.org/html/2609.21554#S3.p1.1)\.
- Mnihet al\.\(2015\)V\. Mnihet al\.Human\-level control through deep reinforcement learning\.Nature518\(7540\),pp\. 529–533\.Cited by:[§A\.1\.4](https://arxiv.org/html/2609.21554#A1.SS1.SSS4.Px3.p1.1)\.
- Newell and Simon \(1972\)A\. Newell and H\. A\. SimonHuman problem solving\.Prentice\-Hall\.Cited by:[§A\.1\.7](https://arxiv.org/html/2609.21554#A1.SS1.SSS7.Px1.p1.1)\.
- nlile \(2025\)nlileGame of 24 dataset\.Note:[https://huggingface\.co/datasets/nlile/24\-game](https://huggingface.co/datasets/nlile/24-game)1,362 puzzles scraped from 4nums\.com; access date May 14, 2025Cited by:[§5\.4](https://arxiv.org/html/2609.21554#S5.SS4.p1.1)\.
- Nyeet al\.\(2022\)M\. Nye, A\. Andreassen, G\. Gur\-Ari, H\. Michalewski, J\. Austin, D\. Bieber, D\. Dohan, A\. Lewkowycz, M\. Bosma, and D\. LuanShow your work: scratchpads for intermediate computation with language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Nyeet al\.\(2021\)M\. I\. Nye, A\. J\. Andreassen, G\. Gur\-Ari, H\. Michalewski, J\. Austin, D\. Bieber, D\. Dohan, A\. Lewkowycz, M\. Bosma, D\. Luan, C\. Sutton, and A\. OdenaShow your work: scratchpads for intermediate computation with language models\.CoRRabs/2112\.00114\.Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1)\.
- OpenAI \(2024\)OpenAIChatGPT\-4o\.Note:[https://openai\.com/blog/chatgpt\-4o](https://openai.com/blog/chatgpt-4o)Accessed: 2025\-05\-16Cited by:[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- OpenStax \(2023\)OpenStaxCollege physics\.Note: urlhttps://openstax\.org/books/college\-physics/pages/1\-introductionOpen textbookCited by:[§A\.1\.6](https://arxiv.org/html/2609.21554#A1.SS1.SSS6.Px2.p1.1)\.
- Pearl \(1988\)J\. PearlProbabilistic reasoning in intelligent systems: networks of plausible inference\.Morgan Kaufmann\.Cited by:[§A\.1\.4](https://arxiv.org/html/2609.21554#A1.SS1.SSS4.Px1.p1.1)\.
- Rittle\-Johnson and Star \(2007\)B\. Rittle\-Johnson and J\. R\. StarDoes comparing solution methods facilitate conceptual and procedural knowledge? an experimental study on learning to solve equations\.Journal of Educational Psychology99\(3\),pp\. 561–574\.Cited by:[§3](https://arxiv.org/html/2609.21554#S3.p3.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2302.04761)Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Shannon \(1948\)C\. E\. ShannonA mathematical theory of communication\.Bell System Technical Journal27\(3\),pp\. 379–423\.Cited by:[§A\.1\.4](https://arxiv.org/html/2609.21554#A1.SS1.SSS4.Px2.p1.1)\.
- Shenet al\.\(2023\)Y\. Shen, K\. Song, X\. Tan, D\. Li, W\. Lu, and Y\. ZhuangHuggingGPT: solving ai tasks with chatgpt and its friends in hugging face\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2303.17580)Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, B\. Labash, A\. Gopinath, D\. Krasheninnikov, A\. S\. Das, R\. Ahuja, and F\. MuellerReflexion: an autonomous agent with dynamic memory and self\-reflection\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Siegler \(1996\)R\. S\. SieglerEmerging minds: the process of change in children’s thinking\.Oxford University Press\.Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p3.1),[§3](https://arxiv.org/html/2609.21554#S3.p1.1),[§3](https://arxiv.org/html/2609.21554#S3.p2.1)\.
- Silveret al\.\(2016\)D\. Silveret al\.Mastering the game of go with deep neural networks and tree search\.Nature529\(7587\),pp\. 484–489\.Cited by:[§A\.1\.7](https://arxiv.org/html/2609.21554#A1.SS1.SSS7.Px2.p1.1),[§A\.1\.7](https://arxiv.org/html/2609.21554#A1.SS1.SSS7.Px3.p1.1)\.
- Spiroet al\.\(1988\)R\. J\. Spiro, R\. L\. Coulson, P\. J\. Feltovich, and D\. K\. AndersonCognitive flexibility theory: advanced knowledge acquisition in ill\-structured domains\.Technical reportTechnical ReportTechnical Report No\. 441,ERIC\.External Links:[Link](https://files.eric.ed.gov/fulltext/ED302821.pdf)Cited by:[§3](https://arxiv.org/html/2609.21554#S3.p1.1),[§3](https://arxiv.org/html/2609.21554#S3.p2.1)\.
- Spivak \(2014\)D\. I\. SpivakCategory theory for the sciences\.MIT Press\.Cited by:[§A\.1\.1](https://arxiv.org/html/2609.21554#A1.SS1.SSS1.Px3.p1.1)\.
- Sutton and Barto \(1998\)R\. S\. Sutton and A\. G\. BartoReinforcement learning: an introduction\.MIT Press\.Cited by:[§A\.1\.4](https://arxiv.org/html/2609.21554#A1.SS1.SSS4.Px3.p1.1)\.
- Team \(2024a\)D\. TeamDeepSeek\-v3 technical report\.Note:[https://arxiv\.org/abs/2412\.19437](https://arxiv.org/abs/2412.19437)Accessed: 2025\-05\-16Cited by:[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- Team \(2024b\)Q\. TeamQwen2\.5\-14b\-instruct\.Cited by:[§4\.2](https://arxiv.org/html/2609.21554#S4.SS2.SSS0.Px1.p1.1.3)\.
- Team \(2024c\)Q\. TeamQwen2\.5: a party of foundation models\.Cited by:[§4\.2](https://arxiv.org/html/2609.21554#S4.SS2.SSS0.Px1.p1.1.2),[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- Tenenbaumet al\.\(2000\)J\. B\. Tenenbaum, V\. de Silva, and J\. C\. LangfordA global geometric framework for nonlinear dimensionality reduction\.Science290\(5500\),pp\. 2319–2323\.Cited by:[§A\.1\.2](https://arxiv.org/html/2609.21554#A1.SS1.SSS2.Px3.p1.1)\.
- Tenenbaumet al\.\(2011\)J\. B\. Tenenbaum, C\. Kemp, T\. L\. Griffiths, and N\. D\. GoodmanHow to grow a mind: statistics, structure, and abstraction\.Science331\(6022\),pp\. 1279–1285\.Cited by:[§A\.1\.1](https://arxiv.org/html/2609.21554#A1.SS1.SSS1.Px1.p1.1),[§A\.1\.4](https://arxiv.org/html/2609.21554#A1.SS1.SSS4.Px1.p1.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§1](https://arxiv.org/html/2609.21554#S1.p2.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Wanget al\.\(2024\)Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, T\. Li, M\. Ku, K\. Wang, A\. Zhuang, R\. Fan, X\. Yue, and W\. ChenMMLU\-pro: a more robust and challenging multi\-task language understanding benchmark\.External Links:2406\.01574Cited by:[§5\.3](https://arxiv.org/html/2609.21554#S5.SS3.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1),[§5](https://arxiv.org/html/2609.21554#S5.p1.1)\.
- Yaoet al\.\(2023a\)S\. Yao, D\. Yu, J\. Zhao, I\. Shafran, K\. Narasimhan, and Y\. CaoTree of thoughts: deliberate problem solving with large language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Yaoet al\.\(2023b\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Yaoet al\.\(2024\)W\. Yao, H\. Mi, and D\. YuHDFlow: enhancing llm complex problem\-solving with hybrid thinking and dynamic workflows\.arXiv preprint arXiv:2409\.17433\.External Links:[Link](https://arxiv.org/abs/2409.17433)Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Zelikmanet al\.\(2022\)E\. Zelikman, Y\. Wu, J\. Mu, and N\. D\. GoodmanSTaR: bootstrapping reasoning with reasoning\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24899–24912\.External Links:[Link](https://arxiv.org/abs/2203.14465)Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Zhanget al\.\(2025\)X\. Zhang, J\. Xu, and M\. HuangAtomic reasoner: fine\-grained cognitive routing for large language models\.arXiv preprint arXiv:2504\.06789\.Cited by:[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.
- Zhouet al\.\(2023\)D\. Zhou, N\. Schärli, L\. Hou, J\. Wei, N\. Scales, X\. Wang, D\. Schuurmans, O\. Bousquet, and Q\. V\. LeLeast\-to\-most prompting enables complex reasoning in large language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.21554#S1.p1.1),[§2](https://arxiv.org/html/2609.21554#S2.p1.1)\.

## Appendix AJustification and Illustrative Examples of Conceptual Reasoning Perspectives

### A\.1Diverse Reasoning Perspectives in the MIRAGE Framework

In this section, we justify the selection of twenty reasoning perspectives incorporated into the MIRAGE framework\. These perspectives grounded in cognitive science and AI literature\.

#### A\.1\.1Symbolic and Formal Reasoning Perspectives

##### Algebraic & Symbolic Logical Reasoning\.

Humans and AI alike benefit from formal symbolic reasoning strategies\. Some problem\-solvers prefer manipulating equations or applying formal logic rules, while others use more visual means\([Tenenbaum et al\., 2011](https://arxiv.org/html/2609.21554#bib.bib38)\)\. Cognitive studies on syllogistic puzzles show that many people naturally employ logical algebraic strategies, indicating the importance of an algebraic and symbolic logic perspective\([Johnson\-Laird and Byrne, 1991](https://arxiv.org/html/2609.21554#bib.bib39)\)\. This underscores that algebraic equation\-solving and logical deduction are foundational modes of reasoning that MIRAGE should support\.

##### Set\-Theoretic Reasoning\.

A set\-theoretic perspective \(e\.g\., thinking in terms of sets, Venn/Euler diagrams\) offers an intuitive way to tackle logic and categorization problems\([Johnson\-Laird and Byrne, 1991](https://arxiv.org/html/2609.21554#bib.bib39)\)\. Diagrams explicitly preserve topological relations \(e\.g\., overlap, containment\) that are only implicit in sentences\([Larkin and Simon, 1987](https://arxiv.org/html/2609.21554#bib.bib40)\), helping reduce cognitive effort in reasoning\.

##### Category\-Theoretic Reasoning\.

Category theory provides a high\-level formal perspective that can unify and connect concepts across domains\. It supports compositional reasoning and abstraction\([Spivak, 2014](https://arxiv.org/html/2609.21554#bib.bib42)\), which are increasingly recognized in machine learning as tools for reasoning about analogies and structural similarity\.

#### A\.1\.2Spatial and Geometric Reasoning Perspectives

##### Geometric & Visual Reasoning\.

Diagrams and spatial representations help reduce reasoning complexity by encoding constraints visually\([Larkin and Simon, 1987](https://arxiv.org/html/2609.21554#bib.bib40);[Kirsh and Maglio, 1994](https://arxiv.org/html/2609.21554#bib.bib41)\)\. Many geometry proofs, physics diagrams, and engineering schematics rely on spatial intuition that cannot be replaced by symbolic manipulation alone\.

##### Topological Reasoning\.

Topology abstracts away metric details and focuses on connectivity or continuity\. It has applications in qualitative spatial reasoning, robotics, and topological data analysis\. Euler’s solution to the Königsberg bridge problem exemplifies how topology reveals structure in problems\.

##### Differential Geometry\.

This perspective allows reasoning on smooth manifolds and curvature\. Tenenbaum et al\. introduced Isomap to uncover low\-dimensional manifolds in high\-dimensional data, showing that many real\-world problems benefit from a differential\-geometric lens\([Tenenbaum et al\., 2000](https://arxiv.org/html/2609.21554#bib.bib43)\)\.

#### A\.1\.3Graph and Network Reasoning Perspectives

##### Graph\-Based Reasoning\.

Graph representations enable relational reasoning and have proven effective in cognitive problem solving \(e\.g\., family trees, dependencies\) and AI\([Battaglia and others, 2018](https://arxiv.org/html/2609.21554#bib.bib44)\)\.

##### Network Flow Reasoning\.

Network flow models handle constraints and optimization in allocation and routing problems\. This view encourages constraint satisfaction through graph structures and complements relational graph reasoning\([Ford and Fulkerson, 1956](https://arxiv.org/html/2609.21554#bib.bib45);[Kuhn, 1956](https://arxiv.org/html/2609.21554#bib.bib46)\)\.

#### A\.1\.4Probabilistic and Information\-Theoretic Perspectives

##### Probabilistic Reasoning\.

Probabilistic reasoning allows managing uncertainty\. Bayesian networks introduced by Pearl\([Pearl, 1988](https://arxiv.org/html/2609.21554#bib.bib47)\)and Bayesian models of human cognition\([Tenenbaum et al\., 2011](https://arxiv.org/html/2609.21554#bib.bib38)\)exemplify this perspective\.

##### Information\-Theoretic Reasoning\.

Shannon’s theory of information\([Shannon, 1948](https://arxiv.org/html/2609.21554#bib.bib48)\)guides exploration, compression, and uncertainty reduction in AI and cognitive science\.

##### Stochastic Process Reasoning\.

Stochastic models like Markov chains and MDPs capture sequential decision making under uncertainty, critical in reinforcement learning\([Sutton and Barto, 1998](https://arxiv.org/html/2609.21554#bib.bib49);[Mnih and others, 2015](https://arxiv.org/html/2609.21554#bib.bib50)\)\.

#### A\.1\.5Analytical and Transformational Perspectives

##### Fourier/Frequency Reasoning\.

Frequency domain analysis simplifies convolution, periodicity, and PDE solutions\. Fourier Neural Operators demonstrate the efficacy of frequency\-based reasoning in AI\([Li and others, 2020](https://arxiv.org/html/2609.21554#bib.bib51)\)\.

##### Tensor/Matrix Reasoning\.

Linear algebra supports embeddings, transformations, and high\-dimensional computation in AI and human reasoning\([Mikolov and others, 2013](https://arxiv.org/html/2609.21554#bib.bib52)\)\.

#### A\.1\.6Physical and Dimensional Reasoning Perspectives

##### Physics\-Based Reasoning\.

Humans often simulate physical processes mentally\. AI systems also learn physics\-based stability and control from visual data\([Battaglia and others, 2013](https://arxiv.org/html/2609.21554#bib.bib53)\)\.

##### Dimensional Analysis\.

This method helps validate units, derive formulas, and catch errors without full derivations\([Buckingham, 1914](https://arxiv.org/html/2609.21554#bib.bib54);[OpenStax, 2023](https://arxiv.org/html/2609.21554#bib.bib55)\)\.

#### A\.1\.7Computational, Learning, and Optimization Perspectives

##### Optimization Reasoning\.

AI and humans alike solve problems via optimization \(e\.g\., shortest paths, maximizing utility\)\([Newell and Simon, 1972](https://arxiv.org/html/2609.21554#bib.bib56)\)\.

##### Machine Learning & Computational

Learning from data to generalize patterns is key in modern AI\. Neural reasoning solvers and AlphaGo’s hybrid architecture exemplify this\([Silver and others, 2016](https://arxiv.org/html/2609.21554#bib.bib57);[Lample and Charton, 2020](https://arxiv.org/html/2609.21554#bib.bib58)\)\.

##### Game\-Theoretic Reasoning\.

Strategic reasoning about agents, adversaries, or incentives is modeled effectively using game theory, central in multi\-agent systems\([Silver and others, 2016](https://arxiv.org/html/2609.21554#bib.bib57)\)\.

### A\.2Why Multi\-Perspective Reasoning Matters\.

Consider the classic*task scheduling problem*, where we are given a set of tasks along with constraints such as ”Task A must precede Task B”, ”Task C and D cannot overlap”, and so on\.

Algebraic perspective:Representing these constraints algebraically leads to a system of inequalities over task start times \(e\.g\.,xA\+dA≤xBx\_\{A\}\+d\_\{A\}\\leq x\_\{B\}\), forming a linear programming formulation\. While mathematically precise, this system can grow rapidly in complexity and become hard to inspect or solve intuitively\.

Graph perspective:Alternatively, we can model the problem as a*directed acyclic graph*\(DAG\), where each task is a node and each precedence constraint is a directed edge\. Solving the scheduling problem now reduces to finding a*topological sort*of the DAG—a well\-known algorithmic problem that can be solved efficiently in linear time\.

This example illustrates how reframing a problem from an algebraic to a graph perspective can yield simpler, more interpretable solutions\. Such shifts are central to MIRAGE’s success: by dynamically choosing among conceptual perspectives, it adapts the reasoning lens to the problem structure, often simplifying complex tasks\.

### A\.3An example of solving one problem in different perspective \- Model: CatGPT\-4o

An example of solving one problem in different perspectives \- Model: ChatGPT\-4oProblem:Prove that the sum of the firstnnodd numbers is equal ton2n^\{2\}, i\.e\.,Sn=1\+3\+5\+⋯\+\(2​n−1\)=n2S\_\{n\}=1\+3\+5\+\\cdots\+\(2n\-1\)=n^\{2\}1\. Algebraic perspective:Use the identity∑k=1n\(2​k−1\)=n2\\sum\_\{k=1\}^\{n\}\(2k\-1\)=n^\{2\}\.Final Answer: <n2\>2\. Geometric perspective:Build ann×nn\\times nsquare layer\-by\-layer; each layer adds one odd number\.Final Answer: <n2\>3\. Graph perspective:Model a path graph with edge weights2​k−12k\-1; total cost isn2n^\{2\}\.Final Answer: <n2\>4\. Probabilistic perspective:LetXk=2​k−1X\_\{k\}=2k\-1; sum of expectations overnnterms yieldsn2n^\{2\}\.Final Answer: <n2\>5\. Topological perspective:Use nested sets of cardinality growing as odd numbers; Euler characteristic isn2n^\{2\}\.Final Answer: <n2\>6\. Symbolic Logic perspective:Inductive proof: base case1=121=1^\{2\}; assumeSn=n2S\_\{n\}=n^\{2\}, thenSn\+1=\(n\+1\)2S\_\{n\+1\}=\(n\+1\)^\{2\}\.Final Answer: <n2\>7\. Fourier/Frequency perspective:Represent the step function with odd coefficients; DC component integrates ton2n^\{2\}\.Final Answer: <n2\>8\. Information\-Theoretic perspective:Sequence of odd numbers encodes linearly; cumulative bits∼n2\\sim n^\{2\}\.Final Answer: <n2\>9\. Tensor/Matrix perspective:Vector dot product:v=\[1,3,…,2​n−1\]v=\[1,3,\.\.\.,2n\{\-\}1\], thenv⋅𝟏=n2v\\cdot\\mathbf\{1\}=n^\{2\}\.Final Answer: <n2\>10\. Physics\-Based perspective:Sequential forces \(odd units\) over unit time steps → total energy isn2n^\{2\}\.Final Answer: <n2\>11\. Optimization perspective:Objective: maximize∑\(2​k−1\)\\sum\(2k\-1\)under sizenn; max value isn2n^\{2\}\.Final Answer: <n2\>12\. Computational perspective:Code confirms identity:sum\(\[2k\-1 for k in range\(1,n\+1\)\]\) == n\*\*2\.Final Answer: <n2\>13\. Machine Learning perspective:Model trained on\(n,Sn\)\(n,S\_\{n\}\)learns mappingSn=n2S\_\{n\}=n^\{2\}\.Final Answer: <n2\>14\. Set\-Theoretic perspective:Set of odd integers\{1,3,…,2​n−1\}\\\{1,3,\.\.\.,2n\{\-\}1\\\}has sumn2n^\{2\}\.Final Answer: <n2\>15\. Game\-Theoretic perspective:Players alternate adding odd numbers; game sum equilibrium isn2n^\{2\}\.Final Answer: <n2\>16\. Network Flow perspective:Each stage has capacity2​k−12k\{\-\}1; total max flow acrossnnstages =n2n^\{2\}\.Final Answer: <n2\>17\. Category\-Theoretic perspective:Morphisms betweennnobjects mapped to odd indices; functor sum yieldsn2n^\{2\}\.Final Answer: <n2\>18\. Dimensional Analysis perspective:Odd steps interpreted as time growth; scaling law yields square relationship\.Final Answer: <n2\>19\. Stochastic Process perspective:Random walk with odd\-step increments; mean position afternnsteps isn2n^\{2\}\.Final Answer: <n2\>20\. Differential Geometry perspective:Geodesic arc\-length built from discrete odd steps; cumulative length =n2n^\{2\}\.Final Answer: <n2\>

## Appendix BTraining of the Selector

Algorithm 2REINFORCE Training of the Selector \(with Fixed Reasoner\)0:Training set

\{\(q,a∗,choices\)\}\\\{\(q,a^\{\*\},\\text\{choices\}\)\\\}, Reasoner

ℛ\\mathcal\{R\}, perspective set

𝒮\\mathcal\{S\}, penalty coefficient

λ\\lambda
1:foreach mini\-batch of

BBquestionsdo

2:foreach question

qqin the batchdo

3:

pi←Selector​\(q,si\)p\_\{i\}\\leftarrow\\textsc\{Selector\}\(q,s\_\{i\}\)for all

si∈𝒮s\_\{i\}\\in\\mathcal\{S\}
4:Sample binary mask

mi∼Bernoulli​\(pi\)m\_\{i\}\\sim\\text\{Bernoulli\}\(p\_\{i\}\)
5:Ensure at least one perspective is selected

6:endfor

7:Construct reasoning prompts from selected perspectives

8:Generate answers

aia\_\{i\}for each

\(q,si\)\(q,s\_\{i\}\)using fixed Reasoner

ℛ\\mathcal\{R\}
9:Aggregate answers per question via majority voting

10:foreach questiondo

11:Compute reward:

r=𝕀\[a^=a∗\]−λ⋅k\|𝒮\|r=\\mathbb\{I\}\[\\hat\{a\}=a^\{\*\}\]\-\\lambda\\cdot\\frac\{k\}\{\|\\mathcal\{S\}\|\}where

kkis the number of selected perspectives

12:endfor

13:Compute REINFORCE loss:

ℒ=−1B∑j=1Brj⋅logPr\(m\(j\)∣q\(j\)\)\\mathcal\{L\}=\-\\frac\{1\}\{B\}\\sum\_\{j=1\}^\{B\}r\_\{j\}\\cdot\\log\\Pr\(m^\{\(j\)\}\\mid q^\{\(j\)\}\)
14:Update selector parameters via gradient descent

15:endfor

## Appendix CPrompt Templates Used in Experiments

Chain\-of\-Thought Prompt TemplateDescription:This prompt guides the model to solve math problems step\-by\-step before stating the final answer\.Prompt Format:``` Solve the following math problem step by step. Explain each step clearly before giving the final answer. Question: <question_text> Final Answer: ```

Multi\-perspective Reasoning Prompt TemplateDescription:This prompt guides the reasoner model to solve the problem using a specified reasoning approach that selected by selector model and to explicitly state its confidence\. Later steps include prior answers and confidence for refinement\.First perspective Format:``` Question: <question_text> Approach: <perspective_name> Solve step by step, then output exactly two lines: 1) Answer: <final numeric answer> 2) Confidence: <0-100%> ``` Subsequent perspective Format:``` Question: <question_text> Previous approach: <perspective_name> Previous Answer: <answer_from_previous_step> Previous Confidence: <confidence_from_previous_step> Now apply approach: <current_perspective_name> to refine or confirm. Again output exactly two lines: 1) Answer: <final numeric answer> 2) Confidence: <0-100%> ```

Diverse Prompting Strategies TemplateDescription:This template generates multiple diverse prompts for a single question by invoking different reasoning strategies \(e\.g\., analogy, logic, inversion\)\. Used in sampling\-based prompting or ensembling\. All these prompts are mentioned in our baseline\([Lau et al\., 2024](https://arxiv.org/html/2609.21554#bib.bib2)\)\.Prompt Variations:1\) \*\*Break Down the Problem\*\*: Divide the question into smaller, manageable parts and tackle each part individually before synthesizing the overall answer\. \[QUESTION\]2\) \*\*Apply Mathematical Logic\*\*: Use mathematical principles and logic to solve the problem, even if it’s not a math question\. \[QUESTION\]3\) \*\*Use Analogies\*\*: Relate the question to a familiar concept or situation to better understand and solve it\. \[QUESTION\]4\) \*\*Consider the Opposite\*\*: Think about what the answer would be if the opposite were true, to gain a different perspective\. \[QUESTION\]5\) \*\*Consider Cause and Effect\*\*: Identify potential causes and their effects to understand the question better\. \[QUESTION\]

Gemini Answer Equivalence Judge TemplateDescription:This template is used to verify whether a model’s predicted answer matches the ground truth using a strict Gemini\-based equivalence check\. The model is instructed to respond with exactlyTrueorFalse\.Prompt Format:[⬇](data:text/plain;base64,WW91IGFyZSBhbiBhbnN3ZXIgY2hlY2tlci4gUmVzcG9uZCB3aXRoIGV4YWN0bHkgVHJ1ZSBpZiB0aGUgcHJlZGljdGVkCmFuc3dlciBtYXRjaGVzIHRoZSBncm91bmQgdHJ1dGgsIG9yIEZhbHNlIG90aGVyd2lzZS4KCkdyb3VuZCB0cnV0aDogPGdyb3VuZF90cnV0aF9hbnN3ZXI+ClByZWRpY3RlZCAgICA6IDxwcmVkaWN0ZWRfYW5zd2VyPgpFcXVpdmFsZW50Pw==)Youareananswerchecker\.RespondwithexactlyTrueifthepredictedanswermatchesthegroundtruth,orFalseotherwise\.Groundtruth:<ground\_truth\_answer\>Predicted:<predicted\_answer\>Equivalent?System Instruction:[⬇](data:text/plain;base64,QW5zd2VyIG9ubHkgVHJ1ZSBvciBGYWxzZS4=)AnsweronlyTrueorFalse\.

## Appendix DBroader Impacts

Potential Benefits\.By selectively invoking domain\-specific “reasoning perspectives,” MIRAGE can turn a mid\-sized LLM such as Qwen2\.5\-7B into a stronger solver for STEM and logical problems without any additional fine\-tuning\. This may \(i\) lower the compute barrier for building intelligent tutoring systems that give explicit, step\-by\-step solutions; \(ii\) assist researchers who need rapid but transparent first\-pass proofs or derivations; and \(iii\) improve accessibility for learners in low\-resource regions by reducing the number of expensive LLM calls required for high accuracy\.

Societal Risks\.

- •*Misuse for persuasive or deceptive reasoning\.*The same multi\-view strategy that helps derive correct answers can be steered toward generating convincing but false arguments\. Careful prompt\-level safeguards and usage policies are needed, especially for domains like finance or politics\.
- •*Academic integrity\.*MIRAGE lowers the cost of automated problem\-solving on benchmarks that closely resemble homework and exam questions\. Institutions should pair such tools with honor\-code education and detection systems\.

Mitigations\.We expose the confidence score for every perspective and allow users to inspect intermediate rationales, which makes it easier to audit errors\. We will publish the full training code, random seeds, and hyper\-parameters in camera\-ready version to encourage third\-party stress tests\.

Environmental Considerations\.MIRAGE significantly reduces inference cost by selectively invoking only a small subset of reasoning perspectives per query\. Compared to ensemble\-style reasoning methods that query all available paths \(e\.g\., Full\-20 or Self\-Consistency with largenn\), our method requires up to 5×\\timesfewer model calls while maintaining comparable or better accuracy\. This efficiency translates to lower carbon emissions and compute requirements, making MIRAGE especially practical for deployment in resource\-constrained or environmentally conscious settings\.

## Appendix EAblation Study

##### Effect of Dynamic perspective Selection\.

To assess the importance of our dynamic selection mechanism, we compare our full approach against two baseline conditions: \(a\)Random perspective Selection, where reasoning perspectives are selected randomly without Selector guidance; and \(b\)Fixed\-Order Selection, where the reasoning perspectives are always queried in a predetermined, fixed order\. Results indicate substantial performance degradation in both baselines, demonstrating that dynamic, context\-aware selection of reasoning perspectives significantly contributes to overall effectiveness\.

##### Impact of Multi\-perspective Aggregation\.

We evaluate the contribution of our aggregation mechanism by comparing our approach against a variant where aggregation is removed entirely—relying exclusively on the first high\-confidence single\-perspective solution\. Additionally, we test a simpler aggregation strategy \(simple majority voting\)\. The results clearly show that our aggregation strategy significantly boosts performance, especially on challenging datasets like MATH500 and MMLU\-Pro, underscoring the robustness gained from synthesizing insights across multiple conceptual perspectives\.

##### Sensitivity to Number of Reasoning perspectives\.

To investigate sensitivity to the number of reasoning perspectives, we progressively vary the maximum allowed number of queried perspectives \(kk\)\. We systematically analyze model performance as a function of this maximum number\. Results reveal a performance\-complexity trade\-off: while using more reasoning perspectives typically yields accuracy improvements, substantial gains are achieved even with very few queried perspectives \(e\.g\.,k≤3k\\leq 3\)\. This suggests that our Selector effectively prioritizes highly relevant reasoning perspectives early, efficiently balancing accuracy and computational cost\.

Similar Articles

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models

arXiv cs.AI

MIRAGE is a framework for mobile GUI agents that replaces verbose chain-of-thought reasoning with compact continuous latent representations, incorporating a generative world model perspective to predict future screen states before acting. On AndroidWorld and AndroidControl benchmarks, it achieves competitive or superior performance while reducing generated tokens by over 75%.

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Hugging Face Daily Papers

We introduce MIPI (Monotonic Inference Policy Improvement) and its instantiation MIPU, a two-step RL framework for LLMs that addresses the training-inference mismatch by explicitly aligning optimization with inference-policy improvement. Under FP8-quantized rollout, MIPU achieves improved reasoning performance and training stability across Qwen3-1.7B and Qwen3-4B models.