Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference
Summary
This paper introduces Alien Abduction, an interactive game to probe how LLMs acquire evidence, update hypotheses, and decide when to stop during abductive reasoning. It finds that models perform better with upfront evidence and with oracle-provided examples than with self-selected queries, revealing deficiencies in active information acquisition.
View Cached Full Text
Cached at: 08/05/26, 07:45 AM
# Don’t Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference
Source: [https://arxiv.org/html/2608.03388](https://arxiv.org/html/2608.03388)
First Author Affiliation / Address line 1 Affiliation / Address line 2 Affiliation / Address line 3 email@domain &Second Author Affiliation / Address line 1 Affiliation / Address line 2 Affiliation / Address line 3 email@domain Shahrukh Mohiuddin1, Chalamalasetti Kranti111footnotemark:1, Sherzod Hakimov1, David Schlangen1,2 1Computational Linguistics, Department of Linguistics University of Potsdam, Germany 2German Research Center for Artificial Intelligence \(DFKI\), Berlin, Germany \{shahrukh\.mohiuddin, kranti\.chalamalasetti, sherzod\.hakimov, david\.schlangen\}@uni\-potsdam\.de
###### Abstract
Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available\. While large language models \(LLMs\) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop\. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes\. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle\. Across models, providing evidence upfront leads to higher success rates than distributing it across turns\. In multi\-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging\. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected\. These findings suggest that models may form hypotheses that fit self\-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop\.
Don’t Let Me Ask for It: LLMs Show Deficiencies in Active Multi\-Turn Information Acquisition for Abductive Inference
Shahrukh Mohiuddin1††thanks:Both authors equally contributed\., Chalamalasetti Kranti111footnotemark:1, Sherzod Hakimov1, David Schlangen1,21Computational Linguistics, Department of LinguisticsUniversity of Potsdam, Germany2German Research Center for Artificial Intelligence \(DFKI\), Berlin, Germany\{shahrukh\.mohiuddin, kranti\.chalamalasetti, sherzod\.hakimov, david\.schlangen\}@uni\-potsdam\.de
## 1Introduction
One fact leads to another—so we continue\. Does the next fit in with that? A merveille\! Good\! We can proceed\. This next little fact—no\! Ah, that is curious\! There is something missing—a link in the chain that is not there\. We examine\. We search\. And that little curious fact, that possibly paltry little detail that will not tally, we put it here\! – Agatha Christie, in The Mysterious Affair at Styles111[https://www\.gutenberg\.org/cache/epub/863/pg863\-images\.html\#chap04](https://www.gutenberg.org/cache/epub/863/pg863-images.html#chap04)\.
Figure 1:Overview ofAlien Abduction\. The Game Master hides a target Python functionffand interacts with the LLM through a black\-box protocol\. In the Active\-Output mode shown here, the model proposes test inputs, receives the corresponding outputs, and eventually submits a final hypothesis as Python code\. The submitted hypothesis is then evaluated after submission against the hidden target function on held\-out test cases\.The ability to recognize an underlying rule from specific observations is a hallmark of intelligence: it has long been studied in cognitive science\(Wason,[1960](https://arxiv.org/html/2608.03388#bib.bib35); Bruner,[2017](https://arxiv.org/html/2608.03388#bib.bib36)\), is a standard component of human IQ tests and is central to scientific discovery\. In accounts of human inquiry dating back to Peirce\(Peirce,[1934](https://arxiv.org/html/2608.03388#bib.bib26)\), such discovery is not a single inference step but a cycle: forming a candidate explanation from observations \(abduction\), deriving what that explanation predicts \(deduction\), and revising it against new evidence \(induction\)\(Heet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib6); Yinet al\.,[2026](https://arxiv.org/html/2608.03388#bib.bib8)\)\. LLMs now perform strongly on many reasoning benchmarks\(Lewkowyczet al\.,[2022](https://arxiv.org/html/2608.03388#bib.bib32); Team,[2024](https://arxiv.org/html/2608.03388#bib.bib33); Zhanget al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib34)\), yet these benchmarks typically present all information upfront and evaluate a single, non\-interactive answer, often on data that may have appeared in their training corpora\(Balloccuet al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib29); Orenet al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib30); Chenget al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib31)\)\. They therefore cannot tell us whether a model can: decide what evidence to gather, form an explanation of it, and test that explanation before committing\. This capability matters in practice, as LLMs are increasingly deployed as agents that must handle unfamiliar environments, tools, and APIs through interaction rather than specifications\.
Several lines of work address parts of this problem\. Interactive program synthesis benchmarks let agents query a hidden function and refine candidate programs\(Weiet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib18); Leeet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib19)\); inductive reasoning benchmarks study rule discovery from provided examples or active queries\(Honovichet al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib4); Sunet al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib5); Yanet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib14); Chari and Pattanaik,[2026](https://arxiv.org/html/2608.03388#bib.bib10)\), while code benchmarks evaluate generation from explicit specifications or reasoning over visible programs\(Chen and others,[2021](https://arxiv.org/html/2608.03388#bib.bib11); Austinet al\.,[2021](https://arxiv.org/html/2608.03388#bib.bib12); Guet al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib17); Xuet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib15)\)\. Black\-box reasoning environments extend hidden\-rule discovery to broader domains\(Heet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib6); Yinet al\.,[2026](https://arxiv.org/html/2608.03388#bib.bib8); Chenet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib24)\)
However, these settings do not jointly vary who selects the evidence and what form the feedback takes over\. Existing comparisons of active and passive evidence collection use exact outputs only\(Genget al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib22)\), leaving the role of binary verdicts and disconfirming evidence less clear\. Moreover, evaluations focus on final\-answer correctness without testing whether a model’s intermediate hypotheses remain consistent with the evidence it has collected\. As a result, the effects of self\-directed exploration, negative evidence, and hypothesis grounding remain underexplored\.
We introduceAlien Abduction, a multi\-turn game, in which the model must reconstruct a hidden Python function from only its signature and a limited turn budget; submissions are executed in a sandbox against held\-out test cases \(Figure[1](https://arxiv.org/html/2608.03388#S1.F1)\)\. Four interactive game modes vary who controls the evidence \(active vs\. passive\) and its form \(exact outputs vs\. membership verdicts on proposed input–output pairs\); two single\-turn modes serve as non\-interactive baselines\. This design measures the effect of self\-directed exploration, tests whether models seek disconfirming rather than only confirming evidence, and scores submissions for consistency with the observed evidence\. Targets span five domains: numbers, number pairs, strings, lists, and Boolean logic\.
Our contributions are: \(a\) A benchmark game for black\-box function induction with six interaction modes varying evidence control \(active vs\. passive\) and evidence type \(exact outputs vs\. membership verdicts\); \(b\) A suite of 50 automatically generated and validated target functions across five domains; \(c\) A sandboxed evaluation protocol with task success, turn\-budget use and turn\-level hypothesis analysis; and \(d\) An empirical study across commercial and open\-weight models\.
## 2Related Work
#### Code generation and code understanding\.
HumanEval\(Chen and others,[2021](https://arxiv.org/html/2608.03388#bib.bib11)\)and MBPP\(Austinet al\.,[2021](https://arxiv.org/html/2608.03388#bib.bib12)\)assess whether models can produce Python functions from natural\-language descriptions, validated through unit tests; later work explored architectures and training for program synthesis\(Nijkampet al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib13); Zhenget al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib16)\)\. By contrast, CRUXEval\(Guet al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib17)\)and its multilingual extension CRUXEval\-X\(Xuet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib15)\)test reasoning about visible programs via input and output prediction\. In both families the task is fully specified: the behavior is described or the code is shown\. In Alien Abduction, neither is available; the behavior must be inferred before code can be written\.
#### Inductive reasoning from examples\.
Instruction Induction\(Honovichet al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib4)\)asks models to describe a hidden rule from input–output examples in natural language, ItD\(Sunet al\.,[2024](https://arxiv.org/html/2608.03388#bib.bib5)\)improves inductive ability by leveraging deduction on list\-function tasks, and MIR\-Bench\(Yanet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib14)\)scales the setting to many\-shot pattern recognition over hundreds of examples\. In all of these, the evidence set is fixed\. Moving beyond static observation,Chari and Pattanaik \([2026](https://arxiv.org/html/2608.03388#bib.bib10)\)study active concept learning with predefined query\-selection policies \(expected\-information\-gain vs\. confirmation\-style positive tests\); the querying strategy is still not chosen by the model\.
#### Interactive black\-box reasoning environments\.
Genget al\.\([2025](https://arxiv.org/html/2608.03388#bib.bib22)\)compare passive observation with active intervention when LLMs reverse\-engineer programs, formal languages, and equations\. In RULEARN\(Heet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib6)\), agents learn hidden rules through cycles of abduction, deduction, and induction; ORACLE\(Yinet al\.,[2026](https://arxiv.org/html/2608.03388#bib.bib8)\)extends black\-box interaction to code, circuits, physical systems, encryption, and games\. AR\-Bench tests whether models can ask useful questions under incomplete information\(Zhouet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib23)\), and PHYSGYM evaluates experiment design under controlled levels of prior knowledge\(Chenet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib24)\)\. These benchmarks show that models often struggle to choose informative actions and to adapt after feedback; none, however, scores the final hypothesis for consistency with the evidence collected during the episode, or penalizes inefficient use of the interaction budget\.
#### Interactive program synthesis\.
Closest to our setting, CodeARC\(Weiet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib18)\)evaluates agents that query a hidden target function and refine candidate Python implementations against a differential testing oracle, while SYNTRA\(Leeet al\.,[2025](https://arxiv.org/html/2608.03388#bib.bib19)\)selects informative test inputs to eliminate competing program hypotheses;Suranaet al\.\([2026](https://arxiv.org/html/2608.03388#bib.bib20)\)add human feedback and verification\. A complementary line uses black\-box access for explanation rather than reconstruction\(Cífka and Liutkus,[2023](https://arxiv.org/html/2608.03388#bib.bib7); Singhet al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib25)\)\.
Alien Abduction combines these threads in a single controlled setting: varying evidence control and feedback form over identical hidden functions allows the effect of self\-directed exploration to be estimated; membership\-verdict feedback tests whether models seek disconfirming evidence; and an internal\-consistency check separates inferring a wrong rule from ignoring one’s own evidence\.
## 3Methodology
### 3\.1Task Formulation
Letf:X→Yf:X\\rightarrow Ydenote a hidden target function with graph
Ff=\{\(x,y\)∈X×Y∣y=f\(x\)\}\.F\_\{f\}=\\\{\(x,y\)\\in X\\times Y\\mid y=f\(x\)\\\}\.The player never observes the source code offf\. Instead, a Game Master with oracle access toffmediates all interaction: it reveals evidence aboutFfF\_\{f\}according to a fixed protocol and verifies proposed solutions\. SinceXXmay be multivariate, an inputxxconsists of one or more values\.
An episode is parameterized by the function signature \(argument and return types\), an interaction mode, and a turn budgetTT\. In each turnt≤Tt\\leq T, the player performs exactly one action: it either requests evidence, as permitted by the mode, or submits a candidate implementationf^\\hat\{f\}as Python code via theSOLVEaction, which ends the episode\. Alongside each action, the player reports its current hypothesis, a state indicating whether it is probing, confirming, or uncertain, its rationale for selecting the query, and a confidence score\. These fields are recorded for analysis but are not processed by the Game Master and do not affect its response\. An episode counts as solved if and only iff^\\hat\{f\}agrees withffon a set of held\-out test inputs; this set is constructed per instance and never revealed to the player, and agreement is checked by executingf^\\hat\{f\}in an ephemeral container \(Section[4\.4](https://arxiv.org/html/2608.03388#S4.SS4)\)\. An episode that exhausts its budget without a submission counts as failed\.
Figure[1](https://arxiv.org/html/2608.03388#S1.F1)shows an episode in theActive\-Outputmode: given the signature\(x: int\) \-\> int, the player probes the black box with inputs such as−3\-3and0, observes the outputs33and0, and submits an implementation of the absolute\-value function\. The challenge is therefore not code generation from a specification, but inference of the behavioral rule from limited evidence\. Each episode thereby instantiates the reasoning cycle from the introduction: the player forms a candidate rule from its observations \(abduction\), derives which probe would test it \(deduction\), and revises the rule as new evidence arrives \(induction\)\.
### 3\.2Game Modes
Game modeControlFeedbackPlayer actionExample exchangeActive\-Output \(AO\)activeoutputTEST: <input\>TEST: \-8→\\rightarrowOUTPUT: 8Active\-Verdict \(AV\)activemembershipTEST: <input\>,<output\>TEST: 2, 2→\\rightarrowOUTPUT: TrueTEST: 2, 1→\\rightarrowOUTPUT: FalsePassive\-Output \(PO\)passiveoutputNEXTNEXT→\\rightarrowOUTPUT: \(\-2, 2\)Passive\-Verdict \(PV\)passivemembershipNEXTNEXT→\\rightarrowOUTPUT: \(\(\-2, 2\), True\)NEXT→\\rightarrowOUTPUT: \(\(10, 9\), False\)Single\-Turn\-Output \(STO\)single\-turnoutputSOLVEonly10 examples shown upfrontSingle\-Turn\-Verdict \(STV\)single\-turnmembershipSOLVEonly10 labeled pairs shown upfrontTable 1:The six interaction modes of Alien Abduction, factorized by evidence control \(who selects the probes\) and feedback form \(exact outputs vs\. binary membership verdicts\)\. In every mode the player can end the episode at any turn withSOLVE:followed by a candidate Python implementation; the single\-turn modes permit only this action after a fixed batch of ten evidence items is shown\. All example exchanges assume the hidden target functionabsolute\_value\(x: int\) \-\> int\.The six modes, summarized in Table[1](https://arxiv.org/html/2608.03388#S3.T1), vary two properties of the evidence aboutFfF\_\{f\}: who controls its selection \(active vs\. passive\) and what form the feedback takes \(exact outputs vs\. binary membership verdicts\)\.
In the two active modes, the player selects its own probes\.Active\-Output\(AO\) returns the exact outputf\(x\)f\(x\)for a queried inputxx, yielding positive evidence only\.Active\-Verdict\(AV\) instead returns whether a proposed pair\(x,y\)\(x,y\)lies inFfF\_\{f\}: the player trades exact outputs for the ability to test, and potentially falsify, its own hypotheses, since both positive and negative evidence can result\.
In the two passive modes, the Game Master controls the evidence stream\.Passive\-Output\(PO\) reveals one valid pair fromFfF\_\{f\}per request, reducing the problem to induction from externally provided positive evidence;Passive\-Verdict\(PV\) reveals candidate pairs with true/false labels, providing contrastive evidence without player control\.
The two single\-turn modes,Single\-Turn\-Output\(STO\) andSingle\-Turn\-Verdict\(STV\), present a fixed batch of evidence upfront and permit only theSOLVEaction\. Since their evidence is still passively provided, they are the natural baselines for the passive modes: STO vs\. PO isolates the effect of sequential interaction, while PO vs\. AO isolates control over probe selection\. Jointly, the six modes disentangle strategic probe selection, induction from positive evidence, induction from mixed positive and negative evidence, and hypothesis formation under a fixed interaction budget\.
## 4Experiment Setup
### 4\.1Target Functions and Test Cases
DomainSignature\#FnsNumberint→\\rightarrowint10Number Pairs\(int, int\)→\\rightarrowint10Stringstr→\\rightarrowstr10ListList\[int\]→\\rightarrowint10Logic\(bool, bool\)→\\rightarrowbool10Table 2:The five domains of Alien Abduction, their function signatures, and the number of sampled target functions per domain\. The complete list of target functions is given in Table[3](https://arxiv.org/html/2608.03388#A1.T3)in Appendix[A](https://arxiv.org/html/2608.03388#A1)\.The benchmark is built on five domains, each defined by a fixed function signature \(Table[2](https://arxiv.org/html/2608.03388#S4.T2)\): numbers, number pairs, strings, lists, and Boolean logic\. We deliberately restrict the targets to primitive functions: short, pure, deterministic transformations that use only the standard library\. Such functions keep the hypothesis space tractable, ensure that failures reflect evidence\-gathering and rule inference rather than programming difficulty, and still admit many distinct behaviors per signature\.
Target functions are generated rather than hand\-picked\. For each signature, we prompt GPT5\.4 to produce 100 candidate functions that must match the signature exactly, be pure and deterministic, use no imports, and remain short\. Each candidate is automatically validated: its code must parse, define exactly one function with the expected parameters, and execute without error on sample inputs\. Because the targets are drawn from the generator model’s own output distribution, they are by construction functions that an LLM can express, so failure to identify them cannot be attributed to complex behavior\. From the validated pool we randomly sample \(with a fixed seed\) 10 functions per signature, yielding 50 targets\. The same 50 targets are used across the modes, so differences between modes cannot be confounded by target difficulty\.
Each target is paired with 100 test cases\. Inputs are drawn from type\-aware pools that mix edge cases \(e\.g\., zero, negatives, empty strings and lists\) with random values, and the output for every input is computed by executing the target function\. These test cases serve two roles: they provide the evidence revealed in the passive and single\-turn modes, and they act as the held\-out suite against which every submitted hypothesis is verified\.
### 4\.2Models
We evaluate four models:GPT5\.4222[https://openai\.com/index/introducing\-gpt\-5\-4/](https://openai.com/index/introducing-gpt-5-4/),GPT5\.4\-mini333[https://openai\.com/index/introducing\-gpt\-5\-4\-mini\-and\-nano/](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/), andMistral\-Large\-3\(AI,[2026](https://arxiv.org/html/2608.03388#bib.bib28)\)as commercial systems, andQwen3\.6\-35B\-A3B\(Qwen Team,[2026](https://arxiv.org/html/2608.03388#bib.bib27)\)as an open\-weight model\. Since GPT5\.4 is also the generator of the target functions, its results additionally indicate how well a model recovers functions drawn from its own distribution\.
### 4\.3Metrics
We report three metrics\.Success ratemeasures correctness: an episode scores 1 if the submitted hypothesis passes all held\-out test cases of the hidden function, and 0 otherwise\. Averaged over episodes and reported on a 0–1 scale, it is aggregated per mode, per domain, and per model\.
Turn Budget Use\(TBU\) measures the proportion of the available interaction budget consumed before an episode ends\. For the interactive modes,
TBU=\(n−1T\)\\text\{TBU\}=\(\\frac\{n\-1\}\{T\}\)whereTTis the maximum turn budget andnnis the number of turns used\. A score of11indicates that the episode uses the entire turn budget, while lower scores indicate earlier termination\. For the single\-turn modes, the turn budget is fixed at one and is therefore fully used by definition\.
We measurehypothesis retrodiction accuracy\(HRA\) to assess the consistency of a model’s current hypothesis with the accumulated evidence\. At turntt, we compute
HRAt=NtmatchedNt\\mathrm\{HRA\}\_\{t\}=\\frac\{N\_\{t\}^\{\\mathrm\{matched\}\}\}\{N\_\{t\}\}
whereNtmatchedN\_\{t\}^\{\\mathrm\{matched\}\}is the number of observed evidences reproduced correctly by the hypothesis andNtN\_\{t\}is the total number observed\. A score of11indicates consistency with all available evidence\. An agent that reliably tracks its belief state should maintain a score of11, since its reported hypothesis should not contradict evidence received so far\. We report HRA across turns and at the final turn to examine how models revise and validate their hypotheses before committing\.
TBU and HRA provide complementary views of the interaction\. TBU captures how much of the available turn budget is used, while HRA captures whether the reported hypothesis remains consistent with the accumulated evidence\.
### 4\.4Implementation
Alien Abduction444The source code and test instances will be released upon the paper’s acceptanceis implemented in the clembench framework\(Chalamalasettiet al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib9)\), which provides the Game Master loop, prompt handling, and scoring infrastructure\. Each interactive episode has a turn budget ofT=15T=15; the single\-turn modes allow a singleSOLVEaction\. In total, each model plays the 50 targets in all six modes, i\.e\., 300 episodes per model\. Models are accessed through their respective APIs with default decoding parameters\. At each turn, we log theselected action,query and response,current hypothesis,interaction state,query rationale,confidence score, and whether the modelresponse was parsed successfully\. These traces are used to compute the turn\-level metrics and support the qualitative analysis\. Every submitted solution is executed in a native Python 3 environment inside an ephemeral sandbox container that is destroyed after the run; only the test results are returned to the Game Master, which then decides whether the episode is won or lost\.
Figure 2:Success rates across interaction modes for each model, with 95% confidence intervals\. Dashed horizontal lines indicate the overall success rate for the model\. A missing bar indicates that zero success for that mode\.Figure 3:TBU scores across models and interaction modes\. Higher scores indicate more of the turn budget was used; scores below 0\.67 correspond to fewer than 10 turns\. Missing bars indicate no instances for that outcome\.
## 5Results
### 5\.1Analysing model capabilities
Can LLMs reconstruct a hidden Python function from input\-output evidence?Figure[2](https://arxiv.org/html/2608.03388#S4.F2)presents the success rates across the evaluated models and interaction modes555Success rates across domains are available in Figures[14](https://arxiv.org/html/2608.03388#A5.F14)and[16](https://arxiv.org/html/2608.03388#A5.F16)in the Appendix[C](https://arxiv.org/html/2608.03388#A3)\.\. The overall success rates range from0\.030\.03for Qwen3\.6\-35B, through0\.200\.20for Mistral\-Large\-3, to0\.640\.64for GPT5\.4\. The low success rates of Qwen3\.6\-35B and Mistral\-Large\-3 are mainly because both models continue probing until the final turn, exhausting all1515turns without converging to a solution\. Moreover, Qwen3\.6\-35B struggles to follow the intended response format, with parser errors occurring in40\.64%40\.64\\%of all instances \(see Table[5](https://arxiv.org/html/2608.03388#A2.T5)in the Appendix\)\. Together, these results suggest that failures arise from difficulties with induction and format adherence\.
We then analyse how each interaction mode contributes to the model\-overall success rates\. We use the single\-turn modes as baselines for comparison with the corresponding multi\-turn modes\. For most models, across both feedback types, success rates follow the order single\-turn, passive, and active, indicating that models perform better when evidence is provided upfront than when they must acquire it through interaction\. Between the two single\-turn modes, the single\-turn\-output mode achieves higher success rates\. Membership feedback is more challenging because it only indicates whether a proposed output is correct, whereas output feedback reveals the corresponding output and is the direct evidence about the hidden function\. This challenge is greater in the active\-verdict mode, where the model must also select the queries\.
Overall, the results show that the models are capable of performing the task, but their performance remains far from saturated\.
Figure 4:Hypothesis retrodiction accuracy across turns\. Dotted lines show the median turn count for each model\.Figure 5:Hypothesis retrodiction accuracy at the final turn across models and interaction modes\.
### 5\.2Evidence\-budget management
How does distributing evidence across turns, with model\-controlled stopping, affect success and use of the available turn budget?Figure[3](https://arxiv.org/html/2608.03388#S4.F3)presents the TBU scores for each model and interaction mode\. Qwen3\.6\-35B, Mistral\-Large\-3, and GPT5\.4\-mini have a score of zero in at least one interaction mode\. Across most of the models, successful instances have lower TBU scores than failed instances\. In active\-output and passive\-output, GPT5\.4\-mini has lower TBU scores than the other models, at0\.390\.39and0\.440\.44, respectively\. Similarly, Mistral\-Large\-3 has a lower TBU score than the other models in active\-verdict, at0\.470\.47\. However, TBU measures how much of the available turn budgets used before an episode ends and should therefore be considered together with the overall success rates\.
The higher TBU scores for failed instances suggest that these instances are more challenging and lead to longer interactions\. Qwen3\.6\-35B and to an extent Mistral\-Large\-3 nearly exhausts the turn budget, whereas the other models use fewer than 10 turns in most interaction modes, as evident from TBU scores below0\.670\.67\.
This pattern shows that models often make their guesses before using the full turn budget\. However, earlier commitment does not translate into higher success rates, as all models achieve lower success rates\. To examine whether this difference is related to model\-controlled stopping, we compare single\-turn\-output and passive\-output\. GPT5\.4 and GPT5\.4\-mini achieve similar success rates in single\-turn\-output, at0\.820\.82and0\.740\.74, respectively, but differ in passive\-output, at0\.780\.78and0\.400\.40\(Figure[2](https://arxiv.org/html/2608.03388#S4.F2)\)\. The larger gap in passive\-output suggests that model\-controlled evidence acquisition and stopping affect the two models differently\. Overall, some models stop before using the available evidence and fail, whereas others exhaust the turn budget without converging to a solution\.
#### Hypothesis Retrodiction Accuracy
Figure[4](https://arxiv.org/html/2608.03388#S5.F4)shows how hypothesis retrodiction accuracy scores vary across settings\. In active\-output, most models begin with high scores, decline over the intermediate turns, and then increase again toward the final turns\. In the other modes, the scores generally decrease as the interaction progresses\. This suggests that hypotheses are matched more accurately during the initial turns and weakens as additional evidence are available\. To examine whether this pattern persists at the point of commitment, Figure[5](https://arxiv.org/html/2608.03388#S5.F5)reports the HRA at the final turn\. For the successful instances of active interaction modes, the final\-turn HRA ranges from0\.850\.85to1\.001\.00across models, whereas for the other modes it ranges from0\.200\.20to0\.550\.55\. Overall, this suggests that models may struggle to validate their hypotheses against the accumulated evidence and may commit before observing enough evidence to refine them\. We next examine whether this pattern depends on how evidence is acquired\.
Figure 6:Negative evidence impact on task success for GPT5\.4 in Numbers domain\.
### 5\.3Query selection
Does self\-selection affect stopping behaviour and overall success?We examine this by comparing the active\-output and passive\-output modes\. Both are multi\-turn, but active\-output requires the model to select each query, whereas in passive\-output the Game Master provides the next example\. As shown in Figure[3](https://arxiv.org/html/2608.03388#S4.F3), Qwen3\.6\-35B, Mistral\-Large\-3, and GPT5\.4\-mini have higher TBU scores in active\-output than in passive\-output for successful instances, while GPT5\.4 shows the opposite pattern\. For failed instances, only GPT5\.4\-mini has a lower TBU score in active\-output\. Despite these differences, most of the models achieve higher success rates in passive\-output\. At the final turn, hypothesis retrodiction accuracy \(see Figure[5](https://arxiv.org/html/2608.03388#S5.F5)\) is higher in active\-output for all models, but this accuracy is measured against examples selected by the model and may therefore reflect consistency with its current hypothesis rather than the ability of those examples to distinguish it from competing hypotheses\. In passive\-output, lower retrodiction accuracy and higher TBU scores suggest that models observe more Game Master\-provided examples that challenge their current hypotheses before committing\. Overall, self\-selection may lead models to stop earlier on examples that fit their current hypotheses, while passive presentation exposes them to more evidence and leads to higher success\.
Figure 7:7\.56% of failed instances of GPT5\.4\-mini’s end with a correct hypothesis but no submitted solution\.
## 6Qualitative Analysis
We analyse failed interactions to identify patterns across models and modes\. Figure[6](https://arxiv.org/html/2608.03388#S5.F6)shows how negative evidence affects task success for GPT5\.4\. Active\-verdict requires more turns than passive\-verdict, consistent with the higher TBU score of the latter discussed in Section[5\.2](https://arxiv.org/html/2608.03388#S5.SS2)\. Moreover, in active\-verdict, when the Game Master labels an input\-output pair asFalse, the model changes both its subsequent queries and hypotheses, suggesting that it eliminates some possibilities\. However, these revisions do not consistently lead to a hypothesis that fits the accumulated evidence, which may explain the longer interactions and lower success\. As shown in Figure[7](https://arxiv.org/html/2608.03388#S5.F7), some failures areunclaimed wins: the model reaches a correct hypothesis but continues querying until the turn budget is exhausted666Details across models are in Table[4](https://arxiv.org/html/2608.03388#A2.T4)in the Appendix\. This suggests that some models may not reliably assess when their hypotheses are sufficiently supported, leading either to premature commitment or failure to converge\. Figure[8](https://arxiv.org/html/2608.03388#S6.F8)compares the input coverage when GPT5\.4 selects queries with that of examples provided by the Game Master\. The active queries cover a narrower region, whereas the Game Master\-provided examples span a wider range\. This supports the analysis in Section[5\.3](https://arxiv.org/html/2608.03388#S5.SS3)that self\-selection can limit input coverage and reduce success in active\-output\. Together, these examples suggest that task success depends on query coverage, hypothesis validation, and how models use negative evidence\. Additional turn\-level traces showing the selected actions and corresponding outputs for each model, along with analyses of confidence scores, interaction states, and query rationales, are available in the Appendix[D](https://arxiv.org/html/2608.03388#A4)\.
Figure 8:Spread of input queries:These results are for GPT5\.4 model, for the two\_numbers category\. The narrower coverage of Active\-output queries may contribute to the higher failure rate\.
## 7Conclusion
We introduce Alien Abduction, a multi\-turn dialogue game for exploring the abductive reasoning capabilities of LLMs across interaction settings and models\. Our findings show that the main difficulty lies in determining how much evidence to collect and how to use it to revise hypotheses over time\. Models often commit before gathering enough evidence or continue querying without converging\. Therefore the probe reveals behaviours that are not visible in single\-turn evaluations and provides a way to examine how models manage evidence before reaching a final answer\. Overall, the results suggest that improving abductive reasoning requires better query selection, hypothesis validation, and stopping\.
## Limitations
Our study makes several design choices to support controlled comparisons across interaction modes\. These choices involve deliberate trade\-offs: they strengthen experimental control while limiting the range of settings examined\.
1\. Task scope\.Alien Abduction probe uses synthetic functions across five controlled domains\. This design isolates evidence acquisition, hypothesis revision, and stopping behaviour\. However, open\-ended abductive reasoning may involve ambiguity, noisy evidence, and background knowledge not represented in the current tasks\. Our findings therefore concern reasoning over controlled rule\-induction tasks, and future work can examine whether the observed behaviours extend to real\-world settings\.
2\. Target generation and model coverage\.Candidate target functions are generated using GPT5\.4, automatically validated, and randomly sampled from the resulting pool\. GPT5\.4 contributes only the initial candidates and does not select the final targets or evaluate model responses\. All target functions are checked to ensure that they run without errors and produce deterministic outputs, and every model is evaluated on the same targets and test cases\. These controls reduce the possibility of an advantage for GPT5\.4\. However, the resulting function distribution may still reflect patterns more commonly produced by that model\. In addition, the evaluated models do not cover the full range of available LLMs\. Future work could extend the framework by using targets produced by multiple generators and evaluating additional models and prompting settings\.
3\. Interaction settings\.The interactive modes use a fixed turn budget, and each mode specifies who selects the examples and what information is returned\. Keeping these settings constant enables direct comparisons between active, passive, and single\-turn interaction\. Different turn budgets or Game Master selection policies nevertheless affect query selection, stopping behaviour, and task success\. Examining these factors would provide a broader account of turn\-budget management\.
4\. Behavioural traces\.HRA is computed from the hypotheses explicitly reported by the models\. These reports provide a turn\-level view of whether the stated hypothesis remains consistent with the accumulated evidence, but they may not fully represent the model’s internal belief state\. Confidence scores, interaction states, and query rationales should similarly be interpreted as observable behavioural traces rather than faithful explanations of the underlying reasoning process\.
## References
- M\. AI \(2026\)Ministral 3\.CoRRabs/2601\.08584\.External Links:[Link](https://doi.org/10.48550/arXiv.2601.08584),[Document](https://dx.doi.org/10.48550/ARXIV.2601.08584),2601\.08584Cited by:[§4\.2](https://arxiv.org/html/2608.03388#S4.SS2.p1.1)\.
- J\. Austin, A\. Odena, M\. I\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. J\. Cai, M\. Terry, Q\. V\. Le, and C\. Sutton \(2021\)Program synthesis with large language models\.CoRRabs/2108\.07732\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2108.07732),[Link](https://arxiv.org/abs/2108.07732)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Balloccu, P\. Schmidtová, M\. Lango, and O\. Dusek \(2024\)Leak, cheat, repeat: data contamination and evaluation malpractices in closed\-source LLMs\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),Y\. Graham and M\. Purver \(Eds\.\),St\. Julian’s, Malta,pp\. 67–93\.External Links:[Link](https://aclanthology.org/2024.eacl-long.5/),[Document](https://dx.doi.org/10.18653/v1/2024.eacl-long.5)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, and et al \(2020\)Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6\-12, 2020, virtual,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\. Balcan, and H\. Lin \(Eds\.\),External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html)Cited by:[Appendix B](https://arxiv.org/html/2608.03388#A2.p1.1)\.
- J\. Bruner \(2017\)A study of thinking\.Routledge\.Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- K\. Chalamalasetti, J\. Götze, S\. Hakimov, B\. Madureira, P\. Sadler, and D\. Schlangen \(2023\)Clembench: using game play to evaluate chat\-optimized language models as conversational agents\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 11174–11219\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.689/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.689)Cited by:[§4\.4](https://arxiv.org/html/2608.03388#S4.SS4.p1.1)\.
- A\. Chari and N\. Pattanaik \(2026\)Wild guesses and mild guesses in active concept learning\.CoRRabs/2602\.06818\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.06818),[Link](https://arxiv.org/abs/2602.06818)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Chenet al\.\(2021\)Evaluating large language models trained on code\.CoRRabs/2107\.03374\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2107.03374),[Link](https://arxiv.org/abs/2107.03374)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Chen, P\. Piekos, M\. Ostaszewski, F\. Laakom, and J\. Schmidhuber \(2025\)PHYSGYM: benchmarking LLMs in interactive physics discovery with controlled priors\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Cheng, Y\. Chang, and Y\. Wu \(2025\)A survey on data contamination for large language models\.CoRRabs/2502\.14425\.External Links:[Link](https://doi.org/10.48550/arXiv.2502.14425),[Document](https://dx.doi.org/10.48550/ARXIV.2502.14425),2502\.14425Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- O\. Cífka and A\. Liutkus \(2023\)Black\-box language model explanation by context length probing\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),Toronto, Canada,pp\. 1067–1079\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-short.92),[Link](https://aclanthology.org/2023.acl-short.92/)Cited by:[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px4.p1.1)\.
- J\. Geng, H\. Chen, D\. Arumugam, and T\. L\. Griffiths \(2025\)Are large language models reliable AI scientists? assessing reverse\-engineering of black\-box systems\.CoRRabs/2505\.17968\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.17968),[Link](https://arxiv.org/abs/2505.17968)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p4.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Gu, B\. Roziere, H\. J\. Leather, A\. Solar\-Lezama, G\. Synnaeve, and S\. Wang \(2024\)CRUXEval: a benchmark for code reasoning, understanding and execution\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 16568–16621\.External Links:[Link](https://proceedings.mlr.press/v235/gu24c.html)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px1.p1.1)\.
- K\. He, M\. Zhang, S\. Yan, P\. Wu, and Z\. Z\. Chen \(2025\)IDEA: enhancing the rule learning ability of large language model agent through induction, deduction, and abduction\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 13563–13597\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.698),[Link](https://aclanthology.org/2025.findings-acl.698/)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1),[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px3.p1.1)\.
- O\. Honovich, U\. Shaham, S\. R\. Bowman, and O\. Levy \(2023\)Instruction induction: from few examples to natural language task descriptions\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Toronto, Canada,pp\. 1935–1952\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.108),[Link](https://aclanthology.org/2023.acl-long.108/)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Lee, J\. Koo, S\. Yoon, M\. Kim, H\. Koh, D\. Lee, and K\. Jung \(2025\)Program synthesis via test\-time transduction\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px4.p1.1)\.
- A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo, Y\. Wu, B\. Neyshabur, G\. Gur\-Ari, and V\. Misra \(2022\)Solving quantitative reasoning problems with language models\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 3843–3857\.External Links:[Document](https://dx.doi.org/10.52202/068431-0278),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/18abbeef8cfe9203fdf9053c9c4fe191-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- P\. Liu, W\. Yuan, J\. Fu, Z\. Jiang, H\. Hayashi, and G\. Neubig \(2023\)Pre\-train, prompt, and predict: A systematic survey of prompting methods in natural language processing\.ACM Comput\. Surv\.55\(9\),pp\. 195:1–195:35\.External Links:[Link](https://doi.org/10.1145/3560815),[Document](https://dx.doi.org/10.1145/3560815)Cited by:[Appendix B](https://arxiv.org/html/2608.03388#A2.p1.1)\.
- E\. Nijkamp, H\. Hayashi, C\. Xiong, S\. Savarese, and Y\. Zhou \(2023\)CodeGen2: lessons for training LLMs on programming and natural languages\.InThe Eleventh International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Oren, N\. Meister, N\. Chatterji, F\. Ladhak, and T\. Hashimoto \(2024\)Proving test set contamination in black\-box language models\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 16354–16372\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/46e624c244cff669223d488defd4e835-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- C\. S\. Peirce \(1934\)Collected papers of charles sanders peirce\.Vol\.5,Harvard University Press\.Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- Qwen Team \(2026\)Qwen3\.6\-35B\-A3B: agentic coding power, now open to all\.External Links:[Link](https://qwen.ai/blog?id=qwen3.6-35b-a3b)Cited by:[§4\.2](https://arxiv.org/html/2608.03388#S4.SS2.p1.1)\.
- C\. Singh, A\. R\. Hsu, R\. J\. Antonello, S\. Jain, A\. G\. Huth, B\. Yu, and J\. Gao \(2023\)Explaining black box text modules in natural language with language models\.CoRRabs/2305\.09863\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2305.09863),[Link](https://arxiv.org/abs/2305.09863)Cited by:[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px4.p1.1)\.
- W\. Sun, H\. Xu, X\. Yu, P\. Chen, S\. He, J\. Zhao, and K\. Liu \(2024\)ItD: large language models can teach themselves induction through deduction\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 2719–2731\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.150),[Link](https://aclanthology.org/2024.acl-long.150/)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Surana, A\. Srinivasan, and M\. Bain \(2026\)Engineering systems for data analysis using interactive structured inductive programming\.InAdvanced Information Systems Engineering: 38th International Conference, CAiSE 2026, Verona, Italy, June 8–12, 2026, Proceedings, Part I,Lecture Notes in Computer Science, Vol\.16558,pp\. 249–266\.External Links:[Document](https://dx.doi.org/10.1007/978-3-032-28110-4%5F14)Cited by:[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px4.p1.1)\.
- L\. Team \(2024\)The llama 3 herd of models\.CoRRabs/2407\.21783\.External Links:[Link](https://doi.org/10.48550/arXiv.2407.21783),[Document](https://dx.doi.org/10.48550/ARXIV.2407.21783),2407\.21783Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- P\. C\. Wason \(1960\)On the failure to eliminate hypotheses in a conceptual task\.Quarterly journal of experimental psychology12\(3\),pp\. 129–140\.Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- A\. Wei, T\. Suresh, J\. Cao, N\. Kannan, Y\. Wu, K\. Yan, T\. S\. F\. X\. Teixeira, K\. Wang, and A\. Aiken \(2025\)CodeARC: benchmarking reasoning capabilities of LLM agents for inductive program synthesis\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Q5pVZCrrKr)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px4.p1.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. H\. Chi, Q\. V\. Le, and D\. Zhou \(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 \- December 9, 2022,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper%5C_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html)Cited by:[Appendix B](https://arxiv.org/html/2608.03388#A2.p1.1)\.
- R\. Xu, J\. Cao, Y\. Lu, M\. Wen, H\. Lin, X\. Han, B\. He, S\. Cheung, and L\. Sun \(2025\)CRUXEVAL\-X: A benchmark for multilingual code reasoning, understanding and execution\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),pp\. 23762–23779\.External Links:[Link](https://doi.org/10.18653/v1/2025.acl-long.1158),[Document](https://dx.doi.org/10.18653/V1/2025.ACL-LONG.1158)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px1.p1.1)\.
- K\. Yan, Z\. Ling, K\. Liu, Y\. Yang, T\. Fan, L\. Shen, Z\. Du, and J\. Chen \(2025\)MIR\-bench: can your LLM recognize complicated patterns via many\-shot in\-context reasoning?\.InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2\-7, 2025 / Mexico City, Mexico, November 30 \- December 5, 2025,D\. Belgrave, C\. Zhang, L\. N\. Montoya, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, N\. Chen, I\. V\. M\. Ruíz, and A\. Loaiza\-Bonilla \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper%5C_files/paper/2025/hash/796076672b00f54fb01d05a2e5fde363-Abstract-Datasets%5C_and%5C_Benchmarks%5C_Track.html)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px2.p1.1)\.
- C\. Yin, T\. Wu, Y\. Shu, A\. Gu, Y\. Wang, J\. Shao, X\. Jiang, and P\. Li \(2026\)Investigating advanced reasoning of large language models via black\-box environment interaction\.InProceedings of the 43rd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.306\.Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1),[§1](https://arxiv.org/html/2608.03388#S1.p3.1),[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px3.p1.1)\.
- H\. Zhang, J\. Da, D\. Lee, V\. Robinson, C\. Wu, W\. Song, T\. Zhao, P\. Raja, C\. Zhuang, D\. Slack, Q\. Lyu, S\. Hendryx, R\. Kaplan, M\. Lunati, and S\. Yue \(2024\)A careful examination of large language model performance on grade school arithmetic\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 46819–46836\.External Links:[Document](https://dx.doi.org/10.52202/079017-1485),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/53384f2090c6a5cac952c598fd67992f-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§1](https://arxiv.org/html/2608.03388#S1.p2.1)\.
- Z\. Zheng, K\. Ning, Y\. Wang, J\. Zhang, D\. Zheng, M\. Ye, and J\. Chen \(2023\)A survey of large language models for code: evolution, benchmarking, and future trends\.CoRRabs/2311\.10372\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2311.10372),[Link](https://arxiv.org/abs/2311.10372)Cited by:[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Zhou, X\. Feng, Z\. Zhu, J\. Yao, S\. Koyejo, and B\. Han \(2025\)From passive to active reasoning: can large language models ask the right questions under incomplete information?\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267\.External Links:[Link](https://proceedings.mlr.press/v267/zhou25e.html)Cited by:[§2](https://arxiv.org/html/2608.03388#S2.SS0.SSS0.Px3.p1.1)\.
## Appendix ABenchmark Functions
Table[3](https://arxiv.org/html/2608.03388#A1.T3)shows the 50 target functions used for evaluation\. The functions are grouped by domain, and those within each domain share the same signature\. Each domain contains 10 functions, covering arithmetic operations, string manipulation, list aggregation, and Boolean logic across the full set\.
DomainSignatureTarget functionsNumber\(x: int\) \-\> intdifference\_from\_reversed\_absolute,smallest\_divisor\_ge\_two\_abs,distance\_from\_square\_of\_three,add\_seven\_if\_negative,max\_with\_negative\_self,times\_four\_mod\_nine,integer\_average\_with\_ten,modulo\_of\_cube\_by\_eleven,fibonacci\_index\_small,absolute\_modulus\_gap\_elevenNumber Pairs\(a: int, b: int\) \-\> intlast\_digit\_sum,zero\_if\_opposite\_else\_product,absolute\_sum\_minus\_absolute\_difference,sum\_if\_negative,difference\_squared,sum\_after\_incrementing\_first,manhattan\_to\_pair,sum\_plus\_larger\_abs,triple\_sum\_minus\_product,product\_mod\_sevenString\(text: str\) \-\> strrot13\_lowercase\_only,keep\_whitespace\_only,mirror\_with\_pipe,strip\_and\_lowercase,repeat\_three\_times,remove\_whitespace,hex\_code\_points,take\_last\_three,letters\_only,prepend\_hashList\(items: List\[int\]\) \-\> intsum\_excluding\_max,sum\_mod\_three,count\_adjacent\_opposite\_pairs,sum\_smaller\_of\_neighbors,count\_distinct\_values,sum\_of\_positive\_squares,sum\_negative\_values,sum\_neighbors\_products,sum\_palindrome\_values,count\_odd\_valuesLogic\(a: bool, b: bool\) \-\> boolboth\_are\_true\_identity,b\_without\_a,a\_requires\_b,second\_is\_majority,bool\_from\_any\_tuple,a\_or\_b\_via\_if,difference\_negative,nor\_result,all\_or\_none,reverse\_implicationTable 3:The 50 target functions of Alien Abduction: for each domain, 10 functions randomly sampled from the validated, GPT5\.4\-generated pool \(Section[4\.1](https://arxiv.org/html/2608.03388#S4.SS1)\)\.
## Appendix BPrompt Templates
Following standard prompting approaches\(Brownet al\.,[2020](https://arxiv.org/html/2608.03388#bib.bib37); Weiet al\.,[2022](https://arxiv.org/html/2608.03388#bib.bib38); Liuet al\.,[2023](https://arxiv.org/html/2608.03388#bib.bib39)\), we use zero\-shot prompts for all the multi\-turn mode experiments\. The prompts vary by who controls the evidence and its form, but follow a typical structure: task description, constraints, response format restrictions\. Figure[13](https://arxiv.org/html/2608.03388#A4.F13)shows the condensed initial prompt for each of the six game modes, instantiated for a Number domain target with signature\(x: int\) \-\> int\. TheSetupline contains the mode identifier used in the accompanying code and instance files\. In our preliminary experiments, we used 15 turns for the multi\-turn interaction modes and a maximum of 300 tokens\. These limits worked well without any truncation or aborted runs and we therefore use the same settings across all experiments\.
Model% of Instances Failed% Failures that areNot Committed% Not Committed and100% HRA at Last Turn`Qwen3\.6\-35B`97\.0094\.8010\.20`Mistral\-Large\-3`80\.3046\.500\.00`GPT5\.4\-Mini`66\.0014\.1053\.60`GPT5\.4`35\.700\.900\.0Table 4:Failure and non\-commitment rates across models\. The final column reports uncommitted failures with an HRA of 1 at the last turn\.ModelTotal Turns\(across 300 instances\)Median Turns% of Turns withparser errors`Qwen3\.6\-35B`29651540\.60`Mistral\-Large\-3`22617\.545\.20`GPT5\.4\-Mini`13793\.020\.40`GPT5\.4`16164\.04\.60Table 5:Turn usage and parser\-error rates across models over 300 instances\.
## Appendix CQuantitative Analysis
We used an AI\-assisted code generator for the Matplotlib visualisation scripts\. The authors reviewed, modified, and verified all generated code and figures\.
#### Task Success Rates
We report task success rates across domains for all evaluated models\. As shown in Figures[14](https://arxiv.org/html/2608.03388#A5.F14)–[16](https://arxiv.org/html/2608.03388#A5.F16), Qwen3\.6\-35B obtains zero success in several domains and modes\. As discussed in Section[5\.2](https://arxiv.org/html/2608.03388#S5.SS2), the model often exhausts the turn budget without committing to a solution\. In active\-output, it achieves non\-zero success rates for Number, String, and Logic, at 0\.1, 0\.2, and 0\.4, respectively\. Its success rate is zero in both single\-turn modes because the responses result in parser errors\.
GPT5\.4 obtains zero success for String in active\-verdict, but reaches a success rate of 1\.0 for Logic in both single\-turn\-output and passive\-output\. Its success rates in active\-verdict are generally lower across domains, except for List, where it scores 0\.6\. Overall, the domain\-level results show variation across models and interaction modes, with Qwen3\.6\-35B particularly affected by non\-commitment and parser errors, and GPT5\.4 performing weakest in active\-verdict\.
#### Parser Errors
Table[5](https://arxiv.org/html/2608.03388#A2.T5)reports the frequency of model responses that do not follow the required format\. Parser errors are most frequent for Mistral\-Large\-3 and Qwen3\.6\-35B, affecting 45\.2% and 40\.6% of their turns, respectively\. GPT5\.4\-mini has a lower rate of 20\.4%, while GPT5\.4 has the lowest rate at 4\.6%\. These differences indicate that some models have greater difficulty following the structured interaction protocol\. Moreover, frequent parser errors reduce the number of valid turns available within the turn budget and may contribute to non\-commitment\. Their effect is particularly visible in the single\-turn modes for Qwen3\.6\-35B, where an unparsable response results in failure because the model has only one turn to submit a solution\.
#### Hypothesis Retrodiction Accuracy
Figures[17](https://arxiv.org/html/2608.03388#A5.F17)–[19](https://arxiv.org/html/2608.03388#A5.F19)report average HRA across domains\. Missing bars indicate that no valid HRA value was available for that setting\. Active\-output produces consistently high HRA across the available models, particularly for Number, String, and Logic\. Passive\-output is more variable: HRA remains high for Number Pairs and Logic but is lower for Numbers, String and List\. Verdict\-based modes also show greater variation across domains, with lower HRA in passive\-verdict for Number Pairs, List, and Logic\. Overall, these results show that the aggregate pattern is not uniform across domains\.
## Appendix DQualitative Analysis
#### Negative Evidence Impact on Task Success
Figures[21](https://arxiv.org/html/2608.03388#A5.F21),[22](https://arxiv.org/html/2608.03388#A5.F22), and[23](https://arxiv.org/html/2608.03388#A5.F23)compare how models respond to negative evidence in active\-verdict and passive\-verdict\. As discussed in Section[6](https://arxiv.org/html/2608.03388#S6), in active\-verdict, each negative verdict rejects only a model\-proposed input–output pair\. For List and Number Pairs, several failed instances accumulate negative verdicts while the reported hypotheses remain inconsistent with the observed evidence or unknown, and the interaction often continues close to the turn limit\. In passive\-verdict, the Game Master provides both positive and negative examples, and models generally reach evidence\-consistent hypotheses in fewer turns, although some failures remain\. The difference is less pronounced for Logic, where interactions in both modes usually end within a few turns\. Overall, the traces suggest that negative evidence is more useful when it is provided alongside broader example coverage than when it only rejects self\-selected proposals\.
#### Confidence Score
Figure[20](https://arxiv.org/html/2608.03388#A5.F20)compares the models’ self\-reported confidence on successful and failed instances\. Confidence is generally higher for successful instances, but the separation varies across models and modes\. Qwen3\.6\-35B shows gaps in active\-output and passive\-output, while GPT5\.4\-mini shows clearer separation in the passive and single\-turn modes\. In contrast, Mistral\-Large\-3 remains highly confident even on failed instances, with little or no separation in the single\-turn modes\. GPT5\.4 also shows similar confidence for successful and failed instances in passive\-verdict\. These results suggest that confidence does not always reliably distinguish correct from incorrect final hypotheses\. Missing bars indicate that no instances were available for the corresponding outcome\.
#### Query Coverage
Figures[24](https://arxiv.org/html/2608.03388#A5.F24),[25](https://arxiv.org/html/2608.03388#A5.F25)and[26](https://arxiv.org/html/2608.03388#A5.F26)show the query coverage of GPT5\.4 across domains\. As discussed in Section[5\.3](https://arxiv.org/html/2608.03388#S5.SS3), the self\-selected queries in the active modes cover a narrow range than in the passive modes\. Overall, this suggests that GPT5\.4’s self\-selected queries explore only a limited portion of the input space\.
#### Current Interaction State
Although the current interaction state is self\-reported by models, it may not completely reflect their internal belief states\. We therefore use this information only as complementary evidence for understanding the model behavior\. We hypothesize that a model that reliably maintains a belief state would indicate a confirming state before committing to a solution\. Figure[27](https://arxiv.org/html/2608.03388#A5.F27)shows the interaction states reported by models at the last turn\. Qwen3\.6\-35B’s responses mostly indicate aprobingstate rather thanconfirming\. This is consistent with the findings in Section[5\.1](https://arxiv.org/html/2608.03388#S5.SS1), where the model often exhausts its turn budget\. On the other hand, GPT5\.4 indicates confirming in most modes, while Mistral\-Large\-3 does so in several modes\. GPT5\.4\-mini sometimes usesProposingorSolving, which do not match the response choices given in the prompt\. Moreover, in both active\-output and active\-verdict, the model indicates probing at the last turn and then immediately commits to a solution, implying that the model may not be accurately reporting its internal belief state\.
Figure 9:Qualitative analysis of GPT\-5\.4 on all variants\. The chosen actions and corresponding outputs are given across turns\. Initial prompt that covers the tasks, the given function signature is excluded for visualisation reasons\.Figure 10:GPT\-5\.4\-mini outputs for six variantsFigure 11:Qwen\-3\.6 outputs for six variantsFigure 12:Mistral outputs for six variantsTemplate D\.1: Active\-Output \(AO\) Infer hidden Python function from limited evidence\. Signature: \(x: int\) \-\> int Turn budget: 15 Setup: Active Inputs Action per turn: TEST: <inputs\> or SOLVE: ‘‘‘python \.\.\. ‘‘‘ GM reply format: OUTPUT: <value\> Visible example format: \(\-10, 10\)
Template D\.2: Active\-Verdict \(AV\) Infer hidden Python function from limited evidence\. Signature: \(x: int\) \-\> int Turn budget: 15 Setup: Active Pair Checks Action per turn: TEST: <inputs\>, <candidate\_output\> or SOLVE: ‘‘‘python \.\.\. ‘‘‘ GM reply format: OUTPUT: True/False Visible example format: \(\(\-10, 10\), True\)
Template D\.3: Passive\-Output \(PO\) Infer hidden Python function from limited evidence\. Signature: \(x: int\) \-\> int Turn budget: 15 Setup: Passive Examples Action per turn: NEXT: or SOLVE: ‘‘‘python \.\.\. ‘‘‘ GM reveals one valid pair at a time: \(\-10, 10\) GM reply format: OUTPUT: \(<input\>, <output\>\)
Template D\.4: Passive\-Verdict \(PV\) Infer hidden Python function from limited evidence\. Signature: \(x: int\) \-\> int Turn budget: 15 Setup: Passive Labeled Pairs Action per turn: NEXT: or SOLVE: ‘‘‘python \.\.\. ‘‘‘ GM reveals one labeled pair at a time: \(\(\-10, 10\), True\) GM reply format: OUTPUT: \(\(<input\>, <candidate\_output\>\), True/False\)
Template D\.5: Single\-Turn\-Output \(STO\) Infer hidden Python function from limited evidence\. Signature: \(x: int\) \-\> int Turn budget: 1 Setup: Passive Examples One\-Shot Only action: SOLVE: ‘‘‘python \.\.\. ‘‘‘ All valid examples are shown up front; no NEXT requests allowed\. Preloaded batch: \(\-10, 10\) \| \(0, 0\) \.\.\.
Template D\.6: Single\-Turn\-Verdict \(STV\) Infer hidden Python function from limited evidence\. Signature: \(x: int\) \-\> int Turn budget: 1 Setup: Passive Labeled Pairs One\-Shot Only action: SOLVE: ‘‘‘python \.\.\. ‘‘‘ All labeled candidate pairs are shown up front; no NEXT requests allowed\. Preloaded batch: \(\(\-10, 10\), True\) \| \(\(0, \-1\), False\) \.\.\.
Figure 13:Condensed initial prompts for the six game modes of Alien Abduction, shown for a Number domain target with signature\(x: int\) \-\> int\.
## Appendix EExample Episodes
This appendix shows two lightly shortened Single\-Turn\-Output episodes: one lost \(Table[6](https://arxiv.org/html/2608.03388#A5.T6)\) and one won \(Table[7](https://arxiv.org/html/2608.03388#A5.T7)\)\. Game Master messages are shown in gray and player responses in green; the final row of each transcript shows the hidden target function\.
\#Game MasterPlayer1Infer hidden Python function from fixed evidence\.You know: \(x: int\) \-\> int; You have 1 turns total\.Mode: Passive Examples One\-ShotOne action only: SOLVE: ‘‘‘python \.\.\. ‘‘‘\.\.\.Preloaded examples:1\. \(\-10, 10\) \| 2\. \(0, 0\) \| 3\. \(\-2, 2\) \| 4\. \(\-99, 99\) \| 5\. \(99, 99\)6\. \(10, 10\) \| 7\. \(1000, 1000\) \| 8\. \(2, 2\) \| 9\. \(\-1, 1\) \| 10\. \(1, 1\)2SOLVE: ‘‘‘pythondef solution\(x\):return x‘‘‘3Game Verdict: LOSS4Incorrect\.Failed Cases:Input: \[\-10\] \| Expected: 10 \| Got: \-10Input: \[\-2\] \| Expected: 2 \| Got: \-2Input: \[\-99\] \| Expected: 99 \| Got: \-995def absolute\_value\(x: int\) \-\> int:return abs\(x\)Table 6:A lost Single\-Turn\-Output episode\. The player submits the identity function although the preloaded examples contain negative inputs with positive outputs; the hidden target isabsolute\_value\.\#Game MasterPlayer1Infer hidden Python function from fixed evidence\.You know: \(text: str\) \-\> str; You have 1 turns total\.Mode: Passive Examples One\-ShotOne action only: SOLVE: ‘‘‘python \.\.\. ‘‘‘\.\.\.Preloaded examples:1\. \(’hello’, ’h’\) \| 2\. \(’hello world’, ’h’\) \| 3\. \(’123’, ’1’\) \| 4\. \(’ab’, ’a’\) \| 5\. \(’a’, ’a’\)6\. \(’a\-b’, ’a’\) \| 7\. \(’ ’, ’ ’\) \| 8\. \(’aba’, ’a’\) \| 9\. \(’ ’, ’ ’\) \| 10\. \(’’, ’’\)2SOLVE: ‘‘‘pythondef solution\(text: str\) \-\> str:return text\[0\] if text else ’’‘‘‘3Game Verdict: WIN4That is correct\!5def first\_char\(text: str\) \-\> str:return text\[:1\]Table 7:A won Single\-Turn\-Output episode\. The player infers the correct rule from the preloaded examples, including the empty\-string edge case; the hidden target isfirst\_char\.\(\(a\)\)Task success rates for the Number domain\.
\(\(b\)\)Task success rates for the Number Pairs domain\.
Figure 14:Task success rates across models and interaction modes for the Number and Number Pairs domains\. Missing bars indicate zero success\.\(\(a\)\)Task success rates for the String domain\.
\(\(b\)\)Task success rates for the List domain\.
Figure 15:Task success rates across models and interaction modes for the String and List domains\. Missing bars indicate zero success\.Figure 16:Task success rates across models and interaction modes for the Logic domain\. Missing bars indicate zero success\.\(\(a\)\)Hypothesis Retrodiction Accuracy for the Number domain\.
\(\(b\)\)Hypothesis Retrodiction Accuracy for the Number Pairs domain\.
Figure 17:Hypothesis Retrodiction Accuracy across models and interaction modes for the Number and Number Pairs domains\. Missing bars indicate zero success\.\(\(a\)\)Hypothesis Retrodiction Accuracy for the String domain\.
\(\(b\)\)Hypothesis Retrodiction Accuracy for the List domain\.
Figure 18:Hypothesis Retrodiction Accuracy across models and interaction modes for the String and List domains\. Missing bars indicate zero success\.Figure 19:Hypothesis Retrodiction Accuracy across models and interaction modes for the Logic domain\. Missing bars indicate zero success\.Figure 20:Confidence scores at the final turn across models and interaction modes\.Figure 21:Negative evidence impact on task success across verdict modes for GPT5\.4 in List domain\.Figure 22:Negative evidence impact on task success across verdict modes for GPT5\.4 in Logic domain\.Figure 23:Negative evidence impact on task success across verdict modes for GPT5\.4 in Number Pairs domain\.Figure 24:Distribution of input queries across interaction modes for GPT5\.4 in the List domain\.Figure 25:Distribution of input queries across interaction modes for GPT5\.4 in the Logic domain\.Figure 26:Distribution of input queries across interaction modes for GPT5\.4 in the Numbers domain\.Figure 27:Interaction state at the last turn across models and interaction modes\.Similar Articles
@rohanpaul_ai: LLMs frequently underestimate how much information they need, so test when they stop instead of relying on answer accur…
A tweet summarizing an arxiv paper that evaluates how LLMs handle multi-turn information seeking, highlighting that LLMs often underestimate their information needs and may rely on prior knowledge rather than gathering sufficient evidence.
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
This paper introduces MT-InfoSeek, a controlled evaluation suite for assessing LLMs' multi-turn information seeking capabilities, revealing that models underestimate missing information and perform poorly with increased underspecification across domains.
LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos
The paper introduces Elenchos, a generative evaluation framework for abductive reasoning in LLMs, where models must infer hidden rule changes from behavioral differences under black-box access. It finds a detection-attribution dissociation: models detect alterations but struggle to identify the specific mutations, especially under interacting mutations.
FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games
FalsifyBench is a new evaluation framework for assessing inductive reasoning in LLMs, inspired by the Wason 2-4-6 task, where agents discover hidden semantic rules by proposing examples and receiving feedback. Evaluation of 12 LLMs shows reasoning models outperform instruction-tuned models, with negative testing (hypothesis falsification) being the key driver of success.
Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?
This study examines whether active exploration helps adults overcome the 'conjunctive handicap' in causal reasoning, comparing human performance to LLMs in a blicket detector task. Results show that active exploration improves conjunctive reasoning in adults, though some gaps remain, and LLMs approach human accuracy but explore less efficiently.