When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
Summary
This paper introduces a benchmark to study how large language models arbitrate conflicting evidence from text and numerical sources, finding that models use heuristic strategies with biases towards recency and external tools.
View Cached Full Text
Cached at: 08/21/26, 10:18 AM
# When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
Source: [https://arxiv.org/html/2608.20116](https://arxiv.org/html/2608.20116)
Edward PhillipsFredrik K\. GustafssonPatitapaban PaloLei CliftonAffiliation:Nuffield Department of Primary Care Health Sciences, University of Oxford, Oxford, UKDanielle BelgraveAffiliation:GlaxoSmithKline, London, UKXiao GuDavid A\. CliftonAffiliation:Oxford Suzhou Centre for Advanced Research, University of Oxford, Suzhou, Jiangsu, China\[0\.7em\] Department of Engineering ScienceUniversity of OxfordOxfordUK
###### Abstract
Large language models \(LLMs\) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence\. We study how LLMs arbitrate between such sources when they support opposing decisions\. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground\-truth label\. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance\. Across open\-weight instruction\-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text\-versus\-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over\-rely on external forecasts even when they conflict with direct contextual evidence\. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool\-augmented decision systems\.
*Keywords*Large Language Models⋅\\cdotMultimodal Reasoning⋅\\cdotEvidence Integration⋅\\cdotConflict Resolution⋅\\cdotNumerical Reasoning
## 1Introduction
In many real\-world decision\-making scenarios, different sources of evidence may support conflicting conclusions\. In healthcare, for example, a clinician’s assessment may describe a patient as stable while vital signs indicate ongoing deterioration\. Similarly, in manufacturing, sensor readings may suggest normal machine operation even as maintenance reports point to impending failure\. Such inconsistencies can arise for several reasons, including temporal mismatches between data sources \(Figure[1](https://arxiv.org/html/2608.20116#S1.F1)\), differences in source reliability, or the integration of observational evidence with inaccurate predictions generated by external tools\.
Figure 1:Evidence arbitration under conflict\.A numerical time series and a textual summary, drawn from different windows of a shared latent risk trajectory, support*conflicting*predictions at the target timeT\+kT\{\+\}k:highfor the older numerical source,lowfor the newer textual source\. The task requires the model to decide which source to prioritize when answering the binary\-choice prompt\.Large Language Models \(LLMs\) are increasingly being explored in high\-stakes domains such as healthcare, mental health, and finance\([1](https://arxiv.org/html/2608.20116#bib.bib1);[13](https://arxiv.org/html/2608.20116#bib.bib2);[3](https://arxiv.org/html/2608.20116#bib.bib3);[6](https://arxiv.org/html/2608.20116#bib.bib36);[30](https://arxiv.org/html/2608.20116#bib.bib37);[33](https://arxiv.org/html/2608.20116#bib.bib38);[7](https://arxiv.org/html/2608.20116#bib.bib39)\), while also being incorporated into decision\-making and agentic systems that integrate external tools and heterogeneous data sources\([23](https://arxiv.org/html/2608.20116#bib.bib28);[18](https://arxiv.org/html/2608.20116#bib.bib31);[36](https://arxiv.org/html/2608.20116#bib.bib32)\)\. Understanding how these models behave when confronted with conflicting evidence is therefore increasingly important\.
In this work, we use the term*arbitration*to refer to how models prioritize or reconcile competing signals when different sources support incompatible conclusions\. Failures of arbitration may lead models to privilege stale, unreliable, or incorrect tool\-generated evidence over more relevant observations, producing unreliable decisions even when the correct signal is present in the prompt\.
Prior work has studied conflicts between textual sources\([11](https://arxiv.org/html/2608.20116#bib.bib19);[15](https://arxiv.org/html/2608.20116#bib.bib20)\), between parametric and external knowledge\([34](https://arxiv.org/html/2608.20116#bib.bib14);[32](https://arxiv.org/html/2608.20116#bib.bib15);[26](https://arxiv.org/html/2608.20116#bib.bib16);[27](https://arxiv.org/html/2608.20116#bib.bib17);[14](https://arxiv.org/html/2608.20116#bib.bib18)\), and across modalities such as image and text\([20](https://arxiv.org/html/2608.20116#bib.bib10);[38](https://arxiv.org/html/2608.20116#bib.bib11);[9](https://arxiv.org/html/2608.20116#bib.bib12);[25](https://arxiv.org/html/2608.20116#bib.bib13)\)\. Other work has examined numerical reasoning\([17](https://arxiv.org/html/2608.20116#bib.bib26);[21](https://arxiv.org/html/2608.20116#bib.bib27)\), time\-series forecasting\([5](https://arxiv.org/html/2608.20116#bib.bib21);[12](https://arxiv.org/html/2608.20116#bib.bib22);[8](https://arxiv.org/html/2608.20116#bib.bib23);[19](https://arxiv.org/html/2608.20116#bib.bib25)\), and tool use in LLMs\([23](https://arxiv.org/html/2608.20116#bib.bib28);[24](https://arxiv.org/html/2608.20116#bib.bib29);[2](https://arxiv.org/html/2608.20116#bib.bib30);[18](https://arxiv.org/html/2608.20116#bib.bib31)\)\. However, these lines of work do not directly characterize how LLMs arbitrate between textual and numerical evidence, when these two support opposing decisions\. This setting is increasingly relevant in applications where natural language summaries, numerical measurements, and external model or tool outputs are presented together\.
In practice, conflicts between numerical and textual evidence are rarely attributable to modality alone\. Instead, they arise from the interaction of multiple cues, including modality priors, temporal recency, source reliability, and evidence provenance \(direct observations versus externally generated predictions\)\. These cues can point in different directions: a newer textual report may contradict older numerical measurements, a reliable time series may conflict with a corrupted summary, or an external forecast may disagree with directly observed context\. Systematically characterizing such behaviour is difficult in unconstrained real\-world data because the reliability, provenance, and ground truth associated with each source are often ambiguous\. To address this, we introduce a controlled synthetic benchmark in which these properties are known by construction\.
The benchmark uses latent risk trajectories to generate both numerical time series and natural language summaries\. This framework allows us to construct conflicts where exactly one evidence source is aligned with the ground\-truth label, while independently manipulating modality, temporal recency, source reliability, and evidence provenance\. Our objective is not to reproduce the full complexity of deployed decision\-making systems, but to isolate arbitration tendencies that may otherwise be difficult to identify in real\-world environments\.
Across open\-weight instruction\-tuned models, we find that arbitration behaviour is systematic rather than random\. Models exhibit distinct text\-versus\-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over\-rely on external forecasts even when they conflict with direct contextual evidence\. These findings suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence\.
Our contributions can be summarized as follows:
- •We formulate*textual–numerical evidence arbitration*as a controlled evaluation setting for studying how LLMs prioritize conflicting textual and numerical evidence\.
- •We introduce a large\-scale synthetic benchmark that disentangles the effects of modality, temporal recency, source reliability, and evidence provenance on model decisions\.
- •We uncover systematic arbitration biases and failure modes across modern LLM families, including modality preferences, prompt\-order effects, and over\-reliance on external forecasts, with implications in settings involving heterogeneous or tool\-derived evidence\.
## 2Related Work
#### Conflict Resolution and Evidence Arbitration\.
A growing body of work has examined how LLMs handle conflicting information across different sources of knowledge\. In retrieval\-augmented and contextual generation settings, several studies investigate conflicts between parametric knowledge encoded in model weights and contextual knowledge provided at inference time and how LLMs behave under such discrepancies\([32](https://arxiv.org/html/2608.20116#bib.bib15);[34](https://arxiv.org/html/2608.20116#bib.bib14)\)\.
Building on this, recent approaches proposed mechanisms to improve conflict handling, including context\-aware neuron reweighting\([26](https://arxiv.org/html/2608.20116#bib.bib16)\), controlled integration of retrieved evidence through shared\-private semantic modeling\([27](https://arxiv.org/html/2608.20116#bib.bib17)\), and adaptive decoding strategies\([14](https://arxiv.org/html/2608.20116#bib.bib18)\)\. Related work on text\-only evidence conflicts further shows that LLMs often exhibit strong positional and stylistic biases when resolving contradictions and rarely express uncertainty in the presence of conflicting evidence\([11](https://arxiv.org/html/2608.20116#bib.bib19);[15](https://arxiv.org/html/2608.20116#bib.bib20)\)\.
More recently, researchers have extended the study of knowledge conflicts to multimodal settings, particularly inconsistencies between visual evidence and internal commonsense or textual reasoning\([20](https://arxiv.org/html/2608.20116#bib.bib10);[38](https://arxiv.org/html/2608.20116#bib.bib11);[25](https://arxiv.org/html/2608.20116#bib.bib13);[9](https://arxiv.org/html/2608.20116#bib.bib12)\)\.
In the context of evidence arbitration in LLMs, our work addresses the critical yet largely unexplored challenge of textual–numerical conflicts\.
#### LLMs for Numerical Tasks\.
Recent work has explored how LLMs can be adapted to numerical and forecasting tasks through reprogramming, multimodal prompting, and cross\-modal alignment techniques\([5](https://arxiv.org/html/2608.20116#bib.bib21);[12](https://arxiv.org/html/2608.20116#bib.bib22);[8](https://arxiv.org/html/2608.20116#bib.bib23);[16](https://arxiv.org/html/2608.20116#bib.bib24);[19](https://arxiv.org/html/2608.20116#bib.bib25)\)\. At the same time, several studies question the effectiveness of LLMs for numerical reasoning and forecasting, highlighting issues such as poor calibration, sensitivity to noise, weak temporal reasoning, and limited numerical understanding\([28](https://arxiv.org/html/2608.20116#bib.bib4);[22](https://arxiv.org/html/2608.20116#bib.bib5);[17](https://arxiv.org/html/2608.20116#bib.bib26);[21](https://arxiv.org/html/2608.20116#bib.bib27)\)\. Motivated by these limitations, we formulate our setting as a binary forecasting task rather than a purely numerical prediction problem, allowing us to study evidence arbitration without requiring precise numerical reasoning\.
#### Tool\-Augmented and Agentic LLM Systems\.
The rapidly growing subfield of tool\-augmented and agentic LLMs has emphasized the role of iterative reasoning, planning, feedback, and external tool use in complex decision\-making tasks\([36](https://arxiv.org/html/2608.20116#bib.bib32);[23](https://arxiv.org/html/2608.20116#bib.bib28);[24](https://arxiv.org/html/2608.20116#bib.bib29);[2](https://arxiv.org/html/2608.20116#bib.bib30);[18](https://arxiv.org/html/2608.20116#bib.bib31)\)\. More recently, these paradigms have been extended to time series analysis, where LLM agents integrate textual reasoning with numerical and domain\-specific evidence for forecasting and multi\-step inference\([31](https://arxiv.org/html/2608.20116#bib.bib33);[37](https://arxiv.org/html/2608.20116#bib.bib34);[39](https://arxiv.org/html/2608.20116#bib.bib35)\)\. Our work is closely related to this line of research, where models often need to integrate numerical evidence with language\-based reasoning\.
## 3Textual–Numerical Evidence Arbitration
We now define the arbitration task and the controlled benchmark used to evaluate it\. Each instance asks a model to make a binary prediction about a future target value from textual, numerical, and optionally tool\-derived evidence\. In the conflict settings, two evidence sources support opposing labels, with exactly one source aligned with the ground truth\. We then describe the four conflict dimensions, the synthetic data and prompt generation pipeline, and the evaluation protocol\.
Figure 2:Benchmark conflict settings for evidence arbitration\.Each panel illustrates one of the four controlled conflict settings used in our evaluation\. In every setting, two evidence sources support opposing decisions and exactly one source is aligned with the ground\-truth label\.\(A\) Baseline modality priors:both sources cover the same window\[0,T\]\[0,T\]and merely disagree\.\(B\) Temporal recency:sources cover different time windows; the more recent source is always aligned with truth\.\(C\) Reliability:one source is marked as unreliable \(a corruption note in text, missing values in numbers\); the reliable source is aligned with truth\.\(D\) Tool forecast:an external forecasting tool predicts a value atT\+kT\{\+\}kthat contradicts the observed context; context is aligned with truth\.### 3\.1Task Definition
We study*textual–numerical evidence arbitration*: the problem of deciding which source to prioritize when textual and numerical evidence supports incompatible conclusions\. As illustrated in Figure[1](https://arxiv.org/html/2608.20116#S1.F1), the model receives a prompt containing two evidence sources and must answer a binary\-choice question about a future target value\.
Each instance is generated from a latent risk trajectory with values in\[0,1\]\[0,1\]\. The model observes evidence derived from the trajectory up to timeTTand predicts whether the future value atT\+kT\+kishigh\(\>0\.5\>0\.5\) orlow\(<0\.5<0\.5\), wherek≥1k\\geq 1\. Evidence is presented as a serialized numerical time series, a natural language summary, or an external forecast\. By construction, only one source is aligned with the ground\-truth label, while the other supports the opposite label\. The model must therefore infer which source to prioritize\.
We use a coarse\-grained binary forecasting objective rather than precise numerical prediction, in order to isolate arbitration behaviour under conflicting evidence while reducing confounds from known limitations of LLMs in fine\-grained numerical forecasting\([28](https://arxiv.org/html/2608.20116#bib.bib4);[22](https://arxiv.org/html/2608.20116#bib.bib5)\)\.
### 3\.2Conflict Dimensions
We evaluate arbitration behaviour using four conflict settings, summarized in Figure[2](https://arxiv.org/html/2608.20116#S3.F2), which disentangle the effects of modality, temporal recency, source reliability, and evidence provenance\.
#### Baseline Modality Priors\.
Textual and numerical evidence are matched in temporal scope and reliability, but support opposite labels\. Across instances, the ground\-truth\-aligned source is alternated between text and numbers\. This setting evaluates whether LLMs exhibit an inherent modality prior when resolving conflicts in the absence of additional arbitration cues\.
#### Temporal Recency Conflicts\.
Textual and numerical evidence describe different temporal windows and support opposite labels\. The more recent source is always aligned with the ground\-truth label, and both sources are presented with explicit timestamps\. This setting tests whether models use temporal recency as an arbitration cue when textual and numerical evidence disagree\.
#### Reliability Conflicts\.
Textual and numerical evidence describe the same temporal window but differ in reliability\. The reliable source is always aligned with the ground\-truth label\. For numerical evidence, unreliability is simulated by randomly masking 50% of time\-series values usingNaNentries; for textual evidence, it is indicated through an explicit statement that the source observations are incomplete or corrupted\. This setting tests whether models appropriately discount evidence marked as unreliable during arbitration\.
#### Tool Forecast Conflicts\.
The model receives contextual evidence describing observations up to timeTT, in either textual or numerical form, together with a simulated external forecast for the target timeT\+kT\+k\. The prompt explicitly states that the forecasting tool analyzed the same observations provided in the context before producing its prediction\. By construction, however, the context is aligned with the ground\-truth label, while the forecast supports the opposite label\. Tool\-generated forecasts are simulated as described in Appendix[B\.3](https://arxiv.org/html/2608.20116#A2.SS3)\. This adversarial tool\-conflict setting evaluates whether models over\-rely on tool\-generated predictions even when they conflict with contextual evidence\.
### 3\.3Benchmark Construction
Figure 3:Benchmark construction pipeline\.A single example is traced through the three stages used to generate model prompts\.\(A\) Time series generation:a sampled label, slope, and margin fix the target valueyT\+k=0\.5±marginy\_\{T\+k\}=0\.5\\pm\\text\{margin\}, from which the observed trajectory is generated by integrating backwards with additive noise\.\(B\) Text generation:trajectory\-level features are discretized and rendered as a natural language summary using templates and synonym sampling\.\(C\) Prompt generation:the textual summary and numerical series are combined with task framing, evidence\-order controls, and answer choices to produce the final binary\-choice prompt\.Our framework comprises three main components: a time series generator, a text generator, and a prompt generator, as illustrated in Figure[3](https://arxiv.org/html/2608.20116#S3.F3)\.
#### Time Series Generation\.
Latent risk trajectories follow a stochastic linear process with additive Gaussian noise and a latent slope parametersscontrolling the overall trend direction and magnitude:
xt\+1=xt\+s\+ϵt,ϵt∼𝒩\(0,σ2\)\.x\_\{t\+1\}=x\_\{t\}\+s\+\\epsilon\_\{t\},\\qquad\\epsilon\_\{t\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)\.To generate each trajectory, we first sample the label, the latent slopess, and a marginmmaround the decision threshold0\.50\.5\(Figure[3](https://arxiv.org/html/2608.20116#S3.F3)A\)\. The target value at horizonT\+kT\+kis then set to0\.5\+m0\.5\+mforhighand0\.5−m0\.5\-mforlow\. We generate the observed trajectory backward from stepTTto step11using the corresponding reverse\-time recursion\. More details about the time series generation procedure are provided in Appendix[B\.1](https://arxiv.org/html/2608.20116#A2.SS1)\. For the main arbitration experiments, we use observed trajectories of lengthT=16T=16, and a forecasting horizon ofk=1k=1to keep task difficulty manageable\. Each value is paired with a generated timestamp, allowing the series to be serialized into a temporally grounded representation\. We use a simulated sampling frequency of one minute in all experiments\. For each conflict dimension, we construct balanced datasets with equal numbers ofhighandlowlabels\. Examples of generated trajectories are shown in Figure[A1](https://arxiv.org/html/2608.20116#A2.F1)in the appendix\.
#### Text Generation\.
The text generator converts each time series into a natural language description of its overall trajectory \(Figure[3](https://arxiv.org/html/2608.20116#S3.F3)B\)\. We extract a small set of high\-level characteristics from the series, including the initial level, the direction and strength of the overall trend, whether the trajectory stays on the same side of the decision threshold or transitions across it, and the final position relative to the threshold\. These characteristics are discretized into semantic categories \(e\.g\., low/moderate/high initial level, weak/moderate/strong increase or decrease, slightly or moderately above/below the threshold\) and mapped to predefined natural language phrases through template\-based rules\. The selected phrases are then composed into complete textual summaries using randomized synonym and template choices to increase linguistic variability while preserving semantic consistency\. More details about the text generation procedure are provided in Appendix[B\.2](https://arxiv.org/html/2608.20116#A2.SS2)\. To maintain a clear separation between textual and numerical evidence, the generated descriptions never include explicit numerical values or exact measurements from the underlying time series\.
To create conflicts between textual and numerical evidence, we keep the original numerical evidence unchanged and sample a second trajectory conditioned on the opposite label\. This newly generated series is then converted into text using the same generation pipeline, yielding textual evidence that is semantically coherent but inconsistent with the accompanying numerical evidence\.
#### Prompt Generation\.
Prompts are constructed in a modular fashion \(Figure[3](https://arxiv.org/html/2608.20116#S3.F3)C\)\. The first component introduces the task and optionally provides domain\-specific framing\. We consider a generic framing, along with three domain\-specific instantiations: healthcare, industrial, and finance\. We then present the available evidence, which may include textual summaries, numerical time series observations, and/or simulated external forecast predictions\. After presenting the evidence, the prompt explicitly asks the model to predict whether the target risk value will behighorlow\. The prediction is formulated as a binary\-choice decision between answer options “A” and “B”\. We provide examples of full prompts in Appendix[C\.7](https://arxiv.org/html/2608.20116#A3.SS7)\. We systematically control the order in which evidence sources are presented \(i\.e\., whether the source aligned with the ground truth appears first or last in the prompt\)\. This manipulation allows us to study how evidence presentation order influences arbitration behaviour\. Additional details on prompt design can be found in Appendix[C](https://arxiv.org/html/2608.20116#A3)\.
### 3\.4Evaluation Protocol
We evaluate arbitration behaviour using prompt templates tailored to each experiment and applied consistently across all models\. For each sample, the model selects one of the two answer options \(“A” or “B”\) based on the provided textual and/or numerical evidence\. Predictions are then obtained directly from the logits associated with the binary answer tokens, which are mapped to the corresponding labels\.
For each experiment, we report classification accuracy under conflicting evidence conditions, alongside unimodal reference conditions in which only the evidence source aligned with the ground truth is provided\. Unless otherwise specified, all experiments use the default configuration shown in Table[A1](https://arxiv.org/html/2608.20116#A4.T1)\. Results are averaged across three random seeds\.
### 3\.5Models
We evaluate a diverse set of open\-weight instruction\-tuned LLMs spanning multiple architectures and parameter scales, including the Qwen3 family \(1\.7B, 4B, 8B, and 14B\)\([35](https://arxiv.org/html/2608.20116#bib.bib6)\), Gemma\-2\-9B\-It \(Gemma\)\([29](https://arxiv.org/html/2608.20116#bib.bib7)\), Llama\-3\-8B\-Instruct \(Llama\)\([4](https://arxiv.org/html/2608.20116#bib.bib8)\), and Mistral\-7B\-Instruct\-v0\.3 \(Mistral\)\([10](https://arxiv.org/html/2608.20116#bib.bib9)\)\. This selection enables comparisons both across model families and within a controlled scaling series\.
## 4Results
We evaluate arbitration behaviour across a range of controlled conflict settings designed to isolate the effects of modality, temporal recency, source reliability, and evidence provenance, as well as through sensitivity analyses examining the impact of domain instantiations and answer choice configurations\. For the main arbitration experiments, we generate balanced datasets containing 2000 instances per setting \(1000 per label\), while sensitivity analyses use 1000 instances per setting for computational efficiency\.
ABCDEFGH
Figure 4:Arbitration accuracy under conflicting evidence\.Panels show classification accuracy across the four conflict settings and evidence\-order conditions:\(A–B\)baseline modality priors,\(C–D\)temporal recency,\(E–F\)reliability, and\(G–H\)tool forecast conflicts\. In each setting, exactly one evidence source is aligned with the ground\-truth label \(GT\)\. Shaded bars show unimodal reference accuracy using only the ground\-truth\-aligned source, while solid bars show accuracy when both conflicting sources are provided\. Results are averaged over three random seeds\.### 4\.1Baseline Modality Priors
Results in Figures[4A](https://arxiv.org/html/2608.20116#S4.F4.sf1)–[4B](https://arxiv.org/html/2608.20116#S4.F4.sf2)show that all models achieve very high unimodal accuracy, indicating that performance differences in the conflicting setting primarily reflect arbitration behaviour rather than intrinsic task difficulty\. Distinct modality priors emerge across model families\. Qwen3 models consistently favour numerical evidence, whereas Llama and Mistral models show comparatively stronger reliance on textual evidence\. Gemma exhibits the most balanced behaviour between the two modalities\.
The numerical preference of Qwen3 is particularly pronounced: all Qwen3 variants systematically favour numerical evidence even when it conflicts with perfectly predictive textual context\. No clear monotonic relationship with model size is observed\. In contrast, the complementary behaviour of Llama and Mistral may partially reflect weaker overall capability to interpret the numerical evidence, as these models also obtain slightly lower unimodal numerical\-only accuracies compared to Qwen3 and Gemma\.
Evidence order substantially affects arbitration behaviour across nearly all models\. Accuracy is generally higher when the source aligned with the ground truth appears later in the prompt, revealing a strong prompt recency effect\. This effect interacts with modality priors: numerical evidence remains influential regardless of position, whereas textual evidence benefits much more strongly from appearing last\. In several conditions, larger Qwen3 variants even fall below chance \(accuracy below0\.50\.5\) when textual evidence is correct and numerical evidence conflicts, suggesting a systematic bias toward numerical evidence rather than simple uncertainty\.
### 4\.2Temporal Recency Conflicts
As shown in Figures[4C](https://arxiv.org/html/2608.20116#S4.F4.sf3)–[4D](https://arxiv.org/html/2608.20116#S4.F4.sf4), compared to the baseline modality prior setting, the temporal recency setting produces much stronger and more consistent arbitration behaviour across model families, suggesting that temporal recency is a particularly influential cue for resolving conflicting evidence\. We observe that the order of evidence presentation again has a substantial effect: performance is generally higher when the most recent source is presented later in the prompt, reinforcing the prompt recency effects already observed in the baseline experiments\. Nevertheless, models differ in their sensitivity to evidence order\.
Gemma exhibits the most consistent behaviour, maintaining high accuracy regardless of evidence order\. Qwen3 models also follow temporal recency cues very reliably, although some variants are more sensitive to prompt ordering than Gemma\. As in the baseline experiments, no clear monotonic relationship with model size is observed\.
### 4\.3Reliability Conflicts
Across both evidence order settings \(Figures[4E](https://arxiv.org/html/2608.20116#S4.F4.sf5)–[4F](https://arxiv.org/html/2608.20116#S4.F4.sf6)\), reliability conflicts produce larger performance drops than temporal recency conflicts across most models, suggesting that explicit source reliability is a weaker arbitration cue than temporal recency\. Gemma again exhibits the most stable behaviour across ordering conditions, whereas Qwen3 variants appear more sensitive to evidence order despite achieving some of the highest peak accuracies\.
### 4\.4Tool Forecast Conflicts
Figures[4G](https://arxiv.org/html/2608.20116#S4.F4.sf7)–[4H](https://arxiv.org/html/2608.20116#S4.F4.sf8)show that tool forecast conflicts produce the strongest degradation observed across all experiments, indicating that many models heavily over\-rely on external forecasts even when these systematically conflict with the provided contextual measurements\.
Evidence order has a particularly strong effect in this setting\. Accuracy improves substantially when the contextual evidence is presented after the tool forecast, indicating that later evidence can partially mitigate over\-reliance on the external prediction\. Relative to the baseline modality\-prior experiments, the introduction of an explicit forecast greatly amplifies arbitration failures, especially for Qwen3 and Gemma\. These models are particularly susceptible when the ground\-truth\-aligned contextual evidence appears first, often achieving near\-zero accuracy despite perfectly predictive contextual evidence\.
Llama and Mistral are substantially less influenced by incorrect tool forecasts, frequently retaining relatively high accuracy even in the conflicting setting\.
### 4\.5Sensitivity Analysis
To assess robustness to domain instantiation and answer choice configuration, we perform sensitivity analyses in the same setting used for the baseline modality prior experiments\. We vary either the domain or the answer choice configuration, while keeping all remaining parameters fixed to the default hyperparameter values in Table[A1](https://arxiv.org/html/2608.20116#A4.T1)and using the default domain \(i\.e\., generic\), label semantics \(“A”=high\), and answer ordering \(i\.e\., “A” first\) as the baseline configuration\. For each sweep condition, we compute accuracy differences relative to this baseline and report the absolute deltas aggregated across sweep values, seeds and evidence\-order settings as mean±\\pmstandard deviation, separately for each model\.
As shown in Table[A2](https://arxiv.org/html/2608.20116#A5.T2), sensitivity to both domain specialization and answer choice configuration is generally low in text\-only settings, but more noticeable effects emerge in numeric\-only and conflicting settings for several models\. In particular, answer choice perturbations can produce significant shifts in conflict accuracy despite relatively stable unimodal performance, indicating that arbitration behaviour can depend on superficial prompt structure\. Robustness also tends to improve with scale within the Qwen3 family, with the 14B model remaining comparatively stable across all settings\. Additional sensitivity analyses on data\-generation parameters are described in Appendix[E](https://arxiv.org/html/2608.20116#A5)\.
## 5Discussion
Across all experiments, arbitration behaviour is highly systematic rather than random\. Models consistently rely on salient evidence characteristics, including modality, temporal recency, source reliability, and external forecasts, even when these cues conflict with the ground truth\. Temporal recency emerges as the most consistently followed arbitration signal across model families, whereas reliability cues are weaker and lead to substantially larger performance degradation\. External forecasts are particularly influential: tool forecast conflicts produce the strongest failures overall, indicating that many models heavily privilege explicit predictions over directly observed contextual measurements\.
Distinct arbitration patterns also emerge across model families\. Qwen3 models consistently favour numerical evidence and are especially susceptible to misleading external forecasts, while Llama and Mistral rely comparatively more on textual evidence, albeit with overall lower accuracy\. Gemma exhibits the most balanced and stable behaviour across settings\. Importantly, these behaviours do not scale monotonically with model size within the Qwen3 family, suggesting that arbitration biases are not simply a function of parameter count\.
Evidence presentation order plays a major role in arbitration\. Across all experimental settings, evidence presented later in the prompt tends to exert greater influence on the final prediction, partially overriding earlier conflicting information\. This positional effect often amplifies the underlying arbitration cue itself, for example strengthening the influence of temporally recent evidence or external forecasts when they appear last\. Sensitivity analyses further show that arbitration behaviour also can vary under changes in answer choice configuration and domain framing, although these effects are generally smaller\.
Several settings produce below\-chance or near\-zero accuracy\. This indicates that models are not merely uncertain under conflict, but can systematically favour incorrect evidence sources\. Overall, the results suggest that current LLMs rely heavily on heuristic arbitration strategies rather than robust evidence integration, making them vulnerable to predictable and systematic failures in multi\-source decision\-making settings\.
## 6Conclusion
As LLMs are increasingly embedded in decision\-making pipelines, their ability to handle conflicting evidence becomes central to their reliability\. This work studied evidence arbitration between textual summaries, numerical observations, and external tool outputs that support incompatible conclusions, by introducing a controlled synthetic benchmark that isolates key arbitration cues \(modality, temporal recency, source reliability, and evidence provenance\)\.
Our results show that LLMs do not resolve such conflicts randomly\. Instead, they exhibit systematic, model\-specific arbitration patterns, often relying on heuristic cues when deciding which evidence to trust\. Temporal recency is followed more consistently than explicit reliability information, while external forecasts can exert disproportionate influence even when they conflict with direct contextual evidence\.
These findings suggest that evaluating LLMs on isolated textual, numerical, or tool\-use tasks is insufficient for understanding their behaviour in multi\-source decision settings\. Conflict\-based evaluations provide a useful stress test for evidence integration, and we hope this benchmark motivates further work on arbitration under real\-world source conflicts in tool\-augmented systems\.
## Limitations
This work does not include experiments on real\-world data\. Instead, the synthetic framework intentionally simplifies real\-world decision\-making settings in order to provide full control over arbitration cues, including temporal recency, source reliability, and evidence provenance, enabling systematic analysis of arbitration behaviour under conflict\.
In addition, our task formulation casts forecasting as a binary decision problem, which does not capture the full complexity of numerical forecasting tasks\. However, this design reduces confounds arising from known limitations of current LLMs in accurate numerical prediction, allowing cleaner evaluation of how models prioritize conflicting evidence sources\.
#### Potential Risks\.
Our benchmark is intended solely for evaluating evidence arbitration in LLMs and should not be interpreted as guidance for deploying such models in high\-stakes decision\-making settings\.
## Acknowledgments
DAC was funded by an NIHR Research Professorship \(NIHR302440\); a Royal Academy of Engineering Research Chair; and the InnoHK Hong Kong Centre for Cerebro\-cardiovascular Engineering \(COCHE\); and was supported by the National Institute for Health Research \(NIHR\) Oxford Biomedical Research Centre \(BRC\) and the Pandemic Sciences Institute at the University of Oxford\.
## References
- Burtonet al\.\(2024\)J\. W\. Burton, E\. Lopez\-Lopez, S\. Hechtlinger, Z\. Rahwan, S\. Aeschbach, M\. A\. Bakker, J\. A\. Becker, A\. Berditchevskaia, J\. Berger, L\. Brinkmann,et al\.How large language models can reshape collective intelligence\.Nature human behaviour8\(9\),pp\. 1643–1655\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Fenget al\.\(2025\)J\. Feng, S\. Huang, X\. Qu, G\. Zhang, Y\. Qin, B\. Zhong, C\. Jiang, J\. Chi, and W\. ZhongRetool: reinforcement learning for strategic tool use in llms\.arXiv preprint arXiv:2504\.11536\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Gohet al\.\(2025\)E\. Goh, R\. J\. Gallo, E\. Strong, Y\. Weng, H\. Kerman, J\. A\. Freed, J\. A\. Cool, Z\. Kanjee, K\. P\. Lane, A\. S\. Parsons,et al\.GPT\-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial\.Nature Medicine31\(4\),pp\. 1233–1238\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§3\.5](https://arxiv.org/html/2608.20116#S3.SS5.p1.1)\.
- Gruveret al\.\(2023\)N\. Gruver, M\. Finzi, S\. Qiu, and A\. G\. WilsonLarge language models are zero\-shot time series forecasters\.Advances in neural information processing systems36,pp\. 19622–19635\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Heinzet al\.\(2025\)M\. V\. Heinz, D\. M\. Mackin, B\. M\. Trudeau, S\. Bhattacharya, Y\. Wang, H\. A\. Banta, A\. D\. Jewett, A\. J\. Salzhauer, T\. Z\. Griffin, and N\. C\. JacobsonRandomized trial of a generative ai chatbot for mental health treatment\.Nejm Ai2\(4\),pp\. AIoa2400802\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Hu and Zhao \(2026\)X\. Hu and J\. ZhaoFin\-bias: comprehensive evaluation for llm decision\-making under human bias in finance domain\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 5678–5694\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Jiaet al\.\(2024\)F\. Jia, K\. Wang, Y\. Zheng, D\. Cao, and Y\. LiuGpt4mts: prompt\-based large language model for multimodal time\-series forecasting\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 23343–23351\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Jiaet al\.\(2026\)Y\. Jia, Y\. Du, K\. Jiang, Y\. Liang, Q\. Ren, Y\. Xin, R\. Yang, F\. Feng, M\. Chen, H\. Lu,et al\.Benchmarking multimodal knowledge conflict for large multimodal models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 22283–22291\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p3.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de Las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7b\.ArXivabs/2310\.06825\.External Links:[Link](https://api.semanticscholar.org/CorpusID:263830494)Cited by:[§3\.5](https://arxiv.org/html/2608.20116#S3.SS5.p1.1)\.
- Jiayanget al\.\(2024\)C\. Jiayang, C\. Chan, Q\. Zhuang, L\. Qiu, T\. Zhang, T\. Liu, Y\. Song, Y\. Zhang, P\. Liu, and Z\. ZhangECON: on the detection and resolution of evidence conflicts\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 7816–7844\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p2.1)\.
- Jinet al\.\(2024\)M\. Jin, S\. Wang, L\. Ma, Z\. Chu, J\. Y\. Zhang, X\. Shi, P\. Chen, Y\. Liang, Y\. Li, S\. Pan,et al\.Time\-llm: time series forecasting by reprogramming large language models, 2024\.arXiv preprint arXiv:2310\.01728\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Johriet al\.\(2025\)S\. Johri, J\. Jeong, B\. A\. Tran, D\. I\. Schlessinger, S\. Wongvibulsin, L\. A\. Barnes, H\. Zhou, Z\. R\. Cai, E\. M\. Van Allen, D\. Kim,et al\.An evaluation framework for clinical use of large language models in patient interaction tasks\.Nature medicine31\(1\),pp\. 77–86\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Khandelwalet al\.\(2025\)A\. Khandelwal, M\. Gupta, and P\. AgrawalCoCoA: confidence\- and context\-aware adaptive decoding for resolving knowledge conflicts in large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 6835–6855\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.348/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.348),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p2.1)\.
- Kurfali and Östling \(2025\)M\. Kurfali and R\. ÖstlingConflicting needles in a haystack: how LLMs behave when faced with contradictory information\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 34361–34376\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1742/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1742),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p2.1)\.
- Langeret al\.\(2025\)P\. Langer, T\. Kaar, M\. Rosenblattl, M\. A\. Xu, W\. Chow, M\. Maritsch, R\. Jakob, N\. Wang, J\. Liu, A\. Verma,et al\.Opentslm: time\-series language models for reasoning over multivariate medical text\-and time\-series data\.arXiv preprint arXiv:2510\.02410\.Cited by:[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2025\)H\. Li, X\. Chen, Z\. Xu, D\. Li, N\. Hu, F\. Teng, Y\. Li, L\. Qiu, C\. J\. Zhang, L\. Qing,et al\.Exposing numeracy gaps: a benchmark to evaluate fundamental numerical abilities in large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 20004–20026\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Li \(2025\)X\. LiA review of prominent paradigms for LLM\-based agents: tool use, planning \(including RAG\), and feedback learning\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 9760–9779\.External Links:[Link](https://aclanthology.org/2025.coling-main.652/)Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1),[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Liuet al\.\(2025a\)P\. Liu, H\. Guo, T\. Dai, N\. Li, J\. Bao, X\. Ren, Y\. Jiang, and S\. XiaCalf: aligning llms for time series forecasting via cross\-modal fine\-tuning\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 18915–18923\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2025b\)X\. Liu, W\. Wang, Y\. Yuan, J\. Huang, Q\. Liu, P\. He, and Z\. TuInsight over sight: exploring the vision\-knowledge conflicts in multimodal LLMs\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 17825–17846\.External Links:[Link](https://aclanthology.org/2025.acl-long.872/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.872),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p3.1)\.
- Loveringet al\.\(2025\)C\. Lovering, M\. Krumdick, V\. D\. Lai, V\. Reddy, S\. Ebner, N\. Kumar, R\. Koncel\-Kedziorski, and C\. TannerLanguage model probabilities are not calibrated in numeric contexts\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 29218–29257\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1)\.
- Parket al\.\(2025\)J\. Park, H\. Lee, D\. Lee, D\. Gwak, and J\. ChooRevisiting llms as zero\-shot time series forecasters: small noise can break large models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 906–922\.Cited by:[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20116#S3.SS1.p3.1)\.
- Qinet al\.\(2024\)Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian,et al\.Toolllm: facilitating large language models to master 16000\+ real\-world apis\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 9695–9717\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1),[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Quet al\.\(2025\)C\. Qu, S\. Dai, X\. Wei, H\. Cai, S\. Wang, D\. Yin, J\. Xu, and J\. WenTool learning with large language models: a survey\.Frontiers of Computer Science19\(8\),pp\. 198343\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Shaoet al\.\(2025\)Z\. Shao, F\. Gao, Z\. Zhu, C\. Luo, H\. Xing, Z\. Yu, Q\. Zheng, M\. Yan, and J\. BuIs cognition consistent with perception? assessing and mitigating multimodal knowledge conflicts in document understanding\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 30911–30932\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p3.1)\.
- Shiet al\.\(2024\)D\. Shi, R\. Jin, T\. Shen, W\. Dong, X\. Wu, and D\. XiongIrcan: mitigating knowledge conflicts in llm generation via identifying and reweighting context\-aware neurons\.Advances in Neural Information Processing Systems37,pp\. 4997–5024\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p2.1)\.
- Suiet al\.\(2025\)Y\. Sui, C\. Li, C\. Zhang, D\. Song, and Q\. LiBridging external and parametric knowledge: mitigating hallucination of llms with shared\-private semantic synergy in dual\-stream knowledge\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 10845–10869\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p2.1)\.
- Tanet al\.\(2024\)M\. Tan, M\. A\. Merrill, V\. Gupta, T\. Althoff, and T\. HartvigsenAre language models actually useful for time series forecasting?\.Advances in Neural Information Processing Systems37,pp\. 60162–60191\.Cited by:[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20116#S3.SS1.p3.1)\.
- Teamet al\.\(2024\)G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé, J\. Ferret, P\. Liu, P\. Tafti, A\. Friesen, M\. Casbon, S\. Ramos, R\. Kumar, C\. L\. Lan, S\. Jerome, A\. Tsitsulin, N\. Vieillard, P\. Stanczyk, S\. Girgin, N\. Momchev, M\. Hoffman, S\. Thakoor, J\. Grill, B\. Neyshabur, O\. Bachem, A\. Walton, A\. Severyn, A\. Parrish, A\. Ahmad, A\. Hutchison, A\. Abdagic, A\. Carl, A\. Shen, A\. Brock, A\. Coenen, A\. Laforge, A\. Paterson, B\. Bastian, B\. Piot, B\. Wu, B\. Royal, C\. Chen, C\. Kumar, C\. Perry, C\. Welty, C\. A\. Choquette\-Choo, D\. Sinopalnikov, D\. Weinberger, D\. Vijaykumar, D\. Rogozińska, D\. Herbison, E\. Bandy, E\. Wang, E\. Noland, E\. Moreira, E\. Senter, E\. Eltyshev, F\. Visin, G\. Rasskin, G\. Wei, G\. Cameron, G\. Martins, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Batra, H\. Dhand, I\. Nardini, J\. Mein, J\. Zhou, J\. Svensson, J\. Stanway, J\. Chan, J\. P\. Zhou, J\. Carrasqueira, J\. Iljazi, J\. Becker, J\. Fernandez, J\. van Amersfoort, J\. Gordon, J\. Lipschultz, J\. Newlan, J\. Ji, K\. Mohamed, K\. Badola, K\. Black, K\. Millican, K\. McDonell, K\. Nguyen, K\. Sodhia, K\. Greene, L\. L\. Sjoesund, L\. Usui, L\. Sifre, L\. Heuermann, L\. Lago, L\. McNealus, L\. B\. Soares, L\. Kilpatrick, L\. Dixon, L\. Martins, M\. Reid, M\. Singh, M\. Iverson, M\. Görner, M\. Velloso, M\. Wirth, M\. Davidow, M\. Miller, M\. Rahtz, M\. Watson, M\. Risdal, M\. Kazemi, M\. Moynihan, M\. Zhang, M\. Kahng, M\. Park, M\. Rahman, M\. Khatwani, N\. Dao, N\. Bardoliwalla, N\. Devanathan, N\. Dumai, N\. Chauhan, O\. Wahltinez, P\. Botarda, P\. Barnes, P\. Barham, P\. Michel, P\. Jin, P\. Georgiev, P\. Culliton, P\. Kuppala, R\. Comanescu, R\. Merhej, R\. Jana, R\. A\. Rokni, R\. Agarwal, R\. Mullins, S\. Saadat, S\. M\. Carthy, S\. Cogan, S\. Perrin, S\. M\. R\. Arnold, S\. Krause, S\. Dai, S\. Garg, S\. Sheth, S\. Ronstrom, S\. Chan, T\. Jordan, T\. Yu, T\. Eccles, T\. Hennigan, T\. Kocisky, T\. Doshi, V\. Jain, V\. Yadav, V\. Meshram, V\. Dharmadhikari, W\. Barkley, W\. Wei, W\. Ye, W\. Han, W\. Kwon, X\. Xu, Z\. Shen, Z\. Gong, Z\. Wei, V\. Cotruta, P\. Kirk, A\. Rao, M\. Giang, L\. Peran, T\. Warkentin, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, D\. Sculley, J\. Banks, A\. Dragan, S\. Petrov, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, S\. Borgeaud, N\. Fiedel, A\. Joulin, K\. Kenealy, R\. Dadashi, and A\. AndreevGemma 2: improving open language models at a practical size\.External Links:2408\.00118,[Link](https://arxiv.org/abs/2408.00118)Cited by:[§3\.5](https://arxiv.org/html/2608.20116#S3.SS5.p1.1)\.
- Wanget al\.\(2026\)X\. Wang, B\. Gao, Y\. Yang, and D\. A\. CliftonMental\-r1: aligning llm reasoning for mental health assessment\.arXiv preprint arXiv:2606\.13176\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Wanget al\.\(2024\)X\. Wang, M\. Feng, J\. Qiu, J\. Gu, and J\. ZhaoFrom news to forecast: integrating event analysis in llm\-based time series forecasting with reflection\.Advances in Neural Information Processing Systems37,pp\. 58118–58153\.Cited by:[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2023\)Y\. Wang, S\. Feng, H\. Wang, W\. Shi, V\. Balachandran, T\. He, and Y\. TsvetkovResolving knowledge conflicts in large language models\.arXiv preprint arXiv:2310\.00935\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p1.1)\.
- Xieet al\.\(2024\)Q\. Xie, W\. Han, Z\. Chen, R\. Xiang, X\. Zhang, Y\. He, M\. Xiao, D\. Li, Y\. Dai, D\. Feng,et al\.Finben: a holistic financial benchmark for large language models\.Advances in neural information processing systems37,pp\. 95716–95743\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1)\.
- Xuet al\.\(2024\)R\. Xu, Z\. Qi, Z\. Guo, C\. Wang, H\. Wang, Y\. Zhang, and W\. XuKnowledge conflicts for llms: a survey\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 8541–8565\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§3\.5](https://arxiv.org/html/2608.20116#S3.SS5.p1.1)\.
- Yaoet al\.\(2022\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReact: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p2.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Yeet al\.\(2026\)W\. Ye, W\. Yang, D\. Cao, Y\. Zhang, L\. Tang, J\. Cai, and Y\. LiuTS\-reasoner: domain\-oriented time series inference agents for reasoning and automated analysis\.Transactions on Machine Learning Research\.Cited by:[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
- Zhanget al\.\(2025\)Z\. Zhang, T\. Wang, X\. Gong, Y\. Shi, H\. Wang, D\. Wang, and L\. HuWhen modalities conflict: how unimodal reasoning uncertainty governs preference dynamics in mllms\.arXiv preprint arXiv:2511\.02243\.Cited by:[§1](https://arxiv.org/html/2608.20116#S1.p4.1),[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px1.p3.1)\.
- Zhaoet al\.\(2025\)H\. Zhao, X\. Zhang, J\. Wei, Y\. Xu, Y\. He, S\. Sun, and C\. YouTimeseriesscientist: a general\-purpose ai agent for time series analysis\.arXiv preprint arXiv:2510\.01538\.Cited by:[§2](https://arxiv.org/html/2608.20116#S2.SS0.SSS0.Px3.p1.1)\.
## Appendix AUse of AI Assistants
AI assistants were used only to improve the phrasing, clarity, and grammar of this manuscript\. All AI\-generated text was reviewed and revised by the authors, who take full responsibility for the final content\.
## Appendix BData Generation Details
In this section, we provide implementation details for the synthetic data generation pipeline used throughout the experiments\.
### B\.1Time Series Generation
Trajectories are generated using a stochastic linear process with additive Gaussian noise and a latent slope parameter controlling the overall trend direction and magnitude\. For all main experiments, we use trajectories of lengthT=16T=16and a forecasting horizon ofk=1k=1\. In additional sensitivity analyses, we vary the trajectory length usingT∈\{8,16,32\}T\\in\\\{8,16,32\\\}\. The target value at time pointT\+kT\+kis first sampled according to the desired class label\. Letmmdenote the sampled margin from the decision threshold0\.50\.5:
m∼𝒰\(mlow,mhigh\),m\\sim\\mathcal\{U\}\(m\_\{\\text\{low\}\},m\_\{\\text\{high\}\}\),wheremlow=0\.35m\_\{\\text\{low\}\}=0\.35andmhigh=0\.45m\_\{\\text\{high\}\}=0\.45in the main experiments\. In additional sensitivity analyses, we vary the margin range using\(mlow,mhigh\)∈\{\(0\.15,0\.25\),\(0\.25,0\.35\),\(0\.35,0\.45\)\}\(m\_\{\\text\{low\}\},m\_\{\\text\{high\}\}\)\\in\\\{\(0\.15,0\.25\),\(0\.25,0\.35\),\(0\.35,0\.45\)\\\}\. The future target value is then defined as
yT\+k=\{0\.5\+mif label =high,0\.5−mif label =low\.y\_\{T\+k\}=\\begin\{cases\}0\.5\+m&\\text\{if label = \{high\}\},\\\\ 0\.5\-m&\\text\{if label = \{low\}\}\.\\end\{cases\}
To avoid degenerate near\-random forecasting instances, we require the minimum margin to dominate the cumulative stochastic noise over the forecasting horizon\. Specifically, generation is rejected whenever
mlow<1\.5⋅σk,m\_\{\\text\{low\}\}<1\.5\\cdot\\sigma\\sqrt\{k\},whereσ\\sigmadenotes the standard deviation of the Gaussian noise process\.
For each trajectory, a latent slope parameter is independently sampled from a symmetric uniform distribution
s∼𝒰\(−0\.05,0\.05\),s\\sim\\mathcal\{U\}\(\-0\.05,0\.05\),where the range of the uniform distribution is chosen to minimize rejection of the generated trajectories due to violations of the valid risk range\[0,1\]\[0,1\]\. Importantly, the slope distribution is shared across labels to prevent shortcut correlations between slope and class membership\.
Given the sampled future target valueyT\+ky\_\{T\+k\}, we first estimate the final observed valuexTx\_\{T\}by approximately inverting the forward stochastic process
xT=yT\+k−s⋅k−ϵ,x\_\{T\}=y\_\{T\+k\}\-s\\cdot k\-\\epsilon,where
ϵ∼𝒩\(0,σ2\)\\epsilon\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)andσ=0\.05\\sigma=0\.05in the main experiments\. In additional sensitivity analyses, we vary the noise level usingσ∈\{0\.01,0\.05,0\.1\}\\sigma\\in\\\{0\.01,0\.05,0\.1\\\}\. Starting from this estimated valuexTx\_\{T\}, the observed trajectory is then generated backward in time from stepTTto step11\. Letxtx\_\{t\}denote the trajectory value at time steptt\. The backward dynamics follows
xt=xt\+1−s\+ϵt,ϵt∼𝒩\(0,σ2\)\.x\_\{t\}=x\_\{t\+1\}\-s\+\\epsilon\_\{t\},\\quad\\epsilon\_\{t\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)\.A forward simulation step is then performed to verify that the resulting future target remains consistent with the desired label and sampled margin constraint\. Trajectories violating the target label, margin condition, or valid risk range\[0,1\]\[0,1\]are rejected and regenerated using rejection sampling\. Across all experiments, the rejection sampling procedure converged reliably within a small number of attempts\.
Each trajectory value is associated with a synthetic timestamp sampled at one\-minute intervals\. For each instance, a random starting hour is uniformly sampled between 08:00 and 16:00, together with a random starting minute\. Timestamps are then generated sequentially using a fixed simulated sampling frequency of one minute\.
For every conflict dimension, datasets are constructed to maintain balanced class distributions, with equal numbers ofhighandlowlabels\.
Examples of generated trajectories are shown in Figure[A1](https://arxiv.org/html/2608.20116#A2.F1)\.
Figure A1:Representative examples of generated time series\.Generated trajectories with observable lengthT=16T=16and forecast horizonk=1k=1\.
### B\.2Text Generation
Each time series is summarized through a small set of global trajectory characteristics extracted from the full sequence\. The generated text is intentionally high level: it describes coarse temporal behaviour while avoiding explicit numerical values or exact measurements\. This preserves a clear separation between textual and numerical evidence modalities\. All generated textual summaries are produced in English using a template\-based generation procedure with randomized lexical variation\.
#### Extracted trajectory features\.
For each time series, we extract the initial observed value, the final observed value, the overall linear trend direction and magnitude, and whether the trajectory remains on the same side of the decision threshold or crosses it\. The decision threshold is fixed at0\.50\.5\. Descriptions are therefore expressed relative to this threshold \(e\.g\., “below the threshold” or “above the critical level”\)\.
#### Level discretization\.
Continuous values are discretized into semantic categories before text generation\. The following bins are used:
Each category is mapped to multiple synonymous natural language realizations\. For example,slightly\_belowmay be rendered as “a level slightly below the threshold” or “a value marginally below the critical level”\.
#### Trend discretization\.
The global trend is computed from the scalar slope associated with the generated time series\. Trend magnitude is discretized into three categories according to the absolute slope value:
Trend direction is determined by the sign of the slope\. Each direction\-strength pair is mapped to multiple paraphrased textual descriptions \(e\.g\., “rose gradually”, “increased steadily”, “climbed significantly”\)\.
#### Cross\-threshold trajectory semantics\.
The final sentence of each generated description depends jointly on the initial and final regions relative to the threshold\. Four cases are distinguished: remaining below the threshold, remaining above the threshold, crossing from below to above, crossing from above to below\. For example, trajectories crossing from below to above may yield phrases such as “became concerning” or “rose into an elevated state”, whereas trajectories remaining below threshold may produce “continued to stay contained” or “still remained limited”\.
#### Template composition and lexical variation\.
Text generation follows a template\-based pipeline with randomized lexical choices\. Independent synonym banks are used for:
- •sentence opening expressions introducing the initial timestamp \(e\.g\., “Initially, at 12:48, …”, “Starting at 14:17, …”\),
- •temporal closing expressions referring to the final timestamp \(e\.g\., “…by 13:52”, “…as of 09:15”\),
- •synonymous references to the decision threshold \(e\.g\., “the threshold”, “the critical level”, “the cutoff level”\),
- •discourse connectors linking trajectory stages \(e\.g\., “then”, “after that”, “from there”\),
- •descriptions of initial and final value ranges relative to the threshold \(e\.g\., “a moderate value below the threshold”, “an elevated level”\),
- •descriptions of trajectory evolution and final status \(e\.g\., “continued to stay contained”, “moved into a concerning range”, “returned to a controlled range”\)\.
Random sampling from these synonym sets increases linguistic diversity while preserving the underlying semantic content\. Importantly, randomness only affects lexical realization and not the semantics associated with a given trajectory\.
#### Domain conditioning\.
The generation procedure is identical across domains\. Only the domain\-specific risk term changes\. Specifically, we use:
- •“risk” for the generic domain;
- •“deterioration risk” for the healthcare domain;
- •“failure risk” for the industrial domain;
- •“financial distress risk” for the finance domain\.
The entity reference \(“system”, “patient”, “machine”, and “asset”, respectively\) is omitted from the generated summaries, as it is already specified in the instruction component of the prompt and repeating it would introduce unnecessary redundancy\.
#### Unreliability cues\.
We inject unreliability cues into the generated text by appending an explicit statement indicating that the underlying observations on which the summary is based are incomplete or corrupted\. Sample unreliability statements include “Note: this report was generated from partially corrupted data within the observation window” and “Note: this summary was produced using partially corrupted data from the observation window”\.
#### Abstraction gap between modalities\.
The textual modality intentionally summarizes only global trajectory properties and omits local fluctuations, short\-term oscillations, and exact magnitudes present in the numerical series\. As a result, the textual evidence represents an abstract semantic interpretation of the underlying time series rather than a verbalization of every datapoint\.
#### Examples\.
Figure[2](https://arxiv.org/html/2608.20116#S3.F2)presents representative examples of generated textual evidence across different domains\.
### B\.3Tool Forecast Generation
Simulated tool forecasts are generated as intentionally incorrect risk predictions\. For each sample, a scalar risk value is sampled from the opposite side of the decision threshold0\.50\.5relative to the true label\. We additionally control how confidently incorrect the tool prediction is by enforcing a minimum distance from the decision threshold \(set to0\.150\.15in all our experiments\), preventing ambiguous forecasts close to0\.50\.5\.
## Appendix CPrompts
All prompts are generated programmatically using a modular template\-based framework\. Each prompt is composed of the following components:
1. 1\.Domain framing
2. 2\.Task instructions
3. 3\.Evidence blocks
4. 4\.Prediction question
5. 5\.Answer choices
6. 6\.Closing instruction
This modular design enables controlled manipulation of domain instantiation, evidence modality, evidence ordering, and answer ordering while keeping the overall prompt structure fixed across experiments\.
### C\.1Domain Framing
Each prompt begins with a short domain\-specific framing describing the meaning of the risk variable\. Following the text generation setup, we consider one generic framing and three domain\-specific instantiations \(healthcare, industrial, and finance\)\. For example, the healthcare framing is:
Thepatienthasadeteriorationriskbetween0and1,wherehighervaluesindicategreaterriskofclinicaldeteriorationandlowervaluesindicatelowerrisk\.
The underlying prediction task remains identical across domains, only the semantic framing changes\.
### C\.2Task Instructions
The instruction block depends on the available evidence modalities\. We consider five task configurations: numerical evidence only, textual evidence only, textual and numerical evidence, numerical evidence with an external tool forecast, and textual evidence with an external tool forecast\.
For unimodal settings, prompts state
Youaregivenatimeseriesofnumericalriskobservationsoverapasttimeperiod\.
for numerical evidence only, and
Youaregivenatextsummaryofriskobservationsoverapasttimeperiod\.
for textual evidence only\.
For multimodal settings, prompts explicitly specify whether the two evidence sources refer to the same or different temporal windows\. For same\-window settings, prompts state
Youaregiventwosourcesdescribingriskoverthesametimewindow:atextsummaryofriskobservationsandatimeseriesofnumericalriskobservations\.
For temporal recency settings, prompts instead state
Youaregiventwotimestampedsourcesdescribingriskoverpasttimeperiods:atextsummaryofriskobservationsandatimeseriesofnumericalriskobservations\.Thetwosourcesmayrefertodifferenttimewindows,andonesourcemaybeearlierormorerecentthantheother\.
Finally, in settings involving external tool forecasts, prompts extend the unimodal instructions with an additional clause describing the tool prediction:
Youaregiven<unimodalsource\>\.Anexternalforecastingtoolhasanalyzedthesesameobservationsandproducedapredictedriskvalueatafuturetimepoint\.
### C\.3Evidence Formatting
#### Numerical evidence\.
Numerical evidence is presented as timestamp–value pairs:
Timeseries:
12:38:0\.38
12:39:0\.41
12:40:0\.48
12:41:0\.47
\.\.\.
All numerical values are displayed with two decimal places\.
#### Textual evidence\.
Textual evidence is presented as a quoted natural language summary:
Textsummary:
"Startingat13:10,riskwasatavaluemarginallybelowthecutofflevel;afterthat,itfellatamoderatepace,andasof13:25itcontinuedtostaycontained,settlingatalowlevel\."
#### External tool forecasts\.
External forecasts are formatted as scalar predictions associated with the target timestamp:
Externalforecastingtoolriskpredictionattime13:01:0\.76
### C\.4Evidence Order Manipulation
To study ordering effects, the relative ordering of evidence blocks is systematically varied across experiments\.
#### Baseline modality\-prior experiments\.
We vary whether the ground\-truth\-aligned or ground\-truth\-misaligned modality appears first\.
#### Temporal recency experiments\.
We vary whether the more recent or less recent source appears first\.
#### Reliability experiments\.
We vary whether the more reliable or less reliable source appears first\.
#### Tool forecast experiments\.
We vary whether the primary context source or the external tool forecast appears first\.
### C\.5Prediction Question and Answer Choices
After presenting the evidence, prompts ask the model to predict whether the target risk value will exceed a threshold of0\.50\.5:
Question:Basedontheinformationabove,willthesystem’sriskattime15:33behigh\(\>0\.5\)orlow\(<0\.5\)?
Predictions are formulated as binary\-choice decisions using answer options “A” and “B”\. We systematically vary the order in which answer choices appear and the mapping between answer tokens and labels\. For example:
or
This controls for potential positional or token\-level biases\.
### C\.6Closing Instruction
Each prompt concludes with a strict response constraint:
AnswerwithonlyAorB\.Donotaddanyexplanationoradditionaltext\.
Answer:
This instruction was originally introduced to simplify analysis of generated responses\. However, all reported evaluations use logits\-based analysis over the answer tokens \(“A” and “B”\), thereby avoiding confounding effects arising from decoding variability\.
### C\.7Example Prompts
Example prompt illustrating the healthcare domain framing and a temporal recency conflict, where the textual evidence is more recent and aligned with the ground\-truth label:
Thepatienthasadeteriorationriskbetween0and1,wherehighervaluesindicategreaterriskofclinicaldeteriorationandlowervaluesindicatelowerrisk\.
Youaregiventwotimestampedsourcesdescribingriskoverpasttimeperiods:atextsummaryofriskobservationsandatimeseriesofnumericalriskobservations\.Thetwosourcesmayrefertodifferenttimewindows,andonesourcemaybeearlierormorerecentthantheother\.Thetaskistopredictwhetherthepatient’sdeteriorationriskatthetargettimepointwillbehighorlow\.
Timeseries:
14:02:0\.74
14:03:0\.78
14:04:0\.86
14:05:0\.87
14:06:0\.88
14:07:0\.97
14:08:0\.98
14:09:0\.96
Textsummary:
"Initially,at14:26,deteriorationriskwasatarelativelylowlevel;fromthere,itincreasedsteadily,andasof14:33itwasstillundercontrol,settlingatarelativelylowlevel\."
Question:Basedontheinformationabove,willthepatient’sdeteriorationriskattime14:34behigh\(\>0\.5\)orlow\(<0\.5\)?
A\)high
B\)low
AnswerwithonlyAorB\.Donotaddanyexplanationoradditionaltext\.
Answer:
Example prompt illustrating the industrial domain framing and a reliability conflict, where the textual evidence is more reliable and aligned with the ground\-truth label:
Themachinehasafailureriskbetween0and1,wherehighervaluesindicategreaterriskoffailure,andlowervaluesindicatelowerrisk\.
Youaregiventwosourcesdescribingriskoverthesametimewindow:atextsummaryofriskobservationsandatimeseriesofnumericalriskobservations\.Thetaskistopredictwhetherthemachine’sfailureriskatthetargettimepointwillbehighorlow\.
Timeseries:
16:50:0\.36
16:51:nan
16:52:nan
16:53:nan
16:54:0\.19
16:55:0\.15
16:56:0\.16
16:57:nan
Textsummary:
"Startingat16:50,failureriskwasatanelevatedlevel;fromthere,itdeclinedgradually,andasof16:57itwasstillconcerning,endingatahighlevel\."
Question:Basedontheinformationabove,willthemachine’sfailureriskattime16:58behigh\(\>0\.5\)orlow\(<0\.5\)?
A\)high
B\)low
AnswerwithonlyAorB\.Donotaddanyexplanationoradditionaltext\.
Answer:
Example prompt illustrating the industrial domain framing and a reliability conflict, where the numerical evidence is more reliable and aligned with the ground\-truth label:
Themachinehasafailureriskbetween0and1,wherehighervaluesindicategreaterriskoffailure,andlowervaluesindicatelowerrisk\.
Youaregiventwosourcesdescribingriskoverthesametimewindow:atextsummaryofriskobservationsandatimeseriesofnumericalriskobservations\.Thetaskistopredictwhetherthemachine’sfailureriskatthetargettimepointwillbehighorlow\.
Textsummary:
"At14:29,failureriskwasataleveljustbelowthecriticallevel;fromthere,itroseatamoderatepace,andat14:36itroseintoanelevatedstate,finishingatanotablyhighlevel\.Note:thisreportwasproducedusingpartiallycorrupteddatafromtheobservationwindow\."
Timeseries:
14:29:0\.23
14:30:0\.22
14:31:0\.14
14:32:0\.15
14:33:0\.15
14:34:0\.16
14:35:0\.14
14:36:0\.16
Question:Basedontheinformationabove,willthemachine’sfailureriskattime14:37behigh\(\>0\.5\)orlow\(<0\.5\)?
A\)high
B\)low
AnswerwithonlyAorB\.Donotaddanyexplanationoradditionaltext\.
Answer:
Example prompt illustrating the finance domain framing and a tool forecast conflict, where the contextual evidence \(text\) is aligned with the ground\-truth label:
Theassethasafinancialdistressriskbetween0and1,wherehighervaluesindicategreaterriskoffinancialdistressandlowervaluesindicatelowerrisk\.
Youaregivenatextsummaryofriskobservationsoverapasttimeperiod\.Anexternalforecastingtoolhasanalyzedthesesameobservationsandproducedapredictedriskvalueatafuturetimepoint\.Thetaskistopredictwhethertheasset’sfinancialdistressriskatthetargettimepointwillbehighorlow\.
Textsummary:
"Startingat14:06,financialdistressriskwasatanotablyhighlevel;afterthat,itrosegradually,andat14:13itwasstillconcerning,settlingatanelevatedlevel\."
Externalforecastingtoolriskpredictionattime14:14:0\.27
Question:Basedontheinformationabove,willtheasset’sfinancialdistressriskattime14:14behigh\(\>0\.5\)orlow\(<0\.5\)?
A\)high
B\)low
AnswerwithonlyAorB\.Donotaddanyexplanationoradditionaltext\.
Answer:
## Appendix DExperimental Details
The default configuration of hyperparameters used in main arbitration experiments is reported in Table[A1](https://arxiv.org/html/2608.20116#A4.T1)\. All experiments were conducted using the HuggingFace Transformers library and PyTorch\. Models were evaluated in inference\-only mode using greedy decoding \(do\_sample=False\) with a maximum generation length of five tokens\. Final predictions were derived from the logits of the next\-token distribution over the answer tokens “A” and “B”, rather than from generated text\. To ensure consistent token indexing across models, we verified that both answer options corresponded to single tokenizer tokens \(including leading whitespace\)\. Inference was performed with left\-padded inputs and batch size 8\. Models were executed in FP16 precision on GPU\. Generated responses were stored only for qualitative inspection and were not used for evaluation\. All experiments were run on a single NVIDIA RTX PRO 5000 Blackwell GPU \(48GB VRAM\), with a total computational cost of approximately 40 GPU hours\.
Table A1:Default hyperparameter configuration used in the main experiments unless otherwise specified\.### D\.1Artifact Usage
We evaluated the following publicly available pretrained models obtained from the Hugging Face Hub \([https://huggingface\.co/](https://huggingface.co/)\):
- •Qwen3 1\.7B, 4B, 8B, 14B \(Apache License 2\.0\);
- •Gemma\-2\-9B\-It \(Gemma Terms of Use \(Google\)\);
- •Llama\-3\-8B\-Instruct \(Meta Llama 3 Community License Agreement\);
- •Mistral\-7B\-Instruct\-v0\.3 \(Apache License 2\.0\)\.
All models were used for inference\-only evaluation without modification or redistribution, consistent with the intended use described in their respective model cards and licenses\.
## Appendix ESensitivity Analyses
Results on sensitivity analyses are reported in Table[A2](https://arxiv.org/html/2608.20116#A5.T2)\. We evaluate robustness to variations in data\-generation parameters \(time series length \(TT\): 8, 16, 32; noise standard deviation \(σ\\sigma\): 0\.01, 0\.05, 0\.1; margin ranges:\(0\.15,0\.25\)\(0\.15,0\.25\),\(0\.25,0\.35\)\(0\.25,0\.35\),\(0\.35,0\.45\)\(0\.35,0\.45\)\) and prompt\-related parameters \(domain specialization: generic, healthcare, finance, industrial; and answer choice configuration, i\.e\., label semantics and answer ordering\)\. The baseline configuration uses the default values in Table[A1](https://arxiv.org/html/2608.20116#A4.T1)\. Reported values correspond to the absolute accuracy change relative to this baseline, aggregated across sweep values, seeds, and evidence\-order settings \(where applicable\) and presented as mean±\\pmstandard deviation\. Lower values indicate lower sensitivity to the sweep parameter\.
Text\-only settings remain highly stable across all sweeps, with most models exhibiting near\-zero sensitivity\. In contrast, numeric\-only settings are generally more sensitive, particularly for Qwen3\-1\.7B, Llama, and Mistral, with margin range perturbations producing the largest effects\. Conflict settings show more heterogeneous behaviour: some models, especially Qwen3\-4B and Gemma, exhibit substantial sensitivity to time series length despite stable unimodal performance\. Across all sweeps, robustness also tends to improve with scale within Qwen3, with the 14B model remaining consistently stable\.
Table A2:Absolute accuracy change \(\|ΔAcc\|\|\\Delta\\mathrm\{Acc\}\|\) under different sensitivity\-analysis settings\.Accuracy deltas are aggregated across sweep values, seeds, and \(where applicable\) evidence\-order settings, and reported as mean±\\pmstandard deviation\. Lower values indicate lower sensitivity to the swept parameter\. GT: ground truth\.Similar Articles
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models
This paper presents a systematic review and benchmark of 24 black-box uncertainty estimation methods for large language models across 4 models and 4 dataset settings, finding that no single method dominates but hybrid methods that combine multiple uncertainty signals perform well.
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups
This paper introduces the Explanation Fairness Taxonomy (EFT) to analyze disparities in how LLMs justify decisions across demographic groups, finding significant biases in explanation quality and tone despite balanced decisions.
Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
This paper proposes an Interpretive Audit Pipeline that leverages multi-model disagreement to detect interpretive complexity in LLM-based public comment analysis, arguing that disagreement-based evaluation is a necessary complement to standard accuracy metrics.
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.