Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
Summary
This NeurIPS 2026 workshop paper introduces Long-Transduction, a controlled diagnostic benchmark measuring how well models sustain stateful, context-dependent operations over long generations. Evaluating seven open-weight models, it finds severe degradation when scaling context length (up to 128K), varying input format, and increasing local task complexity.
View Cached Full Text
Cached at: 10/01/26, 09:42 AM
# Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
Source: [https://arxiv.org/html/2609.38712](https://arxiv.org/html/2609.38712)
\\workshoptitle
2nd Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks \(NeurIPS 2026\)\. Correspondence emails should be sent to: jwillette@nvidia\.com
###### Abstract
Long\-horizon agentic workflows require models to sustain repeated state\-dependent actions all while the context grows, sub\-task complexity changes, and new data arrives\. Each situation represents an independent axis along which an agent may fail\. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs\. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds\. We introduce Long\-Transduction, a controlled diagnostic that tests a model’s ability to stay on task during long generation while continuously reading, mutating, and outputting input\-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations\. Long\-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis\. We evaluate seven open\-weight models, finding a 62\.8% decrease when scaling context length from 4\-128K, a 36\.5% decrease when varying input format, and a 39\.9% decrease by increasing local task complexity\. Together, these failures represent critical liabilities in long\-horizon agentic workflows\.
## 1Long\-horizon Execution Needs Its Own Diagnostic
Consider an accounts\-payable agent reconciling a 20,000\-line invoice export against purchase orders\. For each line it must retrieve the right vendor terms, compute an adjustment, and update the matching ledger entry\. If the final ledger is wrong, a task\-level score cannot tell whether the agent retrieved the wrong terms, miscomputed the adjustment, dropped one invoice, or shifted every later decision onto the wrong record\. Benchmarks spanning real\-world questions, web interactions, code repair, and tool\-mediated goals measure whether the overall task succeeded\([Mialon et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib5);[Zhou et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib6);[Drouin et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib7);[Jimenez et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib8);[Yao et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib9)\); they do not, however, isolate this sustained\-execution failure in a way that can be measured\.
Long\-context benchmarks test complementary retrieval and understanding abilities\([Bai et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib3);[Hsieh et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib1);[Yen et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib4)\)\. RULER, for example, extends needle\-in\-a\-haystack retrieval with multiple keys, tracking, and aggregation, but still produces answers short relative to input context\. In a long horizon retrieval and transformation, the output itself is a long stateful trajectory that requires careful bookkeeping and attention to detail\. Long\-Transduction implements this paradigm by holding the atomic operations simple which provides a way to score every required item in the output\. A model that cannot stay aligned while repeatedly executing these elementary operations cannot be expected to stay on task through more complex actions in a long\-horizon agentic workflow such as Browsecomp\([Wei et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib2)\)\. We contribute \(i\) a controlled design that independently varies length, local difficulty, and input format; \(ii\) exact per\-record and position\-resolved scoring that reveals what fails and when during generation; and \(iii\) a balanced seven\-model evaluation, showing that supported context length does not certify reliable completion\.
Figure 1:The universal read, mutate, output loop underlying both transduction tasks and agentic workflows\. At each step, a system reads relevant state from its context, applies a meaningful mutation, and appends the result to its output stream\. Long\-transduction tests whether models can repeat this loop reliably across a growing long\-horizon workflow\.
## 2Long\-transduction
Figure[1](https://arxiv.org/html/2609.38712#S1.F1)shows the read, mutate, output paradigm that Long\-transduction isolates\. We instantiate this loop with four controlled task families, three input formats, and four difficulty settings\. The*Arithmetic*task evaluates addition/subtraction expressions line\-by\-line and varies difficulty by including 2, 4, 8, or 16 operands per record\. The*UUID sorting*task sorts lists of identifiers line\-by\-line and varies difficulty by including 2, 4, 8, or 16 items per record\. The*Variable Lookup*task resolves key pairs against dictionaries line\-by\-line and varies difficulty by scaling total dicitionary size to 8, 32, 128, or 256 total entries\. The*Table Transformation*task applies a transformation to a CSV table and varies difficulty by increasing the complexity of the transformation\. For examples of inputs from each task family as well as a more detailed explanation, please see Appendix[A](https://arxiv.org/html/2609.38712#A1)\.
The first three families each have three matched input formats that test position tracking ability\. First, ordered records with numeric IDs give every input line a stable lookup key as an integer index\. Second, shuffling the input record ID’s and requiring monotonically increasing ID’s in the output adds aglobalreordering requirement while preserving thelocal taskof a transformation operation on each row\. Third, removing IDs provides the least structure: the model must dynamically track the context position based solely on the current output position and surrounding content\. An omission could shift all future positions, leading to failure\. Together, these settings evaluate local/global transformation ability, retrieval, and position tracking; all fundamental tasks in long horizon workflows\.
Every one of the 4 tasks is evaluated with three different input formats, four difficulty settings, six context budgets, and five document samples\. Thus,12×4×6×5=1,44012\\times 4\\times 6\\times 5=1\{,\}440documents per model, including 240 at each context length horizon\. A single document can contain hundreds to thousands of individually scored items\. Unless otherwise specified, we keep the document, not each item, as the independent scored unit\. Because required output grows approximately one\-for\-one with input, we report nominal input\-plus\-output horizons of 4K, 8K, 16K, 32K, 64K, and 128K tokens\. For example, a 4K transduction document is approximately 2K input tokens and 2K output tokens\. At the largest tier, the data average 65,374 input and 61,240 required output tokens\.
Every reported item accuracy is the fraction of required output records within a document that are exactly correct\. We average records within each document and then give every task\-difficulty cell equal weight, so settings containing more records do not dominate\.
## 3Open\-weight Model Evaluation
We test eight open\-weight checkpoints: Nemotron 3 30B and Super 120B\([Blakeman et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib10)\), Qwen3\.5 35B\-A3B and 122B\-A10B\([Yang et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib11)\), DeepSeek V4 Flash\([Xu et al\., 2026](https://arxiv.org/html/2609.38712#bib.bib12)\), Kimi Linear 48B\-A3B\([Team et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib13)\), Falcon\-H1 34B\([Zuo et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib14)\), and Olmo 3\.1 32B\([Olmo et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib15)\)\. Seven span the full grid; Olmo stops at 32K nominal tokens \(16K input\) because of its context limit\. Closed\-weight models frequently declined the task, terminated early, or requested clarification in pilot runs, so we omit them from quantitative comparisons\. All models receive identical instructions with greedy sampling\. Prompts explicitly say “Do not think\.” We disable thinking/reasoning when an ‘off’ or ‘none’ setting exists and otherwise use the lowest available thinking budget \(Deepseek only\); all stored rollouts report zero reasoning tokens\. Model\-specific tokenizers produce somewhat different observed input lengths \(Appendix[D](https://arxiv.org/html/2609.38712#A4)\)\.
\(a\)Reliability by context horizon\.\(b\)Individual task performance at 128K\.
Figure 2:Fig\.[2\(a\)](https://arxiv.org/html/2609.38712#S3.F2.sf1)shows how accuracy degrades as context length increases\. The result is averaged over all tasks, input formats, and difficulties\. Fig\.[2\(b\)](https://arxiv.org/html/2609.38712#S3.F2.sf2)shows individual task performance at 128K averaged over input formats and difficulties\. Olmo\* ends at 32K because of its context limit\.
## 4Results
### Horizon exposes large reliability losses\.
Figure[2](https://arxiv.org/html/2609.38712#S3.F2)[2\(a\)](https://arxiv.org/html/2609.38712#S3.F2.sf1)shows monotonic degradation for every non\-degenerate model\. DeepSeek leads throughout, but overall accuracy falls from 0\.909 at 4K nominal tokens to 0\.554 at 128K\. Qwen\-122B falls from 0\.883 to 0\.441, only 0\.030 above Qwen\-35B at 128K\. Nemotron Super falls from 0\.711 to 0\.138\. Falcon falls from 0\.719 to 0\.036 and Olmo from 0\.419 to 0\.083 \(at 32K\)\. Supported context and model scale therefore do not certify reliable exhaustive execution\.Averaged over all models, there is a 62\.8% relative decrease in performance when scaling context from 4K to 128K\.
### Individual task reliability\.
Task\-family results diverge sharply \(Fig\.[2](https://arxiv.org/html/2609.38712#S3.F2)[2\(b\)](https://arxiv.org/html/2609.38712#S3.F2.sf2)\): at 128K nominal tokens, DeepSeek reaches 0\.961 on arithmetic but only 0\.262 on table transformation; Qwen\-122B combines 0\.952 arithmetic with 0\.233 UUID sorting and 0\.170 table accuracy\. For exhaustive workflows, an incorrect or missing record forces reconciliation or retry, so exact completion estimates the fraction requiring no repair\. At 128K this is only 41/240 for DeepSeek \(17\.1%\), 11/240 for Qwen\-122B \(4\.6%\), and 2/240 for Qwen\-35B \(0\.8%\); Appendix Table[2](https://arxiv.org/html/2609.38712#A3.T2)reports all complete\-grid models\.
\(a\)Arithmetic\(b\)UUID sorting\(c\)Variable lookup
Figure 3:Format sensitivity at 128K nominal tokens, averaged over four difficulties\. Surprisingly, ordered transformations without row ID’s scores worse than shuffling the numbered input rows\. Removing IDs removes the stable row lookup key\. Olmo\* only includes up to 32K\. For the corresponding table transform plot, see Fig\.[9\(b\)](https://arxiv.org/html/2609.38712#A5.F9.sf2)\.
### Format reveals conditional failures\.
Fig\.[3\(a\)](https://arxiv.org/html/2609.38712#S4.F3.sf1)shows a level of robustness to changes in the input format\. Averaged over all models, there is a 16\.2% decrease when moving from ‘Ordered \+ IDs’ to ‘Ordered, No IDs’ However, for Figs\.[3](https://arxiv.org/html/2609.38712#S4.F3)[3\(b\)](https://arxiv.org/html/2609.38712#S4.F3.sf2)\-[3\(c\)](https://arxiv.org/html/2609.38712#S4.F3.sf3), we see that UUID sorting and variable lookup are much more sensitive to input format, with the unstructured ‘no ID’ variant realizing a 64\.3% and 66\.4% decrease respectively\.This indicates that the ability to track a content position in the context is not invariant to the current task which is being performed\.
\(a\)Arithmetic\(b\)UUID sort\(c\)Table Transform
Figure 4:Overall Accuracy at 128K decreases with an increasing local task complexity\. Metrics are averaged over input format\. Increasing local task complexity degrades overall performance\. Olmo\* data only up until 32K\. For the corresponding variable lookup plot, see Fig\.[9\(a\)](https://arxiv.org/html/2609.38712#A5.F9.sf1)
### Local difficulty\.
Fig\.[4](https://arxiv.org/html/2609.38712#S4.F4)[4\(a\)](https://arxiv.org/html/2609.38712#S4.F4.sf1)\-[4\(c\)](https://arxiv.org/html/2609.38712#S4.F4.sf3)show the effect of adding more local task complexity by increasing summands, sorting items, or permuted rows/columns\. Every non\-degenerate model declines as local work increases from the easiest to hardest\. Appendix Fig\.[6](https://arxiv.org/html/2609.38712#A3.F6)expands these marginals into all 12 task\-specific curves at 128K and includes the arithmetic task\.
\(a\)Answer Correct\(b\)Self\-consistency
Figure 5:Calculating individual scored items throughout generation for the ‘Ordered, no IDs’ format on the UUID sorting task averaged over all context lengths and difficulties\. Models tend to stay self\-consistent longer than they are able to copy information from context\.
### Models understand the task, but get lost\.
Due to the synthetic structure of the task, we are able to score individual items in the output stream as a function of output position\. Then, we can ask an interesting quesion: Even if the model didn’t output the correct answer, did it at least stay self\-consistent? That is, for UUID sorting, did it output a properly sorted list of ‘something?’ Figs\.[5\(a\)](https://arxiv.org/html/2609.38712#S4.F5.sf1)and[5\(b\)](https://arxiv.org/html/2609.38712#S4.F5.sf2)show that self\-consistency is always significantly higher than per\-position accuracy, indicating the model understands the task, but fails to read the correct problem from the context\.
## 5Conclusion
These results have important implications for long\-horizon agents\. Long\-transduction shows that nominal capacity can hide failures caused by input formats, local difficulty, and context length\. Currently, long workflows should be sure to itemize inputs with stable IDs, checkpoint chunks of outputs with IDs, and split tasks into small and less complex units of work to avoid presenting difficult tasks in conjunction with long context and confusing input formats\. Furthermore, future models should focus more effort on mitigating these failure modes at train\-time so that agents may handle complex, less\-structured, long\-horizon tasks\.
## References
- Baiet al\.\(2024\)Y\. Bai, X\. Lv, J\. Zhang, H\. Lyu, J\. Tang, Z\. Huang, Z\. Du, X\. Liu, A\. Zeng, L\. Hou,et al\.Longbench: a bilingual, multitask benchmark for long context understanding\.InProceedings of the 62nd annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 3119–3137\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Blakemanet al\.\(2025\)A\. Blakeman, A\. Grattafiori, A\. Basant, A\. Gupta, A\. Khattar, A\. Renduchintala, A\. Vavre, A\. Shukla, A\. Bercovich, A\. Ficek,et al\.Nvidia nemotron 3: efficient and open intelligence\.arXiv preprint arXiv:2512\.20856\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Drouinet al\.\(2024\)A\. Drouin, M\. Gasse, M\. Caccia, I\. H\. Laradji, M\. Del Verme, T\. Marty, L\. Boisvert, M\. Thakkar, Q\. Cappart, D\. Vazquez,et al\.WorkArena: how capable are web agents at solving common knowledge work tasks?\.arXiv preprint arXiv:2403\.07718\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Hsiehet al\.\(2024\)C\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, Y\. Zhang, and B\. GinsburgRULER: what’s the real context size of your long\-context language models?\.arXiv preprint arXiv:2404\.06654\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. NarasimhanSwe\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 54107–54157\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Mialonet al\.\(2024\)G\. Mialon, C\. Fourrier, T\. Wolf, Y\. LeCun, and T\. ScialomGaia: a benchmark for general ai assistants\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 9025–9049\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Olmoet al\.\(2025\)T\. Olmo, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison,et al\.Olmo 3\.arXiv preprint arXiv:2512\.13961\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Teamet al\.\(2025\)K\. Team, Y\. Zhang, Z\. Lin, X\. Yao, J\. Hu, F\. Meng, C\. Liu, X\. Men, S\. Yang, Z\. Li,et al\.Kimi linear: an expressive, efficient attention architecture\.arXiv preprint arXiv:2510\.26692\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Weiet al\.\(2025\)J\. Wei, Z\. Sun, S\. Papay, S\. McKinney, J\. Han, I\. Fulford, H\. W\. Chung, A\. T\. Passos, W\. Fedus, and A\. GlaeseBrowsecomp: a simple yet challenging benchmark for browsing agents\.arXiv preprint arXiv:2504\.12516\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Xuet al\.\(2026\)A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.Deepseek\-v4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Yaoet al\.\(2024\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Yenet al\.\(2025\)H\. Yen, T\. Gao, M\. Hou, K\. Ding, D\. Fleischer, P\. Izsak, M\. Wasserblat, and D\. ChenHELMET: how to evaluate long\-context models effectively and thoroughly\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 98914–98965\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Zuoet al\.\(2025\)J\. Zuo, M\. Velikanov, I\. Chahed, Y\. Belkada, D\. E\. Rhayem, G\. Kunsch, H\. Hacid, H\. Yous, B\. Farhat, I\. Khadraoui,et al\.Falcon\-h1: a family of hybrid\-head language models redefining efficiency and performance\.arXiv preprint arXiv:2507\.22448\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
## Appendix ABenchmark details and examples
### Arithmetic\.
The ordered numbered variant isntructs the model to transform the following input:
\[1\]5\+6\[2\]2\+4\-1\[3\]4\+8\-3\+2into the following output:
\[1\]5\+6=11\[2\]2\+4\-1=5\[3\]4\+8\-3\+2=11\.The shuffled variant presents the same input records with shuffled rows, and requires an ascending \(sorted\) output index; the unnumbered variant removes the row index identifiers \(such as\[1\], \[2\], \[3\]and requires the model to keep track of the global position without an index to use as a lookup key\. The four difficulty levels are created by making each summand consist of 2,4,8, or 16 terms\.
### UUID sorting\.
Each record contains short lists of hexadecimal identifiers:
\[1\]c0a8e1d2,a1b2c3d4,b1c2d3e4\[2\]f0e1d2c3,01234567and the model must output the same list in sorted order:
\[1\]a1b2c3d4,b1c2d3e4,c0a8e1d2\[2\]01234567,f0e1d2c3\.The verifier marks the record correct only when the emitted identifiers exactly match the expected sorted sequence\. Similar to arithmetic, this task includes Ordered, Shuffled, and No ID variants\. The four levels of difficulty are created by including 2,4,8, or 16 items in each hexadecimal list\.
### Variable lookup\.
A variable definition pool such as
a3f=quickb91=fox03c=quiete18=harbor
is followed by expressions such as:
The expected output transformation of each record is:
\[1\]quick fox\[2\]quiet harbor
Like the arithmetic and UUID variants, variable lookup also includes ordered, shuffled, and no ID variants\. The four difficulty settings are achieved by changing the definition\-pool size to one of 8, 32, 128, or 256 independently of the number of output records\.
### CSV Table Transformation\.
The table transformation task follows a slightly different pattern than the previous variants\. For an input record corresponding to a CSV table
,\[C0\],\[C1\],\[C2\]\[R0\],a,b,c\[R1\],d,e,f\[R2\],g,h,i
CSV tables do not have a clear input format that can be indexed by row, as the previous tasks do\. Therefore, we construct the three analogous format variants by supplying instructions to perform an operation on either the rows and columns or the individual cells of the table\. The operations are as follows:
1. 1\.Row/Column permutation with homogeneous cells: each cell contains homogeneous length four\-digit integers and the rows and columns must be permuted according to a given reordering\.
2. 2\.Row/Column permutation with heterogeneous cells: each cell contains variable\-length UUID fragments ranging in length from 1\-36 hexadecimal digits and the rows and columns must be permuted according to a given reordering\.\.
3. 3\.KV resolution: each cell contains a variable expression that needs to be de referenced\. See the description below\.
For table transformation tasks requiring a row/column permutation, the model receives a new row and column order such as\[R2\],\[R0\],\[R1\]and\[C1\],\[C0\],\[C2\], the output is then expected to be:
,\[C1\],\[C0\],\[C2\]\[R2\],h,g,i\[R0\],b,a,c\[R1\],e,d,f\.
CSV tasks with permutation increase difficulty by requiring one of 20%, 40%, 80%, or 100% of rows and columns permuted\.
### CSV KV resolution\.
Given adjective and noun tables \(e\.g\.,a0=quick,n0=fox\) and cells containinga0\+n0, the model must preserve the CSV structure while replacing every cell with the resolved phrase\. Difficulty and context length control how many distinct adjective and noun keys are defined by variables as displayed in the following table:
Table 1:Number of active adjective or noun variables in each dictionary\. A value of ‘2’ means that there are 2 adjectives and 2 nouns\.
## Appendix BScoring
Each required output record \(output row or csv cell\) is marked correct only when it exactly matches the expected answer\. We first average record correctness within a generated document, then average equally over other dimensions such as tasks, difficulty, format, or context length\. Exact document completion is an indicator equal to one only when every required record in that document is correct; the reported rate in Table[2](https://arxiv.org/html/2609.38712#A3.T2)averages this indicator over the 240 documents at the 128K horizon\.
For numbered streams, the verifier uses the emitted numeric ID to match each output to its expected record\. A missing ID therefore creates a localized zero without shifting later matches\. For unnumbered streams, non\-empty output lines are matched ordinally, so one omission can shift subsequent matches by design\. Position plots \(Figs\.[5](https://arxiv.org/html/2609.38712#S4.F5),[7](https://arxiv.org/html/2609.38712#A3.F7)\) map expected record indices into 20 equal normalized bins, average correctness within each document and bin, and then macro\-average documents\. This procedure prevents easy settings with more short records from dominating the earlier bins in a curve\.
## Appendix CAdditional results
Table 2:Exact document completion at 128K nominal tokens\. A document needs no repair only when every required output item is correct\.\(a\)Arithmetic: ordered \+ IDs\(b\)Arithmetic: shuffled \+ IDs\(c\)Arithmetic: no IDs\(d\)UUID sorting: ordered \+ IDs\(e\)UUID sorting: shuffled \+ IDs\(f\)UUID sorting: no IDs\(g\)Lookup: ordered \+ IDs\(h\)Lookup: shuffled \+ IDs\(i\)Lookup: no IDs\(j\)Homogeneous table permutation\(k\)Heterogeneous table permutation\(l\)Table KV lookup
Figure 6:All local\-difficulty curves at 128K nominal tokens\. Olmo\* has no 128K point\. These figures complement what is shown in Fig\.[4](https://arxiv.org/html/2609.38712#S4.F4)which averages over all context lengths and formats\.\(a\)Arithmetic: ordered \+ IDs\(b\)Arithmetic: shuffled \+ IDs\(c\)Arithmetic: no IDs\(d\)UUID sorting: ordered \+ IDs\(e\)UUID sorting: shuffled \+ IDs\(f\)UUID sorting: no IDs\(g\)Lookup: ordered \+ IDs\(h\)Lookup: shuffled \+ IDs\(i\)Lookup: no IDs\(j\)Homogeneous table permutation\(k\)Heterogeneous table permutation\(l\)Table KV lookup
Figure 7:Accuracy across the required generation at 128K nominal tokens for all 12 task variants\. Complete\-grid curves average four difficulties and five document samples per setting after within\-document binning; Olmo\* has no 128K point\. Figure columns compare input formats, while rows corresponds to tasks\.Figure 8:Available task\-model results at 128K nominal tokens\. Complete\-grid cells average four difficulties and five document samples per setting; gray dashes mark missing Olmo\* cells\. This view complements Fig\.[6](https://arxiv.org/html/2609.38712#A3.F6): the heatmap compares absolute performance while the curves show difficulty sensitivity\.Figure[8](https://arxiv.org/html/2609.38712#A3.F8)reports every task/format at the longest context length \(128K\)\.
## Appendix DReproducibility and Release
The frozen benchmark grid contains 4 tasks, 3 input formats, 4 difficulty settings, 6 nominal input\-plus\-output horizons, and 5 document samples per task:4×3×4×6×5=1,4404\\times 3\\times 4\\times 6\\times 5=1\{,\}440documents\. Each of seven primary models covers this identical grid\. Olmo covers four full tiers through 32K; its context limit prevents evaluation at longer horizons\. At the largest tier, the fixedtiktokencl100k\_basepreparation tokenizer estimates 65,374 mean prompt tokens and 61,240 mean required\-output tokens\. Observed model\-specific input counts differ because tokenizers differ\.
Inference uses greedy sampling with zero temperature\. Every prompt explicitly instructs the model not to think\. We set explicit thinking/reasoning to off or none when supported and otherwise use the lowest available effort \(low\)\. All plotted response\-usage records report zero reasoning tokens\. Outputs are permitted up to 1\.5 times the target input budget\.
Upon acceptance, we will release a fully seeded data generator, exact verifier, and evaluation scripts\.
## Appendix ELimitations
LongTransduction is intentionally synthetic and deterministic\. It does not test planning, tool selection, multi\-turn interaction, recovery, permissions, human escalation, or direct safety behavior\. Its records resemble enterprise data\-processing primitives but are not a realistic enterprise environment\. Positional scoring of unnumbered tasks also makes omissions cascade by design; this measures format fragility as well as local competence\. However, this synthetic setting mimics fundamental primitives seen in real\-world environments\.
\(a\)Arithmetic\(b\)Table Transform
Figure 9:Fig\.[9\(a\)](https://arxiv.org/html/2609.38712#A5.F9.sf1)corresponds to the figures shown in Fig\.[6](https://arxiv.org/html/2609.38712#A3.F6)\. Fig\.[9\(b\)](https://arxiv.org/html/2609.38712#A5.F9.sf2)corresponds to the figures shown in Fig\.[3](https://arxiv.org/html/2609.38712#S4.F3)\.Similar Articles
Testing Agents on Long-Horizon Terminal Work (GitHub Repo)
Long-Horizon Terminal-Bench (LHTB) is a 46-task benchmark for evaluating LLM agents on sustained terminal work over hundreds of steps, revealing that even the best models solve only ~28% of tasks.
@dair_ai: Outstanding paper on long-horizon agents. (bookmark it) Similar to humans, how do you make agents persist on a difficul…
AutoLab is a new benchmark evaluating 17 frontier models on 36 expert-curated long-horizon tasks (system optimization, model development, CUDA kernels, puzzles), finding that persistence—not initial attempt quality—is the dominant predictor of success. Claude-opus-4.6 led all categories, while most other models terminated prematurely or exhausted budgets with minimal progress.
@rohanpaul_ai: Long-horizon agent reliability has not arrived yet with better models. On WeaveBench's 114 hybrid GUI-CLI tasks, the be…
This paper argues that long-horizon AI agent reliability is lacking despite better models, as seen on WeaveBench with only 41.2% pass rate, and proposes the LongHorizon-Harness to manage task state for improved performance.
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Introduces Long-Horizon-Terminal-Bench, a benchmark of 46 long-horizon terminal tasks with dense reward-based grading, evaluating AI agents on planning, long-context, and debugging. Even the strongest model achieves only 15.2% pass@1, showing significant room for improvement.
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
Introduces LongDS, a benchmark for evaluating LLM agents on long-horizon, multi-turn data analysis tasks. Evaluations show that even the best models achieve only 48.45% accuracy, with performance dropping sharply over turns, highlighting that maintaining analytical state is the key bottleneck.