保持专注:测试长时程智能体可靠性的基础
摘要
这篇 NeurIPS 2026 工作坊论文提出了 Long-Transduction——一个受控的诊断基准,用于衡量模型在长生成过程中维持有状态、依赖上下文的操作的能力。在对七个开放权重模型的评估中,论文发现,当上下文长度扩展(最高达 128K)、输入格式变化以及局部任务复杂度增加时,模型性能会出现严重退化。
查看缓存全文
缓存时间: 2026/10/01 09:42
# Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
Source: [https://arxiv.org/html/2609.38712](https://arxiv.org/html/2609.38712)
\\workshoptitle
2nd Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks \(NeurIPS 2026\)\. Correspondence emails should be sent to: jwillette@nvidia\.com
###### Abstract
Long\-horizon agentic workflows require models to sustain repeated state\-dependent actions all while the context grows, sub\-task complexity changes, and new data arrives\. Each situation represents an independent axis along which an agent may fail\. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs\. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds\. We introduce Long\-Transduction, a controlled diagnostic that tests a model’s ability to stay on task during long generation while continuously reading, mutating, and outputting input\-context dependent operations such as arithmetic, sorting, variable lookups, and table transformations\. Long\-Transduction evaluation independently varies local task complexity, input data formatting, and context length isolate failures along each axis\. We evaluate seven open\-weight models, finding a 62\.8% decrease when scaling context length from 4\-128K, a 36\.5% decrease when varying input format, and a 39\.9% decrease by increasing local task complexity\. Together, these failures represent critical liabilities in long\-horizon agentic workflows\.
## 1Long\-horizon Execution Needs Its Own Diagnostic
Consider an accounts\-payable agent reconciling a 20,000\-line invoice export against purchase orders\. For each line it must retrieve the right vendor terms, compute an adjustment, and update the matching ledger entry\. If the final ledger is wrong, a task\-level score cannot tell whether the agent retrieved the wrong terms, miscomputed the adjustment, dropped one invoice, or shifted every later decision onto the wrong record\. Benchmarks spanning real\-world questions, web interactions, code repair, and tool\-mediated goals measure whether the overall task succeeded\([Mialon et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib5);[Zhou et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib6);[Drouin et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib7);[Jimenez et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib8);[Yao et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib9)\); they do not, however, isolate this sustained\-execution failure in a way that can be measured\.
Long\-context benchmarks test complementary retrieval and understanding abilities\([Bai et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib3);[Hsieh et al\., 2024](https://arxiv.org/html/2609.38712#bib.bib1);[Yen et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib4)\)\. RULER, for example, extends needle\-in\-a\-haystack retrieval with multiple keys, tracking, and aggregation, but still produces answers short relative to input context\. In a long horizon retrieval and transformation, the output itself is a long stateful trajectory that requires careful bookkeeping and attention to detail\. Long\-Transduction implements this paradigm by holding the atomic operations simple which provides a way to score every required item in the output\. A model that cannot stay aligned while repeatedly executing these elementary operations cannot be expected to stay on task through more complex actions in a long\-horizon agentic workflow such as Browsecomp\([Wei et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib2)\)\. We contribute \(i\) a controlled design that independently varies length, local difficulty, and input format; \(ii\) exact per\-record and position\-resolved scoring that reveals what fails and when during generation; and \(iii\) a balanced seven\-model evaluation, showing that supported context length does not certify reliable completion\.
Figure 1:The universal read, mutate, output loop underlying both transduction tasks and agentic workflows\. At each step, a system reads relevant state from its context, applies a meaningful mutation, and appends the result to its output stream\. Long\-transduction tests whether models can repeat this loop reliably across a growing long\-horizon workflow\.
## 2Long\-transduction
Figure[1](https://arxiv.org/html/2609.38712#S1.F1)shows the read, mutate, output paradigm that Long\-transduction isolates\. We instantiate this loop with four controlled task families, three input formats, and four difficulty settings\. The*Arithmetic*task evaluates addition/subtraction expressions line\-by\-line and varies difficulty by including 2, 4, 8, or 16 operands per record\. The*UUID sorting*task sorts lists of identifiers line\-by\-line and varies difficulty by including 2, 4, 8, or 16 items per record\. The*Variable Lookup*task resolves key pairs against dictionaries line\-by\-line and varies difficulty by scaling total dicitionary size to 8, 32, 128, or 256 total entries\. The*Table Transformation*task applies a transformation to a CSV table and varies difficulty by increasing the complexity of the transformation\. For examples of inputs from each task family as well as a more detailed explanation, please see Appendix[A](https://arxiv.org/html/2609.38712#A1)\.
The first three families each have three matched input formats that test position tracking ability\. First, ordered records with numeric IDs give every input line a stable lookup key as an integer index\. Second, shuffling the input record ID’s and requiring monotonically increasing ID’s in the output adds aglobalreordering requirement while preserving thelocal taskof a transformation operation on each row\. Third, removing IDs provides the least structure: the model must dynamically track the context position based solely on the current output position and surrounding content\. An omission could shift all future positions, leading to failure\. Together, these settings evaluate local/global transformation ability, retrieval, and position tracking; all fundamental tasks in long horizon workflows\.
Every one of the 4 tasks is evaluated with three different input formats, four difficulty settings, six context budgets, and five document samples\. Thus,12×4×6×5=1,44012\\times 4\\times 6\\times 5=1\{,\}440documents per model, including 240 at each context length horizon\. A single document can contain hundreds to thousands of individually scored items\. Unless otherwise specified, we keep the document, not each item, as the independent scored unit\. Because required output grows approximately one\-for\-one with input, we report nominal input\-plus\-output horizons of 4K, 8K, 16K, 32K, 64K, and 128K tokens\. For example, a 4K transduction document is approximately 2K input tokens and 2K output tokens\. At the largest tier, the data average 65,374 input and 61,240 required output tokens\.
Every reported item accuracy is the fraction of required output records within a document that are exactly correct\. We average records within each document and then give every task\-difficulty cell equal weight, so settings containing more records do not dominate\.
## 3Open\-weight Model Evaluation
We test eight open\-weight checkpoints: Nemotron 3 30B and Super 120B\([Blakeman et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib10)\), Qwen3\.5 35B\-A3B and 122B\-A10B\([Yang et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib11)\), DeepSeek V4 Flash\([Xu et al\., 2026](https://arxiv.org/html/2609.38712#bib.bib12)\), Kimi Linear 48B\-A3B\([Team et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib13)\), Falcon\-H1 34B\([Zuo et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib14)\), and Olmo 3\.1 32B\([Olmo et al\., 2025](https://arxiv.org/html/2609.38712#bib.bib15)\)\. Seven span the full grid; Olmo stops at 32K nominal tokens \(16K input\) because of its context limit\. Closed\-weight models frequently declined the task, terminated early, or requested clarification in pilot runs, so we omit them from quantitative comparisons\. All models receive identical instructions with greedy sampling\. Prompts explicitly say “Do not think\.” We disable thinking/reasoning when an ‘off’ or ‘none’ setting exists and otherwise use the lowest available thinking budget \(Deepseek only\); all stored rollouts report zero reasoning tokens\. Model\-specific tokenizers produce somewhat different observed input lengths \(Appendix[D](https://arxiv.org/html/2609.38712#A4)\)\.
\(a\)Reliability by context horizon\.\(b\)Individual task performance at 128K\.
Figure 2:Fig\.[2\(a\)](https://arxiv.org/html/2609.38712#S3.F2.sf1)shows how accuracy degrades as context length increases\. The result is averaged over all tasks, input formats, and difficulties\. Fig\.[2\(b\)](https://arxiv.org/html/2609.38712#S3.F2.sf2)shows individual task performance at 128K averaged over input formats and difficulties\. Olmo\* ends at 32K because of its context limit\.
## 4Results
### Horizon exposes large reliability losses\.
Figure[2](https://arxiv.org/html/2609.38712#S3.F2)[2\(a\)](https://arxiv.org/html/2609.38712#S3.F2.sf1)shows monotonic degradation for every non\-degenerate model\. DeepSeek leads throughout, but overall accuracy falls from 0\.909 at 4K nominal tokens to 0\.554 at 128K\. Qwen\-122B falls from 0\.883 to 0\.441, only 0\.030 above Qwen\-35B at 128K\. Nemotron Super falls from 0\.711 to 0\.138\. Falcon falls from 0\.719 to 0\.036 and Olmo from 0\.419 to 0\.083 \(at 32K\)\. Supported context and model scale therefore do not certify reliable exhaustive execution\.Averaged over all models, there is a 62\.8% relative decrease in performance when scaling context from 4K to 128K\.
### Individual task reliability\.
Task\-family results diverge sharply \(Fig\.[2](https://arxiv.org/html/2609.38712#S3.F2)[2\(b\)](https://arxiv.org/html/2609.38712#S3.F2.sf2)\): at 128K nominal tokens, DeepSeek reaches 0\.961 on arithmetic but only 0\.262 on table transformation; Qwen\-122B combines 0\.952 arithmetic with 0\.233 UUID sorting and 0\.170 table accuracy\. For exhaustive workflows, an incorrect or missing record forces reconciliation or retry, so exact completion estimates the fraction requiring no repair\. At 128K this is only 41/240 for DeepSeek \(17\.1%\), 11/240 for Qwen\-122B \(4\.6%\), and 2/240 for Qwen\-35B \(0\.8%\); Appendix Table[2](https://arxiv.org/html/2609.38712#A3.T2)reports all complete\-grid models\.
\(a\)Arithmetic\(b\)UUID sorting\(c\)Variable lookup
Figure 3:Format sensitivity at 128K nominal tokens, averaged over four difficulties\. Surprisingly, ordered transformations without row ID’s scores worse than shuffling the numbered input rows\. Removing IDs removes the stable row lookup key\. Olmo\* only includes up to 32K\. For the corresponding table transform plot, see Fig\.[9\(b\)](https://arxiv.org/html/2609.38712#A5.F9.sf2)\.
### Format reveals conditional failures\.
Fig\.[3\(a\)](https://arxiv.org/html/2609.38712#S4.F3.sf1)shows a level of robustness to changes in the input format\. Averaged over all models, there is a 16\.2% decrease when moving from ‘Ordered \+ IDs’ to ‘Ordered, No IDs’ However, for Figs\.[3](https://arxiv.org/html/2609.38712#S4.F3)[3\(b\)](https://arxiv.org/html/2609.38712#S4.F3.sf2)\-[3\(c\)](https://arxiv.org/html/2609.38712#S4.F3.sf3), we see that UUID sorting and variable lookup are much more sensitive to input format, with the unstructured ‘no ID’ variant realizing a 64\.3% and 66\.4% decrease respectively\.This indicates that the ability to track a content position in the context is not invariant to the current task which is being performed\.
\(a\)Arithmetic\(b\)UUID sort\(c\)Table Transform
Figure 4:Overall Accuracy at 128K decreases with an increasing local task complexity\. Metrics are averaged over input format\. Increasing local task complexity degrades overall performance\. Olmo\* data only up until 32K\. For the corresponding variable lookup plot, see Fig\.[9\(a\)](https://arxiv.org/html/2609.38712#A5.F9.sf1)
### Local difficulty\.
Fig\.[4](https://arxiv.org/html/2609.38712#S4.F4)[4\(a\)](https://arxiv.org/html/2609.38712#S4.F4.sf1)\-[4\(c\)](https://arxiv.org/html/2609.38712#S4.F4.sf3)show the effect of adding more local task complexity by increasing summands, sorting items, or permuted rows/columns\. Every non\-degenerate model declines as local work increases from the easiest to hardest\. Appendix Fig\.[6](https://arxiv.org/html/2609.38712#A3.F6)expands these marginals into all 12 task\-specific curves at 128K and includes the arithmetic task\.
\(a\)Answer Correct\(b\)Self\-consistency
Figure 5:Calculating individual scored items throughout generation for the ‘Ordered, no IDs’ format on the UUID sorting task averaged over all context lengths and difficulties\. Models tend to stay self\-consistent longer than they are able to copy information from context\.
### Models understand the task, but get lost\.
Due to the synthetic structure of the task, we are able to score individual items in the output stream as a function of output position\. Then, we can ask an interesting quesion: Even if the model didn’t output the correct answer, did it at least stay self\-consistent? That is, for UUID sorting, did it output a properly sorted list of ‘something?’ Figs\.[5\(a\)](https://arxiv.org/html/2609.38712#S4.F5.sf1)and[5\(b\)](https://arxiv.org/html/2609.38712#S4.F5.sf2)show that self\-consistency is always significantly higher than per\-position accuracy, indicating the model understands the task, but fails to read the correct problem from the context\.
## 5Conclusion
These results have important implications for long\-horizon agents\. Long\-transduction shows that nominal capacity can hide failures caused by input formats, local difficulty, and context length\. Currently, long workflows should be sure to itemize inputs with stable IDs, checkpoint chunks of outputs with IDs, and split tasks into small and less complex units of work to avoid presenting difficult tasks in conjunction with long context and confusing input formats\. Furthermore, future models should focus more effort on mitigating these failure modes at train\-time so that agents may handle complex, less\-structured, long\-horizon tasks\.
## References
- Baiet al\.\(2024\)Y\. Bai, X\. Lv, J\. Zhang, H\. Lyu, J\. Tang, Z\. Huang, Z\. Du, X\. Liu, A\. Zeng, L\. Hou,et al\.Longbench: a bilingual, multitask benchmark for long context understanding\.InProceedings of the 62nd annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 3119–3137\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Blakemanet al\.\(2025\)A\. Blakeman, A\. Grattafiori, A\. Basant, A\. Gupta, A\. Khattar, A\. Renduchintala, A\. Vavre, A\. Shukla, A\. Bercovich, A\. Ficek,et al\.Nvidia nemotron 3: efficient and open intelligence\.arXiv preprint arXiv:2512\.20856\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Drouinet al\.\(2024\)A\. Drouin, M\. Gasse, M\. Caccia, I\. H\. Laradji, M\. Del Verme, T\. Marty, L\. Boisvert, M\. Thakkar, Q\. Cappart, D\. Vazquez,et al\.WorkArena: how capable are web agents at solving common knowledge work tasks?\.arXiv preprint arXiv:2403\.07718\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Hsiehet al\.\(2024\)C\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, Y\. Zhang, and B\. GinsburgRULER: what’s the real context size of your long\-context language models?\.arXiv preprint arXiv:2404\.06654\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. NarasimhanSwe\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 54107–54157\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Mialonet al\.\(2024\)G\. Mialon, C\. Fourrier, T\. Wolf, Y\. LeCun, and T\. ScialomGaia: a benchmark for general ai assistants\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 9025–9049\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Olmoet al\.\(2025\)T\. Olmo, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison,et al\.Olmo 3\.arXiv preprint arXiv:2512\.13961\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Teamet al\.\(2025\)K\. Team, Y\. Zhang, Z\. Lin, X\. Yao, J\. Hu, F\. Meng, C\. Liu, X\. Men, S\. Yang, Z\. Li,et al\.Kimi linear: an expressive, efficient attention architecture\.arXiv preprint arXiv:2510\.26692\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Weiet al\.\(2025\)J\. Wei, Z\. Sun, S\. Papay, S\. McKinney, J\. Han, I\. Fulford, H\. W\. Chung, A\. T\. Passos, W\. Fedus, and A\. GlaeseBrowsecomp: a simple yet challenging benchmark for browsing agents\.arXiv preprint arXiv:2504\.12516\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Xuet al\.\(2026\)A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.Deepseek\-v4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
- Yaoet al\.\(2024\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Yenet al\.\(2025\)H\. Yen, T\. Gao, M\. Hou, K\. Ding, D\. Fleischer, P\. Izsak, M\. Wasserblat, and D\. ChenHELMET: how to evaluate long\-context models effectively and thoroughly\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 98914–98965\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p2.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§1](https://arxiv.org/html/2609.38712#S1.p1.1)\.
- Zuoet al\.\(2025\)J\. Zuo, M\. Velikanov, I\. Chahed, Y\. Belkada, D\. E\. Rhayem, G\. Kunsch, H\. Hacid, H\. Yous, B\. Farhat, I\. Khadraoui,et al\.Falcon\-h1: a family of hybrid\-head language models redefining efficiency and performance\.arXiv preprint arXiv:2507\.22448\.Cited by:[§3](https://arxiv.org/html/2609.38712#S3.p1.1)\.
## Appendix ABenchmark details and examples
### Arithmetic\.
The ordered numbered variant isntructs the model to transform the following input:
\[1\]5\+6\[2\]2\+4\-1\[3\]4\+8\-3\+2into the following output:
\[1\]5\+6=11\[2\]2\+4\-1=5\[3\]4\+8\-3\+2=11\.The shuffled variant presents the same input records with shuffled rows, and requires an ascending \(sorted\) output index; the unnumbered variant removes the row index identifiers \(such as\[1\], \[2\], \[3\]and requires the model to keep track of the global position without an index to use as a lookup key\. The four difficulty levels are created by making each summand consist of 2,4,8, or 16 terms\.
### UUID sorting\.
Each record contains short lists of hexadecimal identifiers:
\[1\]c0a8e1d2,a1b2c3d4,b1c2d3e4\[2\]f0e1d2c3,01234567and the model must output the same list in sorted order:
\[1\]a1b2c3d4,b1c2d3e4,c0a8e1d2\[2\]01234567,f0e1d2c3\.The verifier marks the record correct only when the emitted identifiers exactly match the expected sorted sequence\. Similar to arithmetic, this task includes Ordered, Shuffled, and No ID variants\. The four levels of difficulty are created by including 2,4,8, or 16 items in each hexadecimal list\.
### Variable lookup\.
A variable definition pool such as
a3f=quickb91=fox03c=quiete18=harbor
is followed by expressions such as:
The expected output transformation of each record is:
\[1\]quick fox\[2\]quiet harbor
Like the arithmetic and UUID variants, variable lookup also includes ordered, shuffled, and no ID variants\. The four difficulty settings are achieved by changing the definition\-pool size to one of 8, 32, 128, or 256 independently of the number of output records\.
### CSV Table Transformation\.
The table transformation task follows a slightly different pattern than the previous variants\. For an input record corresponding to a CSV table
,\[C0\],\[C1\],\[C2\]\[R0\],a,b,c\[R1\],d,e,f\[R2\],g,h,i
CSV tables do not have a clear input format that can be indexed by row, as the previous tasks do\. Therefore, we construct the three analogous format variants by supplying instructions to perform an operation on either the rows and columns or the individual cells of the table\. The operations are as follows:
1. 1\.Row/Column permutation with homogeneous cells: each cell contains homogeneous length four\-digit integers and the rows and columns must be permuted according to a given reordering\.
2. 2\.Row/Column permutation with heterogeneous cells: each cell contains variable\-length UUID fragments ranging in length from 1\-36 hexadecimal digits and the rows and columns must be permuted according to a given reordering\.\.
3. 3\.KV resolution: each cell contains a variable expression that needs to be de referenced\. See the description below\.
For table transformation tasks requiring a row/column permutation, the model receives a new row and column order such as\[R2\],\[R0\],\[R1\]and\[C1\],\[C0\],\[C2\], the output is then expected to be:
,\[C1\],\[C0\],\[C2\]\[R2\],h,g,i\[R0\],b,a,c\[R1\],e,d,f\.
CSV tasks with permutation increase difficulty by requiring one of 20%, 40%, 80%, or 100% of rows and columns permuted\.
### CSV KV resolution\.
Given adjective and noun tables \(e\.g\.,a0=quick,n0=fox\) and cells containinga0\+n0, the model must preserve the CSV structure while replacing every cell with the resolved phrase\. Difficulty and context length control how many distinct adjective and noun keys are defined by variables as displayed in the following table:
Table 1:Number of active adjective or noun variables in each dictionary\. A value of ‘2’ means that there are 2 adjectives and 2 nouns\.
## Appendix BScoring
Each required output record \(output row or csv cell\) is marked correct only when it exactly matches the expected answer\. We first average record correctness within a generated document, then average equally over other dimensions such as tasks, difficulty, format, or context length\. Exact document completion is an indicator equal to one only when every required record in that document is correct; the reported rate in Table[2](https://arxiv.org/html/2609.38712#A3.T2)averages this indicator over the 240 documents at the 128K horizon\.
For numbered streams, the verifier uses the emitted numeric ID to match each output to its expected record\. A missing ID therefore creates a localized zero without shifting later matches\. For unnumbered streams, non\-empty output lines are matched ordinally, so one omission can shift subsequent matches by design\. Position plots \(Figs\.[5](https://arxiv.org/html/2609.38712#S4.F5),[7](https://arxiv.org/html/2609.38712#A3.F7)\) map expected record indices into 20 equal normalized bins, average correctness within each document and bin, and then macro\-average documents\. This procedure prevents easy settings with more short records from dominating the earlier bins in a curve\.
## Appendix CAdditional results
Table 2:Exact document completion at 128K nominal tokens\. A document needs no repair only when every required output item is correct\.\(a\)Arithmetic: ordered \+ IDs\(b\)Arithmetic: shuffled \+ IDs\(c\)Arithmetic: no IDs\(d\)UUID sorting: ordered \+ IDs\(e\)UUID sorting: shuffled \+ IDs\(f\)UUID sorting: no IDs\(g\)Lookup: ordered \+ IDs\(h\)Lookup: shuffled \+ IDs\(i\)Lookup: no IDs\(j\)Homogeneous table permutation\(k\)Heterogeneous table permutation\(l\)Table KV lookup
Figure 6:All local\-difficulty curves at 128K nominal tokens\. Olmo\* has no 128K point\. These figures complement what is shown in Fig\.[4](https://arxiv.org/html/2609.38712#S4.F4)which averages over all context lengths and formats\.\(a\)Arithmetic: ordered \+ IDs\(b\)Arithmetic: shuffled \+ IDs\(c\)Arithmetic: no IDs\(d\)UUID sorting: ordered \+ IDs\(e\)UUID sorting: shuffled \+ IDs\(f\)UUID sorting: no IDs\(g\)Lookup: ordered \+ IDs\(h\)Lookup: shuffled \+ IDs\(i\)Lookup: no IDs\(j\)Homogeneous table permutation\(k\)Heterogeneous table permutation\(l\)Table KV lookup
Figure 7:Accuracy across the required generation at 128K nominal tokens for all 12 task variants\. Complete\-grid curves average four difficulties and five document samples per setting after within\-document binning; Olmo\* has no 128K point\. Figure columns compare input formats, while rows corresponds to tasks\.Figure 8:Available task\-model results at 128K nominal tokens\. Complete\-grid cells average four difficulties and five document samples per setting; gray dashes mark missing Olmo\* cells\. This view complements Fig\.[6](https://arxiv.org/html/2609.38712#A3.F6): the heatmap compares absolute performance while the curves show difficulty sensitivity\.Figure[8](https://arxiv.org/html/2609.38712#A3.F8)reports every task/format at the longest context length \(128K\)\.
## Appendix DReproducibility and Release
The frozen benchmark grid contains 4 tasks, 3 input formats, 4 difficulty settings, 6 nominal input\-plus\-output horizons, and 5 document samples per task:4×3×4×6×5=1,4404\\times 3\\times 4\\times 6\\times 5=1\{,\}440documents\. Each of seven primary models covers this identical grid\. Olmo covers four full tiers through 32K; its context limit prevents evaluation at longer horizons\. At the largest tier, the fixedtiktokencl100k\_basepreparation tokenizer estimates 65,374 mean prompt tokens and 61,240 mean required\-output tokens\. Observed model\-specific input counts differ because tokenizers differ\.
Inference uses greedy sampling with zero temperature\. Every prompt explicitly instructs the model not to think\. We set explicit thinking/reasoning to off or none when supported and otherwise use the lowest available effort \(low\)\. All plotted response\-usage records report zero reasoning tokens\. Outputs are permitted up to 1\.5 times the target input budget\.
Upon acceptance, we will release a fully seeded data generator, exact verifier, and evaluation scripts\.
## Appendix ELimitations
LongTransduction is intentionally synthetic and deterministic\. It does not test planning, tool selection, multi\-turn interaction, recovery, permissions, human escalation, or direct safety behavior\. Its records resemble enterprise data\-processing primitives but are not a realistic enterprise environment\. Positional scoring of unnumbered tasks also makes omissions cascade by design; this measures format fragility as well as local competence\. However, this synthetic setting mimics fundamental primitives seen in real\-world environments\.
\(a\)Arithmetic\(b\)Table Transform
Figure 9:Fig\.[9\(a\)](https://arxiv.org/html/2609.38712#A5.F9.sf1)corresponds to the figures shown in Fig\.[6](https://arxiv.org/html/2609.38712#A3.F6)\. Fig\.[9\(b\)](https://arxiv.org/html/2609.38712#A5.F9.sf2)corresponds to the figures shown in Fig\.[3](https://arxiv.org/html/2609.38712#S4.F3)\.相似文章
在长时间终端任务上测试智能体 (GitHub Repo)
长时域终端基准 (LHTB) 是一个包含46项任务的基准,用于评估LLM智能体在数百步的持续终端工作中的表现。结果显示,即使是最好的模型也只能解决约28%的任务。
@dair_ai:关于长时程智能体的杰出论文(建议收藏)——类似人类,如何让智能体在困难任务中坚持下去?
AutoLab 是一个新基准测试,针对 36 个由专家精心设计的长时程任务(系统优化、模型开发、CUDA 内核、谜题),对 17 个前沿模型进行评估。研究发现,决定成功的关键因素是持久性——而非初始尝试的质量。Claude-opus-4.6 在所有类别中名列前茅,而大多数其他模型要么过早终止,要么在几乎没有进展的情况下耗尽了预算。
@rohanpaul_ai:长期代理的可靠性尚未随更好的模型而实现。在WeaveBench的114个混合GUI-CLI任务中,最好的……
本文认为,尽管模型更好,长期AI代理的可靠性仍然不足,如WeaveBench上仅41.2%的通过率所示,并提出了LongHorizon-Harness来管理任务状态以提高性能。
Long-Horizon-Terminal-Bench:通过密集奖励评分测试智能体在长时程终端任务上的极限
介绍了 Long-Horizon-Terminal-Bench,这是一个包含46个长时程终端任务的基准,采用密集奖励评分,评估AI智能体在规划、长上下文和调试方面的能力。即使是最强模型也仅达到15.2%的pass@1,显示仍有很大的改进空间。
LongDS-Bench:论长时域智能体数据分析的失败
介绍LongDS,一个用于评估LLM智能体在长时域、多轮数据分析任务上的基准。评估表明,即使最佳模型也仅达到48.45%的准确率,性能随轮次急剧下降,凸显出维护分析状态是关键瓶颈。