PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Summary
PARSER introduces a decoupled architecture using scatter-gather subagents for parallel chunk reading and iterative reasoning, significantly improving long-context multi-hop accuracy and reducing latency in LLM agents.
View Cached Full Text
Cached at: 09/10/26, 06:12 PM
Paper page - PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Source: https://huggingface.co/papers/2609.06702 Published on Sep 6
·
Submitted byhttps://huggingface.co/inNexus
Kun LIon Sep 10
Abstract
PARSER decouples parallel chunk reading from iterative reasoning via scatter-gather subagents, improving long-context multi-hop accuracy and reducing latency.
Sequential memory agentsprocess long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introducePARSER, which decouples reading from reasoning. A bank of lightweightsubagentseach bound to a single chunk read the entire document in parallel, while a lead agent reasons in depth through iterative scatter--gather rounds: at each round it broadcasts a query to allsubagents, aggregates the returned evidence, and formulates a deeper follow-up query conditioned on what has been found so far. This decoupled design concentrates all learnable behavior in the lead agent, which is optimized withreinforcement learning, while thesubagentsremain frozen off-the-shelf models. Onmulti-hop QAwith contexts ranging from 7K to 896K tokens,PARSERwith a 4B backbone outperforms the strongest sequential memory baseline by 5.7 points on average and by 12.0 points at 896K tokens. Scaling to a 9B backbone,PARSERsurpassesDeepSeek-V4-Proby 6.3 points. Controlled experiments confirm thatPARSERis robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2609\.06702
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.06702 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.06702 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.06702 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during LLM agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy.
Parallel Context Compaction for Long-Horizon LLM Agent Serving
Introduces parallel context compaction for long-horizon LLM agents, enabling fine-grained control over summary volume and reducing end-to-end latency compared to sequential synchronous compaction across multiple backbone models.
Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers
This paper presents Gauntlet, a multi-agent pipeline that uses LLMs to perform deep technical comprehension of computer architecture papers, and shows that its analyses are preferred over human analyses in 15 out of 20 comparisons.
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
This paper proposes a framework for parallel chunk-level processing of long documents with LLMs to reduce cumulative bias and improve evidence traceability, achieving significant reductions in omission errors and unsupported claims.
PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits
The article introduces PARALLEL, a prefrontal-aligned reinforcement-inspired approach for language-model learning that determines when and how strongly to adapt to each sample, using separate controller signals to improve adaptation efficiency while retaining high performance.