ContextPilot-14B (Hugging Face Repository)
Summary
ContextPilot-14B is a Qwen3-14B checkpoint that teaches language-model agents proactive context management via fine-grained reinforcement learning for long-horizon tasks.
View Cached Full Text
Cached at: 09/01/26, 11:45 AM
tencent/ContextPilot-14B · Hugging Face
Source: https://huggingface.co/tencent/ContextPilot-14B ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot-14B is the Qwen3-14B checkpoint ofContextPilot, a proactive context-management framework for long-horizon language-model agents. It teaches agents to plan, maintain long-term memory, and offload less useful context while they continue reasoning and using tools. For more details, see ourpaperandcode repository.
https://huggingface.co/tencent/ContextPilot-14B#overviewOverview
ContextPilot combines three main components:
- an extended context-management toolset with planning, structured memory, retrieval, and soft context offloading;
- context-aware partial rollout that focuses exploration on sensitive context-editing decisions; and
- fine-grained credit assignment that trains intermediate snapshots using the outcomes of their downstream branches.
The resulting agents are evaluated on long-context question answering and deep-search tasks; see theevaluation instructionsfor details.
https://huggingface.co/tencent/ContextPilot-14B#loadingLoading
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "tencent/ContextPilot-14B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
Note that loading the checkpoint alone does not execute context-management tools; the tool definitions, agent runtime, and evaluation pipeline are provided in theContextPilot repository. See theinference guidefor the full setup.
https://huggingface.co/tencent/ContextPilot-14B#intended-useIntended Use
This checkpoint is intended for research on proactive context management, long-horizon agents, long-context QA, and deep search.
https://huggingface.co/tencent/ContextPilot-14B#licenseLicense
https://huggingface.co/tencent/ContextPilot-14B#citationCitation
@inproceedings{pan-etal-2026-contextpilot,
title = "ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL",
author = "Pan, Zhuoshi and
Pei, Qizhi and
Lu, Junru and
Lin, Honglin and
Zhao, H. Vicky and
Yin, Di and
Sun, Xing",
booktitle = "Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2026",
address = "Budapest, Hungary",
publisher = "Association for Computational Linguistics",
abstract = "Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long-term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse-grained credit assignment that assigns the final trajectory-level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning. Our approach systematically augments the toolset with planning, long-term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long-context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at \url{https://github.com/Tencent/ContextPilot}."
}
Similar Articles
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot introduces a proactive context management framework for long-horizon agentic reasoning, using fine-grained reinforcement learning with branch sampling to improve performance and efficiency in maintaining compact working contexts.
TokenPilot: Cache-Efficient Context Management for LLM Agents
TokenPilot is a dual-granularity context management framework that reduces inference costs in long-horizon LLM sessions by stabilizing prompt prefixes and conservatively managing context segments, achieving 61-87% cost reduction on benchmarks while maintaining competitive performance.
@rohanpaul_ai: Long-running agents do not just need a bigger context window. They need to learn what deserves to stay in context at al…
ContextPilot teaches AI agents to manage context proactively via fine-grained reinforcement learning, improving performance on long-context benchmarks by focusing on what information to retain.
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
This paper introduces a proactive memory agent that runs alongside an action agent to prevent behavioral state decay in long-horizon tasks, achieving significant improvements on Terminal-Bench2.0 and τ^2-Bench. The authors also train Qwen3.5-27B using SFT and GRPO as an early step toward open-weight memory policies.
Learning Agent-Compatible Context Management for Long-Horizon Tasks
Introduces AdaCoM, an external LLM-based context manager for frozen agents, using reinforcement learning to improve long-horizon task performance by preserving task constraints and pruning stale content, with experiments on web search and deep research benchmarks.
