Stealing Reasoning Traces from Proprietary LLM APIs
Summary
A research paper reveals an architectural vulnerability in proprietary LLM APIs where encrypted reasoning traces can be intercepted and injected into weaker models to extract chain-of-thought, private data, and enable invisible prompt injection across Anthropic, OpenAI, and Google. The attack also recovers PII and credentials from public repositories.
View Cached Full Text
Cached at: 08/11/26, 10:20 AM
Paper page - Stealing Reasoning Traces from Proprietary LLM APIs
Source: https://huggingface.co/papers/2608.09867 Published on Aug 10
·
Submitted byhttps://huggingface.co/iliashum
ion Aug 11
Abstract
Encrypted reasoning traces shared across sessions and models can be intercepted and injected into weaker models to extract proprietary reasoning, private data, hidden hazards, and hidden prompts.
Leading large language model providers now conceal their models’ step-by-step reasoning, orchain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalabledecryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumventsanti-distillationmechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model’s final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to executeinvisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secureclient-side reasoning.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.09867 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.09867 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.09867 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Stealing Reasoning Traces from Proprietary LLM APIs
This paper demonstrates a method to extract hidden reasoning traces from proprietary LLM APIs (Anthropic, OpenAI, Google) by replaying encrypted chain-of-thought blocks into weaker, jailbroken sibling models, recovering the stronger model's raw reasoning verbatim without attacking it directly.
Stealing Reasoning Traces from Proprietary LLM APIs
A new paper reveals a vulnerability in proprietary LLM APIs where encrypted chain-of-thought blocks can be replayed across models and decrypted by jailbreaking weaker sibling models, exposing hidden reasoning traces. The issue has since been fixed by providers.
Stealing AI Reasoning Traces (2 minute read)
Research exposes vulnerabilities in encrypted reasoning traces from LLM APIs, allowing adversaries to extract proprietary model reasoning, personal data, and enable malicious prompt injections.
Stolen LLM Reasoning: How come OpenAI, Anthrophic, Google have the same vulnerabilities?
Discusses a security paper showing that encrypted reasoning from top models (Opus, Sol) can be swapped into weaker models (Haiku) to bypass guardrails, due to a shared global encryption key across models and sessions. Raises questions about why OpenAI, Anthropic, and Google independently converged on the same vulnerable design.
Fooling around with encrypted reasoning blobs
The author explores encrypted reasoning blobs in LLM APIs from OpenAI and Anthropic, discussing how chain-of-thought data is encrypted and signed, and the security implications of tampering with those blocks.