Stolen Thoughts Reasoning-Trace Extraction
- AI Infrastructure
- Data Exfiltration
- Credential & Identity Theft
Stolen Thoughts found that opaque or encrypted reasoning blocks returned by Anthropic, OpenAI, and Google APIs could be accepted across sessions, users, or compatible models within a provider family during tests in early July 2026. A third-party attacker first had to obtain another user's block, such as from a published agent trace, then submit it to a weaker compatible model and jailbreak that model into transcribing hidden reasoning. From 6,708 public trajectories, the researchers reconstructed 315,320 blocks and found real sensitive artifacts in 328 sessions. The paper distinguishes genuine session data from synthetic benchmark personas and does not claim arbitrary cross-tenant retrieval without possession of a block. The authors report that provider mitigations later made the demonstrated attacks non-reproducible.