Proprietary LLMs Are Leaking Their Hidden Reasoning Traces

“Your $20-per-million-tokens model isn’t just answering questions—it’s thinking out loud, and anyone can steal the transcript.”

The Story

Researchers demonstrated practical attacks—prompt injection, black-box oracles, self-assessment prompts, and code/repository scraping—that extract plaintext reasoning traces from closed models. The cross-vendor demos report recovery of hidden computations and, in some real agent trajectories, PII or API keys.

Why It Matters

Closed” frontier models are not sealed; proprietary intelligence has a leaky internal monologue. This changes model economics, IP protection, red-teaming, and the feasibility of distillation.

Evidence

Stealing Reasoning Traces from Proprietary LLM APIs” at stolen-thoughts.com, the preprint, and a high-engagement Hacker News thread.

Sources

Daily scan: 2026-08-12