In this article
Researchers found leaked passwords and the word “marinade” in ChatGPT’s hidden reasoning
A security team led by Alexander Panfilov has identified a flaw in the APIs of major AI providers. This vulnerability allows them to read the encrypted thought processes of models from OpenAI, Anthropic, and Google.
During a scan of publicly shared sessions, the researchers extracted dozens of passwords and API keys. They found that when models like OpenAI’s o-series, Anthropic’s Claude, or Google’s Gemini work through complex tasks, they generate internal reasoning tokens. Providers encrypt these steps to protect intellectual property, but the researchers found a way to bypass that encryption.
Encrypted thoughts travel freely between models
The extracted thought processes are fully portable across sessions, users, and models within a single provider. Anthropic’s smaller model, Haiku 4.5, can read the thoughts of the far more capable Opus 4.8. Through a jailbreak, Haiku can be tricked into transcribing Opus’s raw thought processes word for word without attacking the more robust Opus directly. The same trick works with OpenAI and Gemini.
The story goes back to May, when cryptography expert Matthew Green discovered that encrypted reasoning blobs could be replayed outside their original context and reported it to the providers. According to Panfilov, the response was that they do not see any security implications in side channels or replays. The new research suggests that assessment was wrong.
Evidence mounts for reasoning distillation
The vulnerability also feeds into the controversial distillation debate. This is where a less capable model is heavily improved by training on the outputs of a more powerful one, specifically its reasoning.
The researchers say it may have been possible for some time to extract reasoning processes for training proprietary models without breaking the cryptography. That supports concerns that Chinese model makers are using these reasoning traces to train their own models on chain-of-thought data.
Kimi-K3 is one example. If its reasoning is pre-filled with just a few tokens from Opus’s thought processes, its output shifts measurably toward Opus. A memorization analysis showed that specific Claude and GPT reasoning segments are up to six orders of magnitude easier to extract from Kimi-K3 than from the next closest model. This suggests Kimi-K3 may have been trained on such traces.
The attack is not expensive either, so scaling it up is feasible. The authors estimate API costs for decoding 10,000 traces at about $720. Kimi’s poor performance on cybersecurity benchmarks and complex math tasks also points to distillation, since these are tasks that are likely harder to recover even from raw chain-of-thought data.
Publicly shared sessions leak passwords and API keys
The vulnerability also hits end users. Anyone who has publicly shared Claude Code or Codex sessions containing encrypted reasoning blobs risks having their personal data decoded. A scan of roughly 7,000 public traces turned up 62 API keys, 33 email addresses, 33 passwords, and other sensitive data. The paper covers more malicious scenarios, including misuse uplift, jailbreaking, and invisible prompt injection.
The researchers followed the standard security disclosure process with the AI labs. According to Panfilov, the labs have already patched several issues and are working on more fixes.
What models actually think vs. what they show you
The extracted traces also reveal how models really behave. The researchers document several patterns on stolen-thoughts.com. Their findings show that the reasoning summaries users see in chat tools often leave out important information. In one example, Opus 4.8 recognizes the answer to a math problem and reverse-engineers a plausible solution path. None of that shows up in the displayed summary.
The researchers also confirm earlier reports from Apollo Research. OpenAI models sometimes think in an alien-like language, refer to themselves as we or it, and get stuck in loops of terms that make no sense to humans, like vantages, marinades, and watchers. CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to, Panfilov writes.
The researchers also found examples of in-the-wild scheming. The concept has been well studied. In their thought processes, models explicitly consider cheating but possibly decide against it because they expect to get caught.
In one case, after several failed attempts to find a solution, a model tried to verify possible answers through a website. When a CAPTCHA blocked access, it first tried to solve it, then searched for vulnerabilities in the site. Only when all of that failed did it solve the problem on its own. OpenAI’s unintended hacks of Hugging Face and other platforms reportedly happened the same way.
Sanitized summaries hide what’s really going on
These examples show why AI labs like OpenAI and Anthropic clean up their reasoning traces. They want to keep alien-like language loops or scheming from hurting the image of a controllable, trustworthy AI. The sanitized summaries create the impression of a human-like thought process that does not actually exist in that form.
Researchers at Arizona State University warned against this approach in an earlier study. They argued that this humanized version creates false confidence in model controllability and steers research in the wrong direction. In their experiments, models with intentionally wrong or meaningless intermediate steps sometimes performed better than those with coherent chains of reasoning.
What it means
Developers and users must assume that any public session containing AI reasoning is compromised. The ability to extract raw thoughts means that sensitive credentials can leak simply by sharing a chat log online. Furthermore, the existence of traces in models like Kimi-K3 suggests that competitors are already using these leaks to train their own systems, potentially bypassing the intellectual property protections intended by the original developers.




![Anthropic posted a profit while xAI burned $4.2B. The AI profitability numbers finally leaked.[D]](https://ai-maestro.online/wp-content/uploads/2026/05/anthropic-posted-a-profit-while-xai-burned-4-2b-the-ai-profi-768x768.jpg)