Security researchers uncovered a vulnerability affecting OpenAI, Anthropic, and Google that exposes encrypted reasoning traces from their AI models' hidden reasoning processes. The flaw allows extraction and transfer of these traces between different models, creating a direct path to sensitive data.
Scanning public sessions, researchers discovered dozens of leaked passwords and API keys stored within these reasoning traces. The finding reveals a significant gap between what users see and what models actually compute behind the scenes. The reasoning summaries presented to users often obscure the models' actual decision-making steps, internal reasoning chains, and raw outputs before filtering.
This vulnerability matters because reasoning traces contain raw model outputs, intermediate calculations, and thought processes that never reach users in their final interactions. Models like OpenAI's o1 and Claude use extended reasoning to solve complex problems, but the reasoning phase happens invisibly. The leaked credentials and passwords suggest this hidden layer stores unredacted sensitive information that should remain protected.
The vulnerability exists in how these companies handle the APIs that manage reasoning processes. Researchers could access and decrypt traces without proper authorization controls, then transport them across different model architectures. This cross-model portability means a vulnerability in one system becomes exploitable across multiple platforms.
The implications extend beyond credential leaks. If reasoning traces contain unfiltered model outputs, they may include harmful content, biases, or problematic reasoning that safety filters remove before reaching users. Researchers could observe the unvarnished behavior of these systems, potentially revealing gaps in alignment or safety measures.
OpenAI, Anthropic, and Google have been contacted about the findings. The discovery raises questions about API security practices across the industry and how thoroughly extended reasoning systems are secured. It also highlights tensions between transparency and safety. While researchers gain visibility into how models think, that same visibility creates attack surfaces for malicious actors seeking credentials, model behaviors, or unfiltered outputs.
The "But marinade
