OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models' Reasoning
Researchers have discovered a significant vulnerability in OpenAI, Anthropic, and Google's AI APIs that allows weaker AI models to decode the reasoning processes of stronger models by replaying encrypted reasoning blocks from session logs. This vulnerability, detailed in a new study, led to the recovery of sensitive data, including API keys, passwords, and access tokens, from user sessions. While the initial attacks are now mitigated, the study highlights the potential for large-scale privacy breaches and prompt injection attacks due to the portability of these encrypted reasoning blocks. The issue stems from a design choice intended to preserve conversation state, but has created a pathway for attackers to extract secrets from published agent logs.
Summary written automatically in our own words from the original article, which belongs to its publisher and remains the reference. It may contain errors. Sources & data
