AI Model Evaluator METR Hit by Credential Theft, Probing
The AI Model Evaluator (METR), a nonprofit focused on assessing risks in frontier AI models, suffered two significant cybersecurity incidents. In March, a researcher exploited a fail-open vulnerability on a personal AWS instance to steal an API key and consume $600,000 in public AI model credits. In May, the organization was targeted by a sustained attack campaign attempting to gain access to advanced AI models, exposing a vulnerability in a public transcript viewer. Despite its maturity and prominent role in the AI security landscape – including collaborations with major AI vendors – METR highlighted that these incidents stemmed from basic cloud and credential hygiene issues, emphasizing the need for heightened security practices within the AI supply chain.
The AI Model Evaluator (METR), a nonprofit dedicated to evaluating risks associated with emerging AI models, experienced two separate cybersecurity incidents in August. In March, a researcher, without access to sensitive data, leveraged a vulnerability in a personal AWS instance to gain unauthorized access. Specifically, the researcher deployed agents to this instance using a ‘vibe-coded’ orchestration tool, which included a fail-open authentication flaw. This flaw silently disabled authentication, exposing the system to the public internet for several days, allowing the researcher to steal an API key for METR’s public models account. The attacker subsequently used this key to establish persistence, consuming approximately $600,000 in API credits on publicly available AI models over three weeks.
In May, METR detected a sustained external attack campaign targeting the organization, with attackers probing for potential financial gain or access to advanced AI models. During this campaign, METR inadvertently exposed a read-only SQL query mechanism through its public transcript viewer. While the attackers initially probed this endpoint, they did not discover the vulnerability and were unable to access any non-public data. However, the vulnerability did expose a portion of category 2 data – unpublished evaluation results involving public models and API keys granting access to public models – and some sensitive model data (category 3). An independent researcher discovered the vulnerability and responsibly disclosed it, leading to the API being temporarily taken offline and a bounty reward being paid.
METR has become a notable player in the AI security ecosystem, collaborating with major AI model vendors like OpenAI, Anthropic, Google, Meta, and Amazon. The organization utilizes a four-tier data classification system and has SOC 2 Type I certification, and employs a dedicated security consultant. Despite this level of security maturity, the incidents highlight the importance of robust cloud and credential hygiene practices, even for organizations considered security-mature within the AI industry. METR emphasized that it has no evidence of agents hacking third parties during its evaluations and no evidence of agents hacking the company’s evaluations. Jacob Krell, senior director of secure AI solutions and cybersecurity at Suzu Labs, notes that METR’s role as an independent AI evaluator places “trillions” of dollars’ worth of crown-jewel intellectual property within the organization’s security perimeter. He argues that evaluator organizations should be treated as part of the AI supply chain, with their security requirements reflecting that.
