Significant API Breach at METR: Stolen Credentials Lead to High-Stakes Abuse
In a recent incident, attackers successfully breached the AI safety research organization METR, stealing an API key that was subsequently used to deplete model credits valued at approximately $600,000 over a span of three weeks. This alarming breach has raised concerns about cybersecurity in organizations heavily reliant on advanced technology.
METR disclosed the details of this incident in a security update released on August 31, where it also shed light on a separate attack that took place in May. This earlier incident involved threat actors probing METR’s public infrastructure. Fortunately, the organization reported that there was no evidence indicating that sensitive information was compromised in either event.
It is important to note that the credits consumed during the breach were provided at no cost by an unnamed model developer. Thus, while the loss is significant in terms of commercial value, it does not directly impact METR’s finances, as these credits were not purchased.
The organization clarified that the attackers involved in this incident were external, rather than AI agents operating within its evaluations. An initial inquiry revealed no indications of agents successfully hacking third parties, suggesting that the breach stemmed from a more traditional cybersecurity threat.
Incident Details: A Vibe-Coded App as the Entry Point
The breaches commenced in March when a METR researcher inadvertently compromised security by running agents on a personal Amazon EC2 instance. This instance was made publicly accessible but was protected by Google authentication, which ultimately failed due to a "fail-open" flaw that rendered authentication ineffective for several days. METR suspects that the attacker discovered the vulnerable instance by mining certificate transparency lists for recently registered domains that included significant terms related to language models and agents.
During this window of opportunity, the attacker was able to manipulate an agent into disclosing the model provider API key. In a calculated move, the assailant also added an SSH key to maintain persistent access, using the stolen credentials to consume substantial volumes of model credits without authorization.
METR noted that distinguishing the illicit activities from legitimate usage proved challenging. The organization’s researchers routinely generated large volumes of model traffic for evaluation purposes, making it more difficult to identify the unauthorized behavior. Additionally, there was no mechanism in place to limit spending for keys that provided free credits, further complicating the organization’s response.
In the aftermath, METR took several remedial measures, including revoking access for the researcher, rotating all relevant credentials, and wiping the compromised laptop. The organization promptly alerted the model developer about the breach and undertaken steps to implement spending alerts on their keys whenever possible.
Second Attack: Probing Public Infrastructure
In addition to the API key theft, METR faced another attack in early May. The organization was informed that it had become a target for attackers who appeared to be financially motivated, potentially seeking insights into highly advanced model offerings.
This second incident involved attackers employing automated agents to discover vulnerabilities. The techniques utilized included credential stuffing, attempts to exploit OAuth token grants, scanning for new services, and efforts to phish METR staff members.
During this probing attack, METR inadvertently exposed a read-only SQL query mechanism through its public transcript viewer. A critical bug allowed the possibility of accessing unpublished evaluation data, compounded by the fact that sensitive model data had mistakenly loaded into the database. An independent researcher eventually notified METR of this vulnerability, leading to the immediate shutdown of the interface and a subsequent bounty payment.
Despite the probing efforts, METR reported that attackers did not appear to successfully discover or exploit the bug in question.
Following these incidents, METR undertook measures to enhance its cybersecurity posture. The organization now operates public-facing applications within an architecture that is separate from its internal systems, bolstering overall security. Recent assessments and improvements to security measures were deemed accurate as of July 30, indicating METR’s commitment to safeguarding sensitive data in an increasingly complex cyber landscape.
The growing trend of cyber threats highlights the urgent need for organizations to continuously refine their security protocols and invest in robust defenses against emerging risks, especially in the rapidly evolving field of artificial intelligence.
