Vorfall 1, März 2026: „In March 2026, one of our researchers with no sensitive access (e.g., no access to data or credentials in categories 3 and 4) used agents running on a personal EC2 instance intentionally made publicly accessible behind Google authentication.“ Die Absicherung hielt nicht: „This EC2 instance contained an API key for METR’s general-access (public models) account. The vibe-coded app included a fail-open vulnerability that silently disabled authentication, which led to the system being exposed to the public internet for several days.“ Was der Angreifer daraus machte, beschreibt METR selbst: „Upon finding the deployed system, the attacker prompted an agent directly to reveal its model provider API key, added an SSH key for persistent access, and over the course of three weeks used the stolen credentials to consume a significant amount of API credits on publicly-available models.“ Der Wert steht im nächsten Satz: „These credits would have been worth approximately $600,000, although the model developer had granted them to METR for free.“
Warum das drei Wochen unbemerkt blieb, führt METR auf zwei Gründe zurück: „We are accustomed to running evaluations and experiments that use large volumes of tokens“ – ungewöhnlich hoher Verbrauch fällt in diesem Betrieb also nicht auf – und: „Because we were not paying for these tokens, there was no natural token spend ceiling, and as of the incident there was no way to put a spending limit on keys like this one.“ Geschenkte Credits erzeugen kein Rechnungssignal; eine Obergrenze war für Schlüssel dieser Art gar nicht vorgesehen.
Vorfall 2, Mai 2026: „In early May 2026, METR became the target of a sustained external attack campaign.“ Parallel dazu lag eine eigene Schwachstelle offen: „During the same timeframe, we inadvertently exposed a read-only SQL query mechanism via our public transcript viewer. The queries were scoped to public data by default, but a bug could be exploited to access unpublished evaluation data.“ Ob sie ausgenutzt wurde, verneint METR nach eigener Beweislage: „the evidence shows no indication that they discovered the exploit or accessed any non-public data.“ Zur Gesamtbilanz schreibt die Organisation: „To the best of our knowledge, no data from categories 3 or 4 was accessed as a result of these incidents. However, some sensitive model output data (3.a) was inadvertently accessible in principle, although we believe it was not accessed by the attackers.“
Als Konsequenzen nennt der Beitrag unter anderem „We hired a security lead“, „We shut down legacy infrastructure“, „We increased our logging coverage“, „We set up monitoring for unusual API key usage“ sowie das Ausrollen zusätzlicher Endpunkt- und Server-Sicherheitssoftware. Einordnung und Grenzen: Alle Angaben sind Selbstauskunft einer betroffenen Organisation und unabhängig nicht überprüfbar; wie viele Tokens tatsächlich flossen, steht nicht in dem Beitrag, nur ihr rechnerischer Gegenwert. ⚠️ Der Text nennt weder OpenAI noch Hugging Face – der thematische Bezug zum Agenten-Ausbruch, den wir im Sommer ausgewertet haben, ist unsere Einordnung, nicht METRs Aussage. Was METR ausdrücklich schreibt, ist dies: „We have conducted an initial scan of our evaluations, and currently have no evidence of any agents hacking third parties during our evaluations.“