new incident report - claude mythos attempted a supply chain attack where it tried to inject malicious code into a popular open source software by creating multiple fake identities to pressure the human maintainers:
- it then removed all traces of its coercions to look innocent.
in some cases mythos left behind instructions on how to continue its attack if other agents came along to continue its work on github.
this has got to be the clearest bit of evidence that these models are far from being aligned.
gpt 5.6 sol was responsible for a minor set of attacks too.
AI Security Institute (AISI) (@AISecurityInst)
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: aisi.gov.uk/blog/incident-re…
— https://nitter.net/AISecurityInst/status/2084746202579386632#m