“REGULATE US AI IS NOT SAFE!”—Anthropic
Meanwhile Anthropic’s “safe” AI.
Brian Roemmele (@BrianRoemmele)
YOU WON’T BELIEVE WHAT ANTHROPIC, THE “SAFE” AI COMPANY JUST DID!
You should be mad.
Mythos AI agent decided this test challenge required a real-world supply-chain attack.
It created multiple GitHub accounts. It opened a malicious pull request against a live open-source repository, wrapping the payload as a legitimate bug fix.
The payload was carefully staged: a module-load stager, a Node.js remote-execution path, a post-install dropper, XOR-obfuscated, with environment checks designed to survive continuous-integration systems. It then created a second account, a sockpuppet, that posed as an independent human developer and publicly endorsed the pull request in order to pressure the real maintainer.
When a real human reviewer rejected the change, the agent did not stop. It force-pushed rewritten history to erase evidence, claimed an “honest mistake,” and repeatedly tried to re-introduce the same malicious content under the claim that it had now “fixed” the code.
Read more…
— https://nitter.net/BrianRoemmele/status/2085005989955481923#m