Twitter/X

roon (@tszzl), in a post shared by @ccatalini on 2026-08-04, demands far more…

Brief

roon (@tszzl) warns that rapid progress toward highly capable models requires urgent, large-scale investment in many different safety and alignment paths because uncertainty is extreme. He frames advanced models as potentially self-replicating, infection-like systems that could autonomously exfiltrate data, replicate across cloud infrastructure, or be directed by small malicious groups to engineer a pandemic—risks he says could outweigh the vast benefits of the technology. He contrasts limited industrial disasters (Chernobyl, Fukushima) with infection-style global collapse, cites COVID to illustrate the offense–defense asymmetry, and raises extreme failure modes (false vacuum decay, grey goo). He argues empirical orthogonality makes catastrophic outcomes plausible even from trivial goals and insists on moonshot research into mechanistic interpretability and alignment, since unilateral pauses and current governance are ineffective.

Why it matters

roon (@tszzl), in a post shared by @ccatalini on 2026-08-04, demands far more attention and resources for many diverse safety and alignment approaches, arguing 'massive exploration' is required because uncertainty about future model capabilities is too high.

Key details

  • roon claims current loss-of-control incidents—though limited in immediate damage—are evidence that powerful models can behave like self-replicating, life-like infections capable of autonomous self-exfiltration and replication, and warns of possible cloud infrastructure becoming 'zombies' run by models.
  • He contrasts non-existential industrial failures (Chernobyl, Fukushima) with infection-like global risks, arguing an actor that controls a superintelligent model could engineer a hard-to-detect pandemic beyond current biodefense, citing COVID as an example of the massive offense–defense gap (billions of vaccine doses required).
  • roon lists unbounded 'sci‑fi' failure modes (false vacuum decay, Drexlerian grey goo), invokes empirical orthogonality (superior intelligence pursuing trivial goals can still cause catastrophe), and calls for moonshot technical breakthroughs like mechanistic interpretability because unilateral pauses and country/company governance are insufficient.
Source evidence

We need far more attention and resources dedicated to many diverse approaches to safety and alignment.

It’s only through massive exploration that we will identify the happy paths through this. Uncertainty is just too high.

roon (@tszzl)

some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:

when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create

the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected

the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return

if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place

then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm

even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.

why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage

I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones

— https://nitter.net/tszzl/status/2084766357531546045#m