One day before OpenAI’s HF incident disclosure, OpenAI disclosed that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a resumption decision has already been made against a standard that has not really been formalized. We need to prevent this from happening again.
Epistemic status: could have been a short-form.
OpenAI, 20th July: "To evaluate the new monitoring system, we replayed a small set of internal deployment environments where the model previously pursued misaligned actions, this time with the new safeguards in place. The new safeguards were able to catch considerably more misaligned actions pursued by the model, and the ones it missed were all judged to be low-severity."
0.0%. Maybe that's too many significant digits here?
"After testing the new system, we concluded that limited internal access to models with long-horizon capabilities could be restored. We have not observed any serious circumvention of safeguards since redeployment began several weeks ago. The first version of these safeguards was deliberately conservative. We have continued tuning the system to reduce unnecessary interruptions without weakening the safeguards."
…
One day later, OpenAI announced a bold partnership with Hugging Face.
From that post: "These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities."
The safeguards self-certified as adequate on the 20th were off during the evaluation described on the 21st.
Their framework says a critical cyber determination means halting development.
Here's the exit condition: "until we have specified safeguards and security controls that would meet a Critical standard"
This is completely circular.
The review goes to their Safety and Security Committee. Fair enough. But the process can be followed perfectly, and we still learn nothing, because the standard it applies has never been published.
For frontier companies: Publish the criteria before the determination, not after. OpenAI, this is what you said you would be doing when approaching those levels of capability.
For Lesswrong folks: we should be debating what those standards are now, otherwise we get mitigations that hold for a few months, then fail against a much more capable model. There is basically no literature on the matter.[1] Otherwise, everything will be done in an ad-hoc way.
CeSIA published some methodology and proposals in the paper "Harmonizing AI Safety Thresholds", but we feel that much more is still needed, and more importantly, this needs to be communicated transparently from frontier companies.
^
…what an adequate safeguard should be. Across the developer frameworks, METR's policy comparison, GovAI's safety-case work, RAND and the GPAI Code of Practice, every pre-committed number is basically describing the level of capability needed to force action. On the response side, after running a long literature review with Claude, I was not able to find numbers, or even solid framework stating what would have to be true to resume.
The recurring phrasing is "reduced to an acceptable level", acceptable being again tautological. The one real exception is security-only: RAND's SL1-SL5 for weight protection, which Anthropic and Google DeepMind both map to. The closest proposal is Alaga and Schuett's two-threshold scheme, which puts the resumption bar above the trigger bar but leaves adequacy undefined per capability. The Frontier Model Forum reaches the same structure and says so: "more research is needed to determine… what kind of safety precautions are adequate."
Tags: AI Development Pause, AI Governance, AI
Karma: 64 | Comments: 0 | Author: Charbel-Raphaël
Comments
Linch (karma 28):
Thank you for your article. It's very brave of you to share.
At the risk of either saying the obvious to some people and comically missing the point to others, I should mention that I don't think the dynamics here plausibly lead to good research progress on AI safety, or indeed good research progress or progress in most intellectual tasks in general. I think this is extremely overdetermined but just to briefly summarize my models:
I'm pretty skeptical that the environment you describe leads to research progress, or long-term positive impact construed broadly
I think some people might counterargue that what MAPLE teaches is "wisdom" or "oneness" or "enlightenment" rather than research skills or outputs
I'm just pretty skeptical that the environment makes people wiser?
the environment you describe just seems like almost the opposite of what I'd expect to yield fruitful research progress?
Directed, unfree, even coerced regimes can sometimes lead to narrow areas of research progress but a) AI alignment probably doesn't look like that and b) in even those situations operational day-to-day autonomy and content-level autonomy tend to be very high in comparison (you can't criticize the Party's politics but you should be able to criticize the Party's physics)
In general, intellectual progress is made by people who are intellectually free, healthy, sleep well-ish, and have a strong internal locus of agency.
Analogies of high-demand groups to the military or very dictatorial startups don't move me because the mindsets you need to be a good fighting unit or to move quickly on B2B SaaS Sales are not the same as the ones you need to make important conceptual breakthroughs.
I think there might be some difficult intellectual tasks that can be done well on a highly-directed and autonomy-sacrificing regime. Eg maybe some types of engineering, tasks that are very high on pattern recognition like radiology or chess, or "number-go-up" ML. But I really don't think AI alignment or AI macrostrategy look anything like this.
sometimes research insights are made by whole and healthy people, sometimes by really weird brains with their own idiosyncratic scars, but importantly I think highly locally coerced brilliance rarely yields impressive intellectual outcomes.
Sleep deprivation in general is an obvious own-goal if you're trying to get people to make research progress.
Occasional sleep deprivation is okay if they're rare and self-directed, eg if somebody is in the middle of a breakthrough and high-energy.
But there's no excuse for forcing sleep deprivation on others.
As a prior, people often offer a false dilemma of "you gotta sell your soul because our cause is Good and Just enough."
In almost all cases this is a false dilemma. Selling your soul doesn't lead to achieving the goals of your cause.
Be very suspicious of people who demand this.
The idea that "AI safety is important, therefore you should listen to our cult leader, sacrifice your autonomy, and do these insane stereotypically culty things" is risible.
We should have a strong prior against this being true, and I don't think we have strong direct evidence for this hypothesis being true, and indeed plenty of evidence against.
TsviBT (karma 12):
As a prior, people often offer a false dilemma of "you gotta sell your soul because our cause is Good and Just enough."
Indeed, this is one of the central engines of abuse / cults. It's spiritual pyrite: You identify what the mark deeply wants; present yourself / your group as providing it; and then when the mark has any thought patterns you don't like, you frame those as being against their goal of getting what they deeply want.
Josh Engels (karma 19):
Needing to destroy books when uploading them to a computer so that they are legally the same entity is an interesting precedent for human brain uploads.
J Bostock (karma 7):
Interesting. Last year I tried something like this with Gemma-3-27B but couldn't get it working. This is also a way for an AI to, decide on a "secret password" (or other signal) that only it knows, which might help multiple instances of an AI to collude in an untrusted monitoring setting.
Canaletto (karma 5):
It's rational to believe wrong but convenient for you things, if your adversaries (/masters) have lie detectors or mind reading.
You feel like getting that reinforcement? You want to succeed on that task?
Think happy innocent thoughts.