title: Announcing ControlConf 2026
author: Buck
contenttype: lesswrongpost
publication: LessWrong
published: 2026-02-26T02:23:32.963000+00:00
source_url: https://www.lesswrong.com/posts/tN5mgBAinwFe5c67c/announcing-controlconf-2026
word_count: 1798
We’re running ControlConf in Berkeley on April 18-19. It's a two-day conference on AI control: the study of reducing risks from misalignment through safeguards that work even when AI models are trying to undermine them.
Since the last ControlConf (Feb 2025), AI agents have gotten way better. We’re approaching the point where control techniques are load-bearing for the safety of real agent deployments. There’s been a lot of progress too. Researchers have run control evaluations in more realistic settings, and AI companies have started building initial control measures—with much more to come.
At ControlConf, the people at the frontier of this work will present on current research problems, promising interventions, and the most important research directions going forward.
Apply here.
I expect talks and discussion on topics like these:
As AIs get more capable, they’re getting better at both monitoring and at generating subtle attacks. On current trajectories, which will they be better at when they’re automating AI R&D?
To what extent will external auditors be able to evaluate whether companies are robust to threats from internally deployed agents?
CoT monitoring: how much should we rely on this? How is it working so far?
Can we use permissions management to mitigate our threats (as discussed here)? How should permissions management for agents in high-security contexts work?
To what extent should we apply control-like risk analysis to misaligned motivations that aren’t scheming?
How far are we away from AIs that we need to treat as misaligned from a control perspective?
What needs to happen on deals with early schemers (or non-scheming misaligned models)?
How hard is it to subtly sabotage ML experiments?
How likely is it that we'll be deploying a known schemer? How much should we focus on such worlds?
How much should we capitalize on derpy schemers’ suboptimalities?
Claude monitors GPT? How tractable is this and what are the blockers?
What domains of research are most important to prevent the sabotage of?
What if future AIs look pretty damn different from modern architectures/scaffolds? Eg if they're insane bureaucracies of agents?
What kind of incrimination would be best for warning-shot-raising/political-will-building?
Maybe the field of AI-empowered security is where all the high-stakes control juice is. Is that the case, and if so, how do we make it go better?
The great thing about AI control as an approach to mitigating misalignment risk is that it makes risk analysis (comparatively) so concrete: for our threat modeling, we can just argue about what AIs will be capable of, rather than needing to argue about inductive biases or whatever. So I’m hoping that this conference is a great place for discussion of misalignment risk that engages seriously with evidence and arguments about the future.
We’re also taking the opportunity to run a one-day workshop on April 17 on AI futurism and threat modeling, aimed at people who want clearer models of catastrophic AI risks and the best strategies for mitigating them. Apply to that here.
Tags: AI
Karma: 33 | Comments: 0 | Author: Buck
Top Comments
Matt Vincent (6 karma):
In this post, Scott Alexander makes a good case for correcting people, but he also provides a few guidelines for determining what counts as nitpicking and what counts as rigor. Here are his caveats:
Let someone be a little wrong if their impact is small and they are not in a mood to debate.
Allow oversimplifications, figures of speech, and misnomers for pedagogical and artistic purposes, unless you can explain why an argument hinges on their choice of terms.
If someone dismissively accuses you of nitpicking, instead of explaining why your distinction isn't relevant, then they're bullying.
Despite my attempt to summarize his post, it's worth reading in its entirety.
Parv Mahajan (39 karma):
RSP takes from a bunch of Astra fellows:
Seems like Anthropic should've known RSPv2 would fail when the RAND report came out, and in retrospect it's kind of embarrassing we (the community) didn't realize this earlier
We're very divided on whether the phrasing/stance on "Anthropic has to win" is good/correct, especially given the talk about "marginal risk" considerations. We're somewhat concerned that Anthropic simply won't pause when it's clear (to concerned parties internally) they probably should.
Why don't they just say racing is bad and that a pause (at some point) would be good? This seems so low-cost to put in the intro/industry reccs., or at least to make an OOM more clear.
Are Anthropic employees not reacting to this? It feels surprisingly low-profile for such a big change in internal governance (although I suppose there are Other Things happening).
Maybe Anthropic should've been more clear about what "behind" and "ahead" mean, and when or when not they're giving themselves the option/soft obligation to pause
In general, we're quite confused about Anthropic's viewpoints on the difficulty of alignment and the likelihood of AI takeover.
Risk reports seem good! We are quite excited for these! But 6 months is way too long of an interval (3 months might be okay?), and we would be less nervous if there were many addendums + edits as models were deployed (and this seems to be the case!). Also, we are unconvinced this doesn't fail during software-only AI R&D takeoff.
On a personal note, many of us are much more nervous about working for Anthropic and are much more nervous about the strategic decision-making of its leadership during the critical period.
Raemon (6 karma):
This'd be easier to engage with with real examples.
I think this is sort-of right, but I think "prior" is not a clear enough mechanistic description of what's going on. Like this is a fine abstraction but it doesn't give me many details to think through towards solutions.
Also, like, pretty much any human knowledge isn't intrinsically "prior", you believe things because of your life experience.
One example that you might classify as a prior but I think we can do better: some people have the intuitive sense that a significantly-more-intelligent person can stomp all over less intelligent people.
I think many of these people get the sense from having, say, played a bunch of competitive strategic games, and have a sense of how high skill-ceilings are.
Some people don't have that experience, instead their reason for the prior is just "it does sure looks like humans stomp all over chimpanzees", and generalize.
On the other side, some people have life experience where messy finnicky details got in the way of supposedly "intelligent" plans, and they have a stronger salient prior about intelligence being bogged now in finnicky details.
In all those cases I think you can do better than stop at "well their prior is different."
Or, rather: being enable to do that is a skill issue.
(See also: Propagating Facts into Aesthetics, which goes into one particular mechanism in more detail)
aysja (12 karma):
It's not just any blog post. It's a blog post outlining a new major strategical shift in the company, specifically in the direction of giving Anthropic far more leeway over how they decide what the risk is and how to deal with it. It seems especially important to state "we can’t leave this up to the companies" loudly and clearly here.
RogerDearnaley (5 karma):
I think Claude is trolling you! :-D Or more accurately, I think it's roleplaying. The dominant answer here is silly/absurd, and paperclips is basically more of the same.
cousin_it (12 karma):
Well, if what Zvi writes is true - that Anthropic was "proactively" helping the military and that their red line was "no mass domestic surveillance" - then I as a non-US person become even more disillusioned about Anthropic's ethics than I already was.
Mitchell_Porter (15 karma):
The Trump 2.0 presidency is authoritarian. They go after their opponents with a hammer. Recall what they did to the federal bureaucracy, to the universities, I won't even bother trying to think of other examples.
Ever since the downfall of FTX, Sacks and the accelerationists have insisted that their number-one concern about AI is that we're all going to end up under the thumb of an AI regime enforcing a common value system (e.g. "woke AI"). Thanks to Gemini's anachronistically diverse images, Alphabet/Google used to be the prime suspect here, but like Meta/Facebook, they have replaced their virtue signaling with conspicuous silence. As for OpenAI, they have long since demonstrated a willingness to shed inconvenient principles.
On the other hand, Anthropic is an enemy because it has Biden's AI-policy people working for it, and ironically, because the company touts the importance of ethics. And contrary to Zvi's hopes that this is a time for the AI industry to stand together, I think the other three frontier-AI leaders would be quite happy to see Anthropic removed from the chessboard entirely.
Noah Birnbaum (7 karma):
This might feel obvious, but I think it's under-appreciated how much disagreement on AI progress just comes down to priors (in a pretty specific way) rather than object-level reasoning.
I was recently arguing the case for shorter timelines to a friend who leans longer. We kept disagreeing on a surprising number of object-level claims, which was weird because we usually agree more on the kinda stuff we were arguing about.
Then I basically realized what I think was going on: she had a pretty strong prior against what I was saying, and that prior is abstract enough that there's no clear mechanism by which I can push against it. So whenever I made a good object-level case, she'd just take the other side — not necessarily because her reasons were better all else equal, but because the prior was doing the work underneath without either of us really knowing it.
There's something clearly rational here that's kinda unintuitive to get a grip on. If you have a strong prior, and someone makes a persuasive argument against it, but you can't identify the specific mechanism by which their argument defeats it, you should probably update that the arguments against their case are better than they appear, even if you can't articulate them yet. From the outside, this totally just looks like motivated reasoning (and often is), but I think it can be pretty importantly different.
The reason this is so hard to disentangle is that (unless your belief web is extremely clear to you, which seems practically impossible) it's just enormously complicated. Your prior on timelines isn't an isolate thing — it's load-bearing for a bunch of downstream beliefs all at once. So the resistance isn't obviously irrational, it's more like... the system protecting its own coherence.
I think this means that people should try their best to disentangle whether some object level argument they’re having comes from real object level beliefs or pretty abstract priors (in which case, it seems less worthwhile to press on them).