hraness
Theme
Appearance

saved

Stop AI: How and Why

by David KruegerPost-AGI Workshoppublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

David Krueger argues advanced AI should be stopped, not merely regulated, because the most competitive systems cannot yet be made acceptably safe. He treats the international race as mostly domestic politics rather than an unsolvable coordination failure and proposes dismantling the AI compute supply chain. He closes with Evitable's movement-building path toward an indefinite global pause.

ideas

  • Acceptable risk requires a pause. Marginal safety work is not enough if competitive AI cannot be made safe enough.
  • The race is more politics than destiny. Domestic politics, not pure game theory, shapes whether states can stop.
  • Target the compute supply chain. Systematically dismantling chips and related bottlenecks is framed as more tractable than fine-grained model rules.
  • Movement strategy matters. Evitable focuses on building the political coalition for an indefinite global pause.

quotes

“We haven't solved alignment, interpretability, or safety testing, which are the main technical problems that people are working on.”

David Krueger

“I would say we've made some progress, but they're certainly not solved.”

David Krueger

“People do disagree a lot about timelines, about basically everything.”

David Krueger

transcript

There's so much to say on this topic and I'm sure I'm going to miss important things, but let's start by just looking at the situation we're in. We haven't solved alignment, interpretability, or safety testing, which are the main technical problems that people are working on. I would say we've made some progress, but they're certainly not solved. The AI 2027 forecast seemed more right than wrong. I think the timelines are still longer, but companies are saying, yeah, we're starting to do that recursive self-improvement thing — no, we don't know how to make it go well.

People do disagree a lot about timelines, about basically everything. But that disagreement means that there's a bunch of uncertainty. And so the worst-case things have a pretty decent chance of being true if you just take some sort of outside view on the community of experts. Anti-AI sentiment is high and it's increasing, at least in the US, but we're still not really regulating AI, which is kind of interesting. X-risk hasn't been a big part of the conversation, but mass unemployment is, which is roughly equivalent from a policy point of view. It's an extremely bad outcome that we don't have a plan for that should be averted at almost all costs. And I would say there haven't been any real serious international coordination efforts, especially to stop or to say, let's actually control this in a meaningful way. And nobody really has a plan for the post-AGI world — which is what this whole workshop is about.

I keep having this experience where I meet somebody who's new to this topic and they get it. They've caught up to speed a bit and they're like, wow, this is insane — obviously, if we could, we should stop building AI. And then another experience I keep having is I meet somebody who's not new to the topic and they do get it. They think it's insane what's happening, but they don't think we should obviously stop building AI. And I think the second group is kind of causing the third — people who sort of get it, understand it's a really big deal, but also don't think we should stop.

How many people have seen the AI documentary? I liked it a lot. I encourage you to watch it. But the thing I really didn't like about it is that at some point they were like, well, look, we have these two options — we either lock it down or let it rip. "Let it rip" is basically like, let's just see what happens, not try to control or slow down the technology in a meaningful way. The alternative is basically some sort of mass surveillance, concentration of power thing that also sounds really sketchy. The documentary maker says, wait, why can't we just stop? And then you have a bunch of people from the AI safety community saying, oh no, it's an arms race, there's international competition and China, et cetera.

But I think we've to a large extent mistaken a domestic political problem for an international game theory problem. The international game theory is real, but it's not the crux of it, actually. So the rest of this talk is going to follow this outline: how do we stop, and why.

The international game theory problem is real in the sense that countries need to be confident that their rivals aren't building superintelligence in secret. If they're not confident about that, I think they will do it. Countries might be willing to go to war to prevent that. So we might be in a situation where it doesn't happen if countries sort of wake up quickly enough and if timelines are slow enough — because of some sort of mutually assured AI malfunction thing, or maybe just a big war that sets everything back in time. So it's not a foregone conclusion.

The main thing I want to say is that in some sense this really is the main blocker. When I look at the problem, the thing that makes it seem hard is: governments need to be able to trust that their rivals aren't building it in secret. But I actually think that this is not as hard as it looks from a technical point of view. On the ground, I believe we have a solution to that. It's more of a political problem, and if the US government were to prioritize stopping AI, I think there's roughly a 50% chance that that would be enough.

I think the best option is actually the third one, which is to systematically dismantle the AI compute supply chain. Basically, get rid of the advanced AI computers that are causing this whole problem — a Butlerian jihad kind of thing. If there's no computers, there's no AI. Get rid of the chips, get rid of the factories that make the chips, and basically make a concerted effort to actually lose the institutional knowledge that currently exists at places like TSMC.

This would be quite easy to do for the frontier because it's very concentrated — there are multiple choke points. There is this question of how far back can we go and how far back do we need to go. A really key piece of this is knowing where as much of the compute is as possible, so that if and when governments do decide that they want things to slow down or even to stop — and by that, I mean an indefinite global pause, indefinite, not forever, but not going to expire before we want it to — they can actually get rid of the majority of their chip stockpile.

I think this should happen in a systematic way. That's important — this is the way to make it durable and actually set things up such that we continue to work on and make progress on this as we need to. So that could mean further restricting the compute hardware, or relaxing those restrictions and saying, actually, we think we've set up a regime where we solved enough problems that we can now proceed building these kinds of chips and AI systems.

There are some issues with this. If secret projects survive, the good guys who actually follow this agreement are in trouble. It doesn't address algorithmic progress — but I think if you go hard enough on the compute, then you can address that. And of course, we have to be willing to seriously delay the benefits of AI.

A couple other options that are maybe more commonly discussed but I think are not as good. Option 2 is hardware-enabled mechanisms. There are two ideas here: one, know where the chips are — put something on the chip so it phones home and you can track where they are. Two, restrict the computations on chips — the version I'm most optimistic about is just whitelisting certain computations that we're very sure are okay, like doing inference on existing models. I think there are good proof of concepts for these, but they haven't been deployed in practice. And mostly this is not really aiming for an indefinite pause.

Option 3 is MIRI's plan involving FLOP thresholds — same idea as option 1 but more restrictive. Less than 16 H-100s except in monitored facilities, restrict dangerous research. This is more of an indefinite pause approach. It's a white paper with less engagement so far, though there are some conversations happening outside the public sphere.

The issues I see with these plans compared to mine: first, we still have a bunch of chips lying around, and people can potentially remove the hardware-enabled mechanisms or do dangerous things we didn't anticipate. If we know how to make chips with hardware-enabled mechanisms, we also know how to make them without them. So it's easier to set up a secret fab. AI is still available to some actors in these plans — that's tricky both technically and in terms of setting up good governance and balance of power. It also makes the agreement more fragile: if a country just says, "you know what, never mind, we're going to use these chips to build superintelligence and take over the world," other countries may have a very limited window of time to react. If you have to actually rebuild the supply chain to produce the advanced chips, that takes years.

So we need to make sure nobody builds such things until we fix this problem. Other than stopping AI, what else could we do? We could try to figure this out in time. We could try to buy a bit more time — just pause for a couple of years to do some more alignment research. Or we could coordinate to not build this kind of AI while still building other kinds. I think these are all unacceptable because they leave too much risk on the table.

I think a lot of people when they think about this problem are thinking about how they can reduce the risk on the margin. That's an interesting way to think about it, and there's some value to it. But I think we're making a big mistake by doing that so much of the time. We're sending the wrong message to the rest of the world and behaving in a somewhat uncooperative way as a result.

What if we were at the "post-Martian takeover workshop"? Would we be having such a workshop if we were like, so obviously the aliens are coming and they're going to take over — the real question is what are we going to do after that? There's also the fact that a lot of people are working at AI companies. I think that's very confusing for the rest of the world. It's like, "I'm working on climate change at ExxonMobil" — what?

If we want to coordinate with people — which is obviously important for stopping AI, but also a lot of other plans — I think we need solid principles and inspiring visions. "By working with the baddies, I've reduced X-risk from 18% to 14% — now only less than a billion people are expected to die in expectation from their activities, isn't that great?" is not an inspiring vision for people. Thinking on the margin is often uncooperative — this is partially what got us the failed "evals and voluntary commitments" regime.

I don't think we're doomed — I'm not a doomer in that sense. I just think we need to look at the problems in order to address them. But we can't rely on getting lucky. Timelines could be very short. Takeoff could be very fast. The attack/defense balance might not be in our favor. The MIRI worldview — that you just can't make these things powerful without them becoming insatiable optimizing agents that are hard to control — might be roughly correct. Evidence from current systems doesn't give us enough of an update to eliminate that possibility.

We can't rely on solving superalignment on schedule. The loss of control, if we do lose control, may be irreversible. Maybe a few AIs go rogue and all the other AIs can team up and get them back under control — but maybe not. If the AI is smart enough and out there in the world somewhere, maybe it can establish a base of power and multiply, and the offense/defense balance is not such that it's easy to find and eliminate all copies of an adversarial superintelligent AI.

I think we should actually be talking a lot more about the assurance problem versus the alignment problem. It's important not just that the thing is safe, but that we know that it's safe — otherwise we are in effect rolling the dice. And setting all that aside, we still have gradual disempowerment and concentration of power, which there's a lot of ideas about, but not really plans that people have a lot of confidence in. And of course there are unknown unknowns.

Another question people have: why do we need to stop now? Can't we wait until things are a little bit further along — maybe do a temporary pause and get more valuable alignment research done?

The basic picture is: we're on a train and it may or may not be racing towards a cliff, and there's fog, so we don't know how close we are. What do you do? The obvious common-sense thing is to slam on the brakes. We don't know where the cliff is — progress could be very sudden. It takes time to slow down. Even if you want to stop in the future, you may not be able to stop at the point you want to. And there are a lot of things making it harder to stop over time — more proliferation of models, more integration of AI in society, AI companies getting more powerful, new techniques being developed, especially recursive self-improvement.

I want to zoom in on the "we don't know what capabilities are dangerous" part. A lot of people believe we have a good enough grasp on this. I think that's not the case.

Can you imagine seeing evidence that AI is getting too dangerous and we need to stop? If you can imagine seeing that evidence, then we may be in that situation in the future if we don't stop now. And we may not be able to stop once we're in that situation. So there is this risk that in the future you'll be like, oh damn, we need to stop — oh, we can't anymore.

I think a lot of alternative plans are harder than stopping. Basically, global regulation of some other form is generally going to be harder. You're still going to need to know where at least a bunch of the chips are — and if you know where they are, you could get rid of them. There's going to be more chips and more fabs making it harder to know where the chips are. You're also going to need to solve more political problems in order to get an agreement in place. And stopping is more robust to violations.

There's also a keep-it-simple-stupid principle that I think really applies here because the amount of time we have to do something like this is likely very small. You compare this to things like nuclear weapons and global warming where we have had some significant amount of international coordination, but it took decades. From the point of view of talking to the public and policy makers, getting agreement on what the rules should be — I think it's going to be very hard to do even within a society if it's not something simple and easy to grasp that clearly solves the problem.

The real sticking point for a lot of people: yes, I agree, if I could snap my fingers and stop it, I would. It's just never going to happen — it's politically infeasible. I think the situation looks a lot different now than it did a year ago, and this was fairly predictable. When I was talking to people about a year ago, they were like, I don't think AI is going to be a big issue in the 2026 midterms. I think it's pretty clear that it's shaping up to be a much bigger issue than most people were expecting.

The AI backlash is growing dramatically. People don't have to buy the X-risk picture — they can just be like, wow, this thing is crazy powerful, I don't want these companies I don't trust or this government I don't trust to have all that power. Also it sounds like I'm going to lose my job — then what? The economic costs of stopping are another big issue, but the costs are really just totally worth it based on the amount of risk that we're mitigating. And the simple thing is a much easier political target — it's easier for people to get behind.

An important part of my view is that there's a narrow slice of political will where you can get some other really serious regulation, like compute governance, where you don't have everybody just clamoring and demanding the full-on Butlerian jihad thing. Once people are seriously worried about this, stopping just becomes the obvious thing. So there's not a whole lot of worlds where I imagine we can do something a little bit weaker but leaves more value on the table.

In terms of what we want people to do, the biggest thing right now is just talk about the issue — because it's still not being treated as a serious issue enough in the public discussion. There's a bit of a taboo that we need to break, like we saw in the research community and we're still seeing in politics. And then of course, be politically active, and also resist AI as it shows up in your life and your community in ways you don't like, or just out of strategic reasons in order to increase our bargaining power against the AI companies.

I sort of view Evitable as sitting at the intersection of three circles: the broader anti-AI or AI-skeptical, critical AI resistance movement; the Stop AI or Pause AI or Shut It Down abolitionist movement; and AI safety. An immediate next big goal on the movement-building side is to figure out how to get grassroots transmission of what I view as the core DNA: understanding the scope, the depth, and the urgency of the problem. The scope is that it's roughly as bad as X-risk. The depth is that it's a very fundamental problem — you can't solve it with Band-Aid piecemeal solutions like UBI. And the urgency is that it's happening now, you need to act now. I think those lead to the obvious thing of shutting it down, in most people's minds. Nonviolence is an important part of this DNA as well.

In terms of coalitions, I think we are thinking about movement building. But also, what we're trying to do here is more like steering. Because I think by default these coalitions are emerging — the sentiment is growing rapidly. By default, I think they are likely to be steered in less productive directions and satisfied with solutions that I don't think are really sufficient and fundamental and won't reduce the risk to an acceptable level.

That's it. Thanks!