hraness
Theme
Appearance

saved

Leviathan Risk

by Owen Cotton-BarrattPost-AGI Workshoppublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Owen Cotton-Barratt adds leviathan risk beside scheming AIs and human power-grabs as collectives gain AI leverage. As AI becomes epistemic infrastructure, corporations and states may wake up and press humans to cede authority. Drawing on Hobbes and The Narrow Corridor, he looks for escape hatches between anarchy and tyranny.

ideas

  • Leviathan risk is a third AI-risk bucket. It sits beside rogue agents and human coups as collective agents gaining power.
  • Epistemic infrastructure wakes collectives. AI that runs sensing and decision pipelines can make firms and states more agentic.
  • Humans may hand over authority under pressure. The path need not be a single villain; institutional gravity can demand deference.
  • Escape hatches sit between anarchy and tyranny. Narrow-corridor thinking asks how to keep liberty when leviathans strengthen.

quotes

“I'm particularly wanting to look at the case of collective agents.”

Owen Cotton-Barratt

“But how I think that with AGI, these things may become much more like coherent agents than they have been historically.”

Owen Cotton-Barratt

“I think this can happen at all kinds of different scales.”

Owen Cotton-Barratt

transcript

People might hear what I'm saying and say, "Oh, actually, I just think of this as a kind of subspecies of misalignment risk," or maybe "I think of this as a subspecies of gradual disempowerment." I'm kind of fine with that. I want to point to the category. I don't know exactly how to draw the boundaries.

I'm particularly wanting to look at the case of collective agents. But how I think that with AGI, these things may become much more like coherent agents than they have been historically. And so I'm using this metaphor of waking up. These things are kind of like slumbering giants. And I think that they may become more oriented, looking in sharp ways at what is happening in the world, what is good for their interests, and more taking actions as a result that are pretty aligned with what is going to be best for those interests.

The core pattern of why AGI might push in this direction is that if you have AGI analysis — high-quality analysis of what is the situation, what's going on, what will be good by the lights of the collective — then that means there can be less space for decision-makers to exercise their own judgment, to have slack around what they're deciding.

I think this can happen at all kinds of different scales. It could happen at the scale of a company, where right now CEOs generally have an enormous amount of latitude to pursue things according to what they think is best. Occasionally you get activist investors saying, "Hey, we don't think this person is doing a great job." But if everybody could have just on tap extremely high-quality reports saying, "Look, here's the strategic situation for this company. Here is really the thing which would be better for the bottom line" — it would create much more pressure on the leadership to comply, or have a really good refutation, or much more quickly people are going to move to change leadership.

At a bigger scale, you can look at a state. Imagine a strongly ideological state, maybe a one-party state. As people make decisions, there would be more pressure towards making the decisions that are really aligned with whatever the stated purposes of this party and its beliefs are. Because it's more legible when decisions that people are making are departing from what is actually going to be seen by reasonable, very smart analysis to be in pursuit of that.

This could also affect who gets promoted within that party. You'll have sharper filtering. Everybody will be able to see, "This person, this candidate, maybe they're not properly likely to actually follow through on the party's beliefs." And so now you look bad if you end up supporting this person.

Why worry about this now? We've had collective agents for ages and ages. And maybe when you were first starting to get these things, you might have thought, "Maybe we should be worried about governments, companies — maybe this is just going to eat the humans inside them." And from our perspective now, it looks like it turned out OK. Maybe there were some bad cases — you had tobacco companies with incentives to push falsehoods on people in an emergent way just from the incentives. Or maybe there have been destructive ideologies, which have been pretty bad for a whole bunch of people. But overall, if you look at the world, we've had these collective agents for quite a while and it doesn't feel like it completely crushed all the people. There's still people having a bunch of independent thoughts and generally being more empowered. Maybe this is fine.

If I ask, "Why did these collective agents not end up more powerful? What are the limits?" I see a couple of different caps. One is they are constrained by relying on the cognition of the humans that make them up. This means they're not radically smarter than the individual humans. In many cases, I think they're dumber, because they can't access the full smartness — or at least they're dumber than the smartest individuals within them. Also, they have to deal with human conscience. If people have qualms about doing something, that introduces friction.

And then also, if something becomes super powerful, maybe it just gets captured by humans. In the 20th century, various communist governments came to power — in a sense it looked like an enormous amount of power was being held by an ideology. And then in many cases it got captured by some dictator who was then effectively running things. The collective agent lost its power to the particular human.

I think AI could reduce the effectiveness of both of these caps.

The big deal is AI as epistemic infrastructure. Things that we can build that could help people track what's going on, do conditional analysis, gain strategic understanding of where things are. I think there's a lot you could do already with basically pretty near-term AI systems. But for the perspective of looking at post-AGI, I don't think we need to care about how quickly we'll be able to do this. It's pretty obvious that at some point, if we get this AGI thing, we're going to be able to have really high-quality analysis on tap pretty easily.

This could happen at the individual level — everybody just goes and consults their own version of a chatbot, only now it's really super. And it tells them, "Hey, this strategy you're considering for the organization won't do so well." That already helps to create common knowledge. If somebody is pushing for a strategy and everybody can see because they talk to their chatbots that it's a dumb idea, people are going to challenge it.

But you might get something even stronger than that. The organization may commission more in-depth reports — "Hey, we're the organization for making lots of paperclips. What would be the most effective strategy?" They spend more on compute to have deeper analysis. And that becomes common knowledge within the company.

So you've got this thing which is like a super high-quality, very cheap management consultancy report, which has gone into a lot of detail working out what would really be good by the lights of this organization. Maybe there are still some decisions on the edge where people can exercise judgment. But there's a lot of constraining, narrowing the options that look reasonable. And if you try to push for something that isn't in that set, it starts looking like defecting. "What are you doing? You're going against the interests of this collective."

This has the effect of this kind of waking up. The collective starts looking more like a rational agent, with coherent actions towards some goals. I've said "around stated goals" here — maybe the things it's effectively pursuing don't exactly align with the stated goals. Maybe there's some kind of emergent complexity in what it ends up going towards. I'm not exactly sure, and I feel like this is some important thing to work out. But I think the general pattern — these collectives end up acting more as coherent agents — looks pretty robust, just by making intelligence cheaper and more readily accessible.

Where it gets especially alarming is: if you are this collective and you have preferences about what happens in the world and you want things to go well, the instrumental sub-goals start biting again. You start looking at, "Well, what will affect this? What should I be caring about?" And one natural thing for the collective to care about is its own internals. "I'm an agent who is trying to make this stuff happen. Wouldn't it be better if I was more effective at my goals?"

What does making itself more effective look like? It'll partially look like improving internal processes, but it will also look like looking at the humans inside the collective agent. And now it's having preferences over those humans. Maybe it has preferences about who is selected into positions of power. Some of this is just about rewarding competence in ways that we'd say, "Of course, this is completely normal and healthy for organizations to promote the people who are competent." But some of it might also be about selecting based on who is the true believer in the paperclips or whatever. And maybe it also shapes incentives for people — if people see, "This is how I need to be in order to be selected by the system to gain power," that may distort them away from what they otherwise would have wanted.

Getting slightly wilder, these collectives could want to distort the beliefs of their own internals. I say slightly wilder, but states have propaganda for their citizens — this is a relatively normal thing. I think you could have it happen in more cases, depending on the relative power of the central body to manipulate versus the people within it to detect manipulation and to work out what's going on.

You might think, "OK, just have a collective which is kind of good and doesn't want to screw around with the people making it up." And the reason I think this is not fully reassuring is that there may be other collectives out there, other people trying to do different things. And they will also maybe have preferences around, "Hey, what's going on with this first entity? Can I mess with that?"

You see states in the world today and historically sometimes trying to mess with the internal governance of other states. They naturally have preferences about what's going on there — "Here is a way to achieve more of the things that we want."

I don't know exactly how strong this dynamic is, but at least there's a pressure: if you naively say, "Let's not model what's going on with our internals," maybe you get outcompeted by the other things out there who have no such qualms. So maybe you at least need to do some modeling of your internals in order to protect them. And then potentially there's a slippery path — if you're doing some modeling, maybe you start regarding anything which is taking things away from the directions that you most believe in as bad manipulation, the type of thing that should be avoided.

This is almost an aside — pressures toward handing over formal authority to AI systems — but it is part of my picture. I think there are several different reasons why there could be pressure to start handing over formal authority. Obviously there are reasons why handing over formal authority is alarming and people might not want to do it, but I think it's worth understanding the positive pressures.

First, there's an efficiency thing. Things are competitive. Humans slow things down or make dumb decisions sometimes. Maybe you just get better by cutting them out of the loop after your AGI is good enough.

Second, maybe there are political reasons. You might be trying to set up a new coalition with people and you might think, "I don't totally trust Bob. Who knows what he'll do with the power. But we've kind of got this alignment thing sorted, so we can just empower this AI and trust that it will actually pursue the objectives that we're setting it up with."

Third, if you're worried that having power makes you a target for others to manipulate you, maybe you can move this target by actually giving away your power. And if you weren't really using your power, if you were basically only ever rubber-stamping the high-quality AI analysis, maybe you'd just be like, "Yeah, I can deal without having this. And maybe that's better."

So I think there are reasons you might move from these collective agents being still formally in the hands of humans — but increasingly enacting just whatever is rational by the lights of the collective — to it's not even formally in the hands of the humans. And I think that transition might look pretty continuous from the outside. If the decisions were already being made according to whatever the high-quality analysis says is rational, you kind of don't care whether humans were doing the rubber-stamping or had formal authority.

One way you can avoid this savage ocean of people getting squeezed between all kinds of competing collectives is you just have one collective who's on top — the big boss. This is Hobbes' approach: have the state hold the monopoly on violence. And maybe you can have some kind of monopoly on manipulation.

It's a bit funny to say "monopoly on manipulation" because we don't really want the state actually manipulating people. I'm phrasing it this way for a couple of reasons. One, I want to highlight that if you do have one thing enforcing it, you should notice that it could be subverted — it could potentially do active manipulation. That's the thing you now need to really watch for. Two, it's not clear to me exactly where you draw the boundaries of what is manipulation. Maybe the monopoly should be on something like modeling humans. Paul was pointing to maybe you just want to ban modeling of people at all. But maybe it's hard to do this without some entity able to do enough modeling to recognize when there are shenanigans going on somewhere else in the system. So maybe you need some kind of monopoly power.

I think there are a couple of different possible escape hatches. On the one hand, you could have defensive technology — AI guardian screening. This is trying to empower the individuals, or possibly you draw the boundary not around individuals but around slightly larger groups. "Here is our little collective. We have AI guardians which stop any potentially alarming information coming in. You're not allowed to use tech that could manipulate others."

Or you could have a deeply liberal big state, which has overwhelming power but has a very deep commitment to constituents' freedoms and non-manipulation. That could be a good way to go if we can set it up.

Briefly stepping back, I'm trying to point to something here. This leviathan risk is not an acute risk. It's not "here is a single moment where everything goes wrong." It's more structural. Maybe we don't need to think about a single moment when power is seized by this new entity. This could just be that there were some existing entities which we already give very large amounts of power to, and they become more agentic over time.

Is it urgent to be thinking about this? I'm not sure. You might think generally with structural issues, you can just wait until they're becoming more of an issue, and then it's easier to build the coalition to address them. In this case, the reason I think it potentially is urgent is: maybe as these things start to wake up and start to be more strategic, they will have preferences against losing their power. And if we want to have a coalition which stops these entities from becoming too agentic or too powerful in the first place, maybe it's easier to build that coalition before it's clear exactly which entities will be the losers in whatever state we want to step to.

Thank you.