hraness
Theme
Appearance

saved

Vertical Alignment

by Raymond DouglasPost-AGI Workshoppublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Raymond Douglas proposes vertical alignment: collectives can hold goals that are not mere averages of members' preferences. Many failures are emergent Leviathan dynamics where no single actor seized power. The frame unifies those risks with what a good multi-level future would require.

ideas

  • Collectives have non-average goals. Vertical alignment treats group-level objectives as real, not just preference aggregates.
  • Peer-only views miss Leviathan dynamics. Horizontal multi-agent analysis underplays upward pressure from organizations and states.
  • Nobody need seize power for risk. Emergent institutional dynamics can disempower humans without a cartoon coup.
  • Good futures need multi-level design. The same vertical lens clarifies both failure modes and desirable equilibria.

quotes

“In this talk, I'm going to attempt to outline a basic theory of good futures and how we reach them.”

Raymond Douglas

“Well, the basic reason is that we want to try and figure out what a good future looks like.”

Raymond Douglas

“In particular, we would like to find good futures that are actually sort of stable and feasible.”

Raymond Douglas

transcript

In this talk, I'm going to attempt to outline a basic theory of good futures and how we reach them.

So why are we all here? Well, the basic reason is that we want to try and figure out what a good future looks like. In particular, we would like to find good futures that are actually sort of stable and feasible. Not merely beautiful moments, but some kind of enduring good, inherit-the-universe type thing. And if we can't do this, we would like to know that quickly, so that we can avoid racing into some horrible pit that's impossible to escape.

And where are we at right now? We have some guesses about bits of what might make for a good future. And we also have this massive laundry list of catastrophes that we desperately need to avoid. And gosh, there are loads of them. I think it's kind of wild to take a step back and think about how many different exciting ways there are that everything could permanently, horribly, unrecoverably go wrong.

And there's this interesting asymmetry where it seems like for a lot of the really bad cases, there's something sticky about them. We kind of get that if an AI were to take over the world and kill everyone, there's no coming back from that. And similarly, we can kind of picture AI enabling a kind of surveillance state where once a government locked in that totalitarianism, there would be no coming back from it. Conversely, it's a bit hard to articulate what it would take for a good world to be stable — what it would take for democracy to be so powerful that it was never going to get thrown off.

And yeah, it seems like the tacit default plan is maybe we take this big list of all of the problems, roughly rank-order by how much do we think this is a problem, how much do we have any leverage over it at all, how much can we afford to punt it or maybe get AI to help with it. And then if we could avoid all of those, hopefully things work out, or something like that.

And what I want to say is: what would be really great, if we could pull it off, is to have something more like a unified account. It would be great if we had some way of collecting together all of these different problems and fitting them in one picture, such that we could define all of them and also the good outcome in the same terms. I think that if we could do that, that would give us a lot more confidence that we had actually understood the deep underlying structure that unified all of the problems, rather than there being a bunch of patchy issues that we kind of understand each in turn. And it would give us a lot more reason to believe that the good future that we could articulate in contrast to them actually was a good future. I'm not at all confident that this is possible, but I am going to try in the next 15 minutes.

So okay, here is the first proposal for a unified account. My story is: it's all about power. What's going on with all of these risks? Well, the risk is we have power now and we might not have power later. And this is sort of definitional, right? An existential risk is losing power over the future. And for a lot of the classic ones, it's very concrete. The thing that's happening is that some other thing is rocking up on the scene and taking the power. We've built the AI. The AI has the power now. It is taking the power from us. The government has gotten strong. The government is taking your power. You do not have any power anymore.

And interestingly, I think that once you apply this lens, you notice that a lot of the wider debate is also framed in these same terms. If you think about what the classic pushbacks are against regulation or pausing, it's often also framed in terms of: oh man, we don't want the government to have too much power over how AI works, we don't want the UN to be controlling the world and bureaucratic hell or something like that. And so it seems like if we can articulate a lot of the considerations in those terms, maybe we would articulate the solutions in the same terms.

But then, I think this power thing has something real going for it. But if you try and just stay on that level and say, okay, the problem is power, the solution is power — and what is the solution? The natural solution that pops out is something like: we need to make sure that humans have the power, humans keep the power, humans have control over the world. And then I think that's just profoundly unsatisfying. And it's profoundly unsatisfying for broadly speaking two reasons.

What is wrong with this basic story of just "let humans have the power"? I think that thinking about power flowing between peer agents is a very good way to describe several of the problems, and also several of the most salient mainstream existential risk problems. If a misaligned AI takes over — very clear, you've got you're here, the AI is here, the power is going like that.

And I do think it's kind of interesting that this emergent dynamics picture seems to capture a lot of the extra stuff. And then once you look at it, this emergent dynamics picture seems like it captures a decent fraction of what's already going on. I think this is a place where people's intuitions really differ — how much of what's going on in the current world is a byproduct of emergent dynamics, coordination problems, whatever. But it does seem notable that we are currently, I think many people in this room would agree, incurring at least some possibility of existential risk. It's not the socially optimal amount. And there are lots of very powerful people who are currently actively building the very powerful AIs and saying things like, "Oh my God, this thing's really dangerous, I think it could kill me and my family and everyone I care about. Nonetheless, I need to work quite hard on it because of forces beyond my control and indeed beyond everyone else's control."

From that perspective, the idea that the thing that we're trying to do is "keep humans in control of the world" seems a lot weirder. Because if this is happening, we're clearly already not entirely in control of the world.

So okay, that's the setup. The setup is: maybe power is some unified story. And then here is my cluster of problems — okay, but what's already going on with us and the world is a little more complex than us having power. And here is my proposed solution.

Here is my best guess for the extra step that you can add that cleans up a lot of the picture. Basically, I think that when we are modeling situations, we are often doing it in what I'm going to call a horizontal frame. And what I mean by that is you pick some level of resolution, and then you analyze what's going on within that level of resolution. And it's very natural. There are loads of different ones you can pick. You can look at the situation and be like, okay, who are the different people here and how are they interacting with each other? You can look at geopolitics and be like, okay, what are the different states, what do they want, what do they believe? And you kind of regard the state or the company or the social group as being a bit like an agent — as wanting some things and believing some things. It's very useful, it's very natural. But typically you just pick one level and you think about what's going on at that level.

And then my bold claim that I want to make on top of that is that it is often useful to think of collectives as having their own goals. And I mean having their own goals in quite a strong sense. What I mean by this is: yes, we already accept that you can talk about what the US government wants. I want to say that what the US government wants does not secretly mean what the people in the government of the US want. I think it means something more complicated. Pinning down exactly what is a bit tricky, but I want to say it is not merely a weighted average of what the people in power want. It's not like we take what the president wants and give that 30% weight and then give 20% weighted by how much money people have and save 40% for the voters. The collective can want things in a way that is not a crisp weighting function of the members.

So if I sort of would guess, maybe 20% of you will be like, "yeah, obviously." And 20% of you will be like, "this is some crazy postmodern bullshit." Actually I want to check. Hands up for "yeah, obviously." Oh, brilliant. Hands up for "weird postmodern bullshit." Oh, okay, great. Well, that was easy. If you disagree and sort of don't want to, you can come find me later.

And I think that once you add in this vertical framework, you get a whole bunch of stuff. So since you all buy it — I will tell you what this gets us.

One thing is I think that this additional dimension on top of the horizontal angle actually makes it possible to talk coherently about a lot of the risks at the same time. Horizontal is sufficient to articulate many, many problems, like misaligned AI or whatever. Once you add in this vertical stuff, suddenly it becomes possible to neatly fold in a bunch of what seems like emergent dynamics. Sometimes the emergent dynamic is basically like a powerful thing but just extremely powerful optimization force without much agency. But sometimes it really is an agentic system that is co-opting people to get something done.

I also think adding in this vertical perspective gives you a slightly weirder notion of what value is and what the good is that we care about for the future. Especially sort of CEV stuff — it's very natural to be like, okay, we have people, people want things, the thing we care about is what people want. But actually empirically, if you look at the world around us, the way our current values get enacted is massively modulated by these larger systems, which do sort of interesting things to elicit preferences from us.

A great example of this is democracy. We all love democracy. But if you think about how democracies actually work, basically all of them heavily depend on the ways in which they restrict people's preferences. When you are building a democracy — as indeed people have in the past — they think very carefully about what are the ways in which we want to limit people's abilities to just decide what happens, what are the ways in which we want to structure this collective to elicit people's preferences in the right ways. And what are the things we want to set in place where, actually, we know that if we asked people, they would say that they wanted to re-institute the death penalty or something like that. So we're just not going to ask them.

I think this is interesting because it tells you something about the way in which even good things we care about are bundled in. Also, since you guys already buy the vertical thing, I think a spicy claim is something like: it makes sense to think of humans as also a bundle of preferences. And you can decompose and think about what the different bits of me are. Again, we talk about this intuitively — what they want, how they interact with each other, what they know. And then from that perspective, the sheer notion of extrapolating preferences gets a lot weirder, because the way I work depends on an internal balance of power that can get thrown off when I am augmented.

The sort of weird flip side of saying that some of the things we care about exist in these higher emergent forces is it also comes with accepting the fact that some things we don't care about and don't like also exist in these higher emergent forces. I currently bite the particular bullet that — to the extent that there exist collectives that are optimizing for things in the world, some of those things they are optimizing for are bad and I don't like them, but they're out there. And this complicates the sort of tacit impression that people sometimes have that: okay, what's the problem? The problem is that we're not coordinating. If everyone sort of held hands and formed a big ring around the world, then we would just pause the AIs or something like that. I'm like, yeah, you know, there's some irreconcilable differences of preference among people, but then there are also some irreconcilable preferences that are embedded in these larger agentic processes, and that's way harder.

I also think this makes it easier to talk about some ways in which humans are irrational. Why don't people want to believe true things? Because emergent forces don't want them to. Lots of large structures are basically bound around people having incorrect beliefs.

I also think this gives you some way to articulate a bit of what a good future would look like, which very briefly is: having enough agency on the top level of organization to have a thing which is kind of capable of basic self-preservation. If the kind of human agent was able to look around and be like, oh geez, if I build AI I'll die — just have it not do that. Which needs to be a little bit reciprocal. One way to tell if this is working is whether it is the case that we would pause if we were all about to die. This could look pretty bleak depending on some contingent facts about how hard different risks are. On the other hand, it might be necessary, and also AI might make it possible. Thank you.