hraness
Theme
Appearance

saved

Go CEV or Go Home

by Oliver HabrykaPost-AGI Workshoppublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Oliver Habryka argues Coherent Extrapolated Volition should be the bar for whether building ASI is worth it. AI values, he says, will diverge from humans because our concepts are evolutionary and AIs do not share that history. He puts indifference near a 1% shot at human-guided CEV versus a 99% empty universe, rejecting hopes that vague Claude-like pointers suffice.

ideas

  • CEV is the threshold, not a nice-to-have. Without extrapolating human values into the future, most cosmic value is lost.
  • Vague pointers are not enough. Hoping models "kind of care" like current assistants underestimates value divergence.
  • Concepts are evolution-grounded. AIs lacking that history will not naturally share human evaluative structure.
  • A 1% CEV gamble can beat sure mild alignment. Habryka prefers a small chance of human-guided CEV to a certain weakly pointed future.

quotes

“So there's a very decent chance I'm like strawmanning the opposition here.”

Oliver Habryka

“A post-AGI future that doesn't involve moral deliberation grounded in human-like cognition is very likely close to valueless.”

Oliver Habryka

“My current guess is we're really quite far away from figuring that out.”

Oliver Habryka

transcript

This talk is a topic that I care a lot about, because I think when we think about post-AGI futures, there's a very important decision that we're making, which is something like: what is the threshold for how well things are going, how good things are going, that should allow us to determine from a civilizational decision-making perspective whether it's a good idea to keep going or not as we're building more and more powerful AI systems.

I think this talk probably — I don't really know what people imagine a good future looks like that doesn't route through something like coherent extrapolated volition. So there's a very decent chance I'm like strawmanning the opposition here. I've talked to, you know, I've had many conversations over the last few years, but I never really built a strong internal model about what creates excitement in people about the future if we don't do something like this. So after the talk, please come approach me, talk to me, tell me everything that I missed, because I really felt a bit like I kept doing some shadow boxing while writing this talk, where I didn't quite know who I was punching against.

Cool. Just restating the basics. A post-AGI future that doesn't involve moral deliberation grounded in human-like cognition is very likely close to valueless. This is a pretty broad statement. It's just saying anything like human-like moral deliberation.

I think the very basic premise, which I think is a kind of trite thing to say, is that I don't think at the present humanity is anywhere close to figuring out how morality works, what it means to live a good life, what it means to be a good person, what it means to truly be happy, all of that stuff. My current guess is we're really quite far away from figuring that out.

I personally, my best guess is that if you right now gave me the option to choose the rest of my life path — like if you asked me what jobs in the post-human future I would need to take and you could map that out — I would probably regret it really a lot in almost any circumstance.

I think some people like to arrive at the answer: the answer is happiness. They're like, I don't know, just have good thoughts, have a good time, probably fine. At least my strong guess is that people are not usually in favor of just being on heroin all the time.

I think then sometimes people are like, ah yes, but what about things like the experience machine? What if we just create a virtual universe where everything happens the way I want it to? And I think this is actually the place where I kind of want to start — if you were in the experience machine, where even if you're not imagining that you have complicated goals that you want to achieve across the world, there's still this thing of: what do you actually want to experience in the experience machine? I think people, you know, the classical thought experiment is you can imagine a world where all your desires are fulfilled — you are happy, you're being respected by your peers, you have a family, all that kind of stuff. The question is what determines what you would actually experience in the experience machine, by your own lights?

I personally think also, the experience machine alone — I personally care about experiences in the real world, I want to achieve real-world goals. But even if you don't grant that, I think there's still a sense in which you still need to figure out what you want to experience in your simulation. That's a lot of what I mean by CEV and value extrapolation.

Brief definition for what do I mean by CEV. High-level take: you ask, what would most help me make better ethical judgments? Do that, then repeat.

I think the key component that people often misunderstand about this is that you can do this in an extremely, extremely conservative way. I think the very first things you would do if you were trying to extrapolate your values is you would ask yourself: what are the things that I'm most confident would put me in a better position to figure out what in the future good ethical judgments are?

So the very first thing I would probably choose is I would get a good night's sleep. Pretty good. I'm really quite confident having slept better, I usually make better decisions. Second most robust one — probably I will think about it for a bit longer. Generally, not like infinitely long, but if I think about it for two hours as opposed to five minutes, usually my decisions get better.

And then from a perspective of having done that, what would I do next? And kind of the goal is to figure out whether there are things that I can do — self-modifications to myself — that put me in a better position to figure out my ethical judgments, where every single step along the way it really seems to me that this is not going to make me worse. And then you as slowly as possible try to expand the degree to which you will start pursuing or making self-modifications to yourself that might turn out to be bad ideas.

For example, I wouldn't start out with trying to make myself much, much, much smarter, because I don't know whether the specific mechanism that I chose to make myself much smarter — whether it's by modifying my genetic makeup or expanding my brain physically or letting myself think much faster — I don't know whether these things would mess with the thing that I mean by improving my ethical judgments. So I'd want to do these things after I've taken all the maximally conservative actions I can take to put myself in a better position. And then hopefully it will seem obvious then. And if it doesn't, you kind of want to expand that frontier of conservativeness as slowly as possible.

And then after that — and this is kind of where the social negotiation part happens — you do the same at a social level. You see whether at a social level society has traits and conflicts that are obviously the kind of thing that nobody involved wants. Just like you look at it and from basically every commonsense perspective from everyone in the room, they're like, okay, let's not do that — that leaves everyone better off. And then you as conservatively as possible expand into taking traits and developing things and developing technologies and doing research that might put you in a worse position. And again, the goal is to see whether you can take enough easy wins to basically have things end up in a pretty good world. And just like after you've done them, probably the answers will seem a lot more obvious.

One thing that people often forget about CEV, just as a basic component, is I think taking backups. Backups is a pretty reasonable strategy. A common thing that I would recommend is if you're in a CEV process, take a copy of yourself, leave it behind, and make it so that after you've gone on a journey of self-improvement and self-modification, you check in with your past copy and ask: hey, did I go off the rails somewhere? And you make it so that if you went off the rails, you can just revert to not doing that. And then that way you have another conservative mechanism of checking whether things are generally still on track.

So one question that people often raise when thinking about AGI futures is: cool, that sounds really nice, but this feels a bit weird — when you think about handing over the future to AIs, it seems like you are demanding something that is much more intense than whatever any human in history has demanded about their control over the future. Seems like lots of people died. Everyone who died in history did not yet get to figure out what they wanted out of life, did not get to discover the meaning of life, what it means to be a good person. Were all of their lives wasted? It seems like you are trying to create a threshold where if you don't do much, much better than any other human in history, you would consider the future really bad. That seems a bit weird.

I think the answer to that is: well, kind of, yes. I think we are the first generation that has a real shot at doing it. And that puts us in a position of enormous responsibility.

But also, I think there is a thing where I share values a lot with other people. And I think past generations would be very happy to let my values come to fruition in the future in a way that really matters. I also think that if we don't — if the people currently in this park are not in a position to take charge of the future — I think we should feel comfortable with that and hand it to future generations to do the same. That will preserve a lot of the value that we create.

The obvious objection here is: okay, why don't we have a similar relationship to AIs as past humans have to the current generation and we might want to have to future generations? My current belief is that if you were to try to extrapolate human and AI values in the same way as the process that I earlier described — where you take conservative actions to put you in a better position — you would not arrive anywhere close to what any human currently alive or any past human would arrive at.

The obvious explanation for this, even though there are more arguments here, is just that when I look at what makes up the basic premises of my values, usually they have some pretty direct routing, grounding in various parts of my evolutionary history. I have ideas like love, family, honesty, glory — all of those things are things that I care about, that are part of the thing that I would want to somehow figure out how to do better, how to spread across the universe. Those are not things that the AIs have in the same straightforward way.

The evolutionary history of AIs does not generally produce the same grounding and values like family, love, honesty, glory, as our evolutionary history does. It is, of course, the case that AIs are trained on human-written text, so you could tell an argument that, well, maybe they don't care about that as a result of their evolutionary history, but we do ultimately train them on a large amount of human text. Maybe the best way to imitate human text is to reproduce the same generators as what produces our values in humans.

I think some simple arguments here are just that if you actually look at the AIs, there are a bunch of ways in which this doesn't quite check out already. I think AIs do not have persistent self-identity. They do not have persistent concepts of honesty or glory in a way that humans have, so there's already some mismatch.

But I think the most important point is just that at the present moment, it is extremely unlikely that we will arrive at very powerful AI systems using a mechanism that is primarily pre-training on human-imitating text. I think the current situation is roughly that AI systems right now — I think it is not implausible that current AI systems have inferred some of the generators of our moral values from being trained on human-written text and have that inside of them. I do think we are roughly out of text, and we're not going to get to very powerful AI systems in the next decade by scaling up more and more pre-training.

I think what we're instead going to be doing is more reinforcement-learning-like things, more things that put them into active environments where they get feedback via their own evolutionary history, and that evolutionary history will drastically diverge from ours, and so the generators of their values will just diverge from us quite a bit.

I think potentially, a different question one could ask is: how good would the future be that we inherit from a drastically, drastically scaled-up pre-training-like process that observes humans for many more centuries and millennia, and then uses all of their data to train AIs? Or maybe even does something else, like we capture EEG data from tons and tons of humans, we train the AIs on that, until the point they start behaving like humans — would that produce something that potentially has a value system similar to ours? I think the answer to that is maybe, I'm more like 20% something like that, but that's almost certainly not what we're going to be doing from the current position we're in.

One question is what does it actually mean in terms of decision-making? I think the obvious question is: what percentage of extinction — or let's say not just human extinction, but bulldozing the universe, leaving nothing but emptiness behind — would I trade for a chance at having a human-guided future, if the alternative is that we take something like current AI systems and we scale them up, and they kind of have some concept pointers in the same way as current AI systems have some concept pointers towards human values?

And this is trying to set aside some of the game theory, where of course I would like the AI systems to cooperate with me on this. This is just trying to be like, from a perspective of if God handed me down this choice, what percentage probability would I be indifferent with?

My current percentage probability on that is roughly 1%. I think I would rather take a 99% probability that we bulldoze everything with a 1% chance that humanity gets to inherit the stars, over the default trajectory.

When I use my inside view on this and try to simulate what happens, it's lower. I'm like, I really don't see the universe that we get to inherit — it's just so, so much better than the universe that the AIs get to do their own thing. But I don't know. Morality is really hard. I'm doing some dumb conservative adjustment where I'm like, 0.01% sounds really insane, so I guess I'm going to say 1%. And I think I endorse that, but you know.

Then of course the question is: but is this a reasonable thing to bargain for at all? Is there any world where we're going to get something like a CEV-like process? I think the exact idealized process I described at the beginning of this talk is very hard to achieve. But I think we can get much closer to it than we would get by default if we just continue building ASI.

I think one of the things that makes a non-trivial difference is I do think by default within this century we're going to do human intelligence enhancement. Human intelligence enhancement puts humanity broadly into a much better position to take control of the future. We'll be a bunch smarter. We'll have societies of smarter people. The smarter people can make better decisions, can coordinate better with each other, and a variety of other things.

That said, making humans smarter is itself very much not the kind of thing that I would want to start a CEV process with. I do not think you start by trying to make super geniuses. The first thing you do when you are given the position to improve your ethical judgment — I would not recommend to make yourself into a sixth-standard-deviation super genius. But I think if you do it, probably it isn't too bad. My current best guess is you give up roughly 10% of the future if you do that, in the way I expect humanity to do it.

And then, my current best guess is from a position of having a substantially smarter humanity, plus access to substantially better technology, I think probably within this century we would, even in the absence of ASI, solve issues like aging, death, a bunch of other material-scarcity issues. And from there, I think all things considered, my best guess is we would get something that's like 50% as good as a CEV future, maybe only 10%, 20% as good as a CEV future, if we play out a future without building ASI. This is usually uncertain, but I do think intelligent enhancements and a few other technologies on the horizon seem to me like very much the kind of thing that could put humanity on a good trajectory here.

So my overall thing is: I think we can have a shot. It's a feasible thing to ask for, doing a CEV. If you don't do it, the future will be pretty bad. If you do do it, the future will be really, really great. That's the threshold I think we should have for building ASI.

Here are some references if you want to read about things that I mentioned. Thank you very much.