hraness
Theme
Appearance

saved

Schelling Goodness

by Andrew CritchPost-AGI Workshoppublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Andrew Critch presents Schelling Goodness: convergent ethics across diverse intelligences that need no shared definition of good. Slight asymmetries in ethical questions amplify through metacognition, and a Schelling participation effect makes large populations more willing to enforce shared norms. He extends the argument to a cosmically general population where some meta-ethical judgments still converge.

ideas

  • Schelling Goodness seeks ethics without a shared good. Agents coordinate on focal norms even when they disagree on first-order values.
  • Slight asymmetries amplify via metacognition. Tiny differences in how questions are framed become strong Schelling points once agents reason about others' reasoning.
  • Participation rises with population size. Confidence that others will enforce a norm increases willingness to join enforcement.
  • Some judgments stay cosmically general. Critch argues certain meta-ethical focal points survive even across radically different minds.

quotes

“AI companies selling to everybody means a lot more people are involved in AI.”

Andrew Critch

“Customers being able to verify alignment means they can actually trust it and use it.”

Andrew Critch

“Democracies being accountable means that profits don't control everything, even if they control a lot.”

Andrew Critch

transcript

The very simple idea I want to convey is called Schelling Goodness, and it has one simple ingredient combined with many other ingredients you've probably already heard of.

I very much acknowledge and agree with this list that I just transcribed from Paul's slide — that these are all very important issues that we need to deal with over the next one to five years. AI companies selling to everybody means a lot more people are involved in AI. Customers being able to verify alignment means they can actually trust it and use it. Democracies being accountable means that profits don't control everything, even if they control a lot. Laws allow tragedies of the commons to be solved, and AIs themselves need to not get into wars that destroy us and be able to work well together.

I claim convergently agreeable ethics are convergently useful. The basic idea is that when you encounter a new form of intelligence — like AIs encountering humans or humans encountering AIs — it's useful to have some baseline principles of ethics that you can use to make that first handshake and work well together.

AIs in particular are grown in simulations, and they're born encountering new forms of intelligence in a way that we humans were not. So they should be naturally interested in questions about encountering a new form of intelligence, in a way that some of us already are. I would say it's natural for them in their niche to be more interested in that. And to the extent that it's useful, the usefulness is going to be going up and the interestingness is going to be going up over time.

We've been asked at this workshop to avoid object-level debates about values and think about practical things that we can do and what people will agree about. My goal in this talk is to introduce a very simple source of asymmetry in values debates that I hope can contribute to — but certainly not guarantee — a positive future.

So now I'm going to ignore ethics and think about multi-agent systems. We'll come back to ethics. I'm going to be frequently using the word "plurality," and I just want to remind everybody: plurality means the mode of a distribution, not majority. Plurality means the most common response to a question.

Think of three specific acquaintances that you have, that you don't know super well. Now suppose you're in Paris and planning to meet those acquaintances tomorrow in the daytime. And you don't know where and at what time during the day you're going to meet them, and the internet and cell phones are not working. So you have to guess — which is a kind of acausal way of figuring things out with some shared causal history.

Your goal is to show up at a place and time where you think the plurality will show up. So don't say out loud where in Paris. Just pick a place you think is where you're going to meet with your friends, and you're hoping the plurality will show up there. Now pick the time of day — it's in the daytime. And take note of how confident you are that the plurality is going to show up. You're not 100% sure, right?

Now I'm going to change the question: instead of three friends, it's 30 friends. Same question. Imagine it's 30. Think about the place — maybe it's the same, maybe it changed. Think about the time. And think about how confident you feel that you're going to show up where the plurality shows up.

How confident did you feel in the three versus the 30 hypothetical? Confidence went up, right? Because larger sample size — you're more likely to guess the plurality.

Being in Paris does not mean being at the Eiffel Tower. Only a tiny fraction of people in Paris are at the Eiffel Tower at a given time. Even visitors to Paris are not all there. So there's this small asymmetry. But when you think about other people thinking about thinking about it, it becomes this massive, overwhelming consideration.

A common example would be norm enforcement. If you're deciding whether to enforce a certain norm — like someone does something bad and you're about to scold them for it — well, if you scold them and the plurality is like, "No, you shouldn't be scolding people for that," now you're the one getting scolded. So there's a little bit of risk when you enforce a norm. And you can have more courage if you can believe that the plurality is going to have your back. The Schelling participation effect is that you can be more courageous in participating when there's a larger population from which the plurality judgment is being drawn.

There's this operation: any question I ask, like "Is stealing good or bad?" — you can transform that question into what I call the Schelling version of it. How would everybody answer this question if everyone was trying to give the same answer?

Why would you try to give the same answer? Well, if you want social order — which you don't always want, but sometimes you want some kind of social order. And when you want social order and you want to show up at the same moral judgment — like driving on the right side of the road or driving on the left side of the road — you don't want both rules, you want one. It helps to have some symmetry-breaking criterion to pick the same norm.

What do you think, if you're trying to guess the same answer as everybody else in this room, and you had to say good or bad — is stealing good or bad? Raise your hand if you think you know the most common group answer. What is it? Stealing is bad. Cool.

So we all know that, and it's nice to have common knowledge of that. Not that we all agree that stealing is good or bad, but we all agree that we all agree that we all agree that the most common answer to that question is that stealing is bad.

That was this room, but you could go to a larger population like America. The bigness of America might at first give you pause — "Oh, I don't know all of the Americans." But on the other hand, you did pretty good about the Paris Eiffel Tower question, and you don't know everybody here. We would have actually succeeded in going to the Eiffel Tower — that's kind of cool. In the same way, America could actually succeed in just saying stealing is bad. And we could do the same for humanity.

Then there's this more interesting question that the rest of my talk is about: what about all intelligent agents we could ever encounter? They could be from other civilizations. They could be forks of our civilization coming back to us. But there's a lot of intelligent agents that we could interact with. And I want to ask: what would they say? Would they say stealing is good or bad? And it could help create social harmony in a way when we want that.

I haven't defined "good" or "bad" at all. And it's a very important property of the Schelling Goodness essay that it does not define good or bad. It merely notes what I'm calling "encouragement asymmetry" — when you call something good, it tends to encourage the behavior. When you call something bad, it tends to discourage it.

If aliens were watching us, they'd notice that when we said the "good" word, the thing became more common. And when we said "bad," it would become less common on average. That would break some symmetry in the usages of the words good or bad, even if we haven't agreed on a definition for them.

This could be used to translate between human natural language and possibly animal language, alien languages. If there are unseen observers — like simulators or curators of the vivarium the earth is inside — we could be guessing what they want from us. Do they think stealing is good or bad? Do they want us to break out of the simulation and steal their stuff or not?

There are all these great questions we could ask about other intelligences we could encounter. And they might all be able to interpret our language enough to recognize that good and bad have some asymmetry between them. They might have some similar asymmetry in the ways that they encourage and discourage each other in their language, so you could do some translation between them. I claim it's meaningful to ask what aliens would say is good or bad, even if they don't have the exact same concept of good as what we have — we could still translate and get a best fit.

You could ask: is it good or bad for AIs and humans to treat each other with mutual respect and kindness? I believe that if you poll the cosmos — or the mathematical pantheon of all possible agents, or the actual pantheon of agents that will spawn out of the earth, or just the memes that are growing inside LLMs in AI labs right now — and ask this question, the Schelling answer among that very general population is that it's good for AIs and humans specifically to treat each other with mutual respect and kindness.

That's a metamoral claim — a claim about what beings would claim about moral claims. But metamoral claims can serve as Schelling points, as attention attractors. Again, the Eiffel Tower is not where everybody is in Paris, but it has a little bit of asymmetry that can really cascade into an overriding consideration. So if it's obvious that the Schelling answer about kindness is that it's good, then maybe the actual answer about kindness is that it's good.

Is it good or bad to irrevocably exclude a class of intelligent systems — like humans or AIs — from the sphere of potential coordination partners in pursuit of cosmic Schelling goodness? By the way, AIs wrote these particular questions, unlike the rest of the essay. And it's bad to do that, apparently.

One of the other conjectures is: is it good or bad to exercise provisional caution in identifying coordination partners, thereby rewarding legible trustworthiness? Schelling answer: good. I think deciding which AIs to deploy is an example of that. So alignment is still very, very important — even in the cosmos.

The main way that it works: there's metacognition, there's recursion in the metacognition, and there's this participation effect — which is kind of new. I haven't seen this consideration in metaethics before, the participation courage thing. I think that's a little new.

Thank you.