hraness
Theme
Appearance

saved

The Growing Gap Between ChatGPT and the AI Inside OpenAI - Noam Brown

by Noam Brown and Dwarkesh PatelDwarkesh Clipspublished

Hraness cites a source capture. The source author remains the source.

gist

Noam Brown tells Dwarkesh that frontier release cycles (often every two months or faster) are colliding with models that can already run week-long tasks and will soon stretch toward month- and three-month horizons. Labs still assume short safety eval windows from the GPT-4 era, so they may not be able to test full-horizon behavior before the next ship. Slowing releases to buy eval time would widen the internal/external model gap—already visible in math—while RSI could tempt labs to keep the strongest systems inside and stop external deployment altogether.

ideas

  • Release cadence outruns horizon eval. Two-month ship cycles cannot cover three-month agent horizons, so full-capability safety and product tests never finish before the next model.
  • GPT-4-era policies are stale. Safety programs still assume short eval windows and have not been updated for long-horizon agents inside or outside the labs.
  • RSI can freeze external releases. If three months of progress collapses into one, labs may skip classifiers and public deploys and recurse internally, concentrating capability.
  • Math is the early gap. An unpublished internal model already solves hard and previously open problems the public cannot reach, creating an unfair external lag.
  • Both slowdown and secrecy have costs. Delaying releases helps eval; it also deepens the lab-vs-world disparity, and Brown says he has no settled weighing of those trade-offs.

quotes

you don't have a way to evaluate the models at the full length of their capabilities

Noam Brown, stating the horizon-vs-release-cycle bind.

tremendous concentration of power by the end of the year

Noam Brown, naming the RSI-plus-internal-only risk.

math is the first domain where we're going to see that we're seeing this pretty clearly

Noam Brown, pointing to math as the early internal/external gap.

I don't have an answer for like how to weigh those trade-offs appropriately

Noam Brown, declining a settled policy on release delay vs access.

transcript

We're in the situation where the model release cycle is extremely fast, right? Like you're seeing new frontier models released like at most every 2 months, sometimes faster. Every week there's like a new AI breakthrough. And um people that look at AI, I mean, sometimes they they they last looked at AI like a year ago or 6 months ago and really dug into like what the models were capable of. And actually the models today are far beyond what was possible even 6 months ago. And so I think if people are skeptical of like a lot of these a lot of these capabilities, like I encourage you to just like try the models today and see what the frontier really is today. Um So we're in this period where like the model release cycle is very fast. And then we're also in this situation where the models are increasingly able to operate over longer and longer horizons. And I think this is an interesting scenario because we before we do any model release, we want to make sure that the models are properly aligned. We want to do safety evaluations. We want to do like very thorough stuff to like make sure that everything is like great, in good shape. Um This has been the case all the way since like, I don't know, GPT-4 or earlier. And implicitly there is this assumption that you can do these like evaluations in like a pretty short period of time. Um but if you have the models operating over longer and longer horizons, are able to operate effectively over longer and longer horizons. Like look, already you can have them like okay, GPT-3 you could loop it to do stuff over long horizons, you just wouldn't do very well at it. But today's models are able to actually do well at operating over very long horizons. Like you want it to do a week-long task, it can do a week-long task. Um We'll probably get to the point where they can do month-long tasks. We'll probably get to the point where they can do three-month-long tasks. If you're in a world where they can operate effectively over 3 months, but the model release cycle is every 2 months, then you don't have a way to evaluate the models at the full length of their capabilities before the model release cycle before the next model release cycle. And so there is this interesting question of well, what do you do in that situation? Like how do you ensure the models are safe and aligned in a period where like actually they can operate over these like extremely long horizons. And who knows, maybe the maybe the capabilities degrade. This isn't even an alignment issue. This is also just like a product issue that like maybe the maybe the product degrades over that time span in ways that like we have not had sufficient time to test. Maybe the alignment degrades. Maybe the the safety stuff degrades. This isn't an issue right now, but it is quickly becoming an issue that we have to figure out a solution for. And I think when you look at a lot of like a lot of the safety and and the policies were put in place in like the GPT-4 era where this was just like not on anybody's radar. Yeah. And it hasn't really been updated for a lot of companies. It hasn't really been updated since then to account for the fact that these agents are operating over these like very long horizons. And so um it is a situation that I think a not not enough people are considering both like within the labs and outside the labs of like how how do you deal with this uh how do you how do you prepare for this like problem that's going to like if you just look at the trend lines, we're going to hit this in like at some point.

What one concern I have is that during RSI, if the amount of progress that currently takes a 3 months happens in 1 month instead, um but you're not like the internal use case of AI is big enough that they're like, "Okay, we can just keep doing RSI. Why are we like going to go through all this extra work to build classifiers and safeguards and whatever and potentially take a bunch of like flak in order to like externally deploy this model. Why don't we just keep doing RSI stronger and stronger?" And so the the not only does the calendar time underrate the capabilities gap between the models, but the maybe like you just like stop externally deploying models altogether doing RSI cuz why do we want to help other people do RSI themselves with our models? You just end up in a situation with like tremendous concentration of power by the end of the year where right now it already is already the case. We'll talk about this with the Millennium Prize problem and other similar problems that the broader world does not have access to the models which are allowing for really cool things to happen, right? Um and uh they're going to be more broadly relevant than just mathematics eventually. They'll be doing more than just like coming up with cool math results. They'll be relevant to like um political leaders who need to make important decisions about the world. They'll be relevant to, I don't know, media of like what what's going on in the world? What do like what what what should uh the public be thinking about this? Um they're just economically relevant. People are running businesses, they want to use these models. And uh I think by default we just don't get the also the external deployment of AIs as progress speeds up significantly lags in qualitative terms the internal deployment of AIs.

Yeah, I think that's absolutely right. I think this is like, you know, it's it's tempting to say like, "Okay, these models are becoming extremely powerful, they're extremely dangerous, like they're operating over these like longer and longer horizons, and we want to make sure that they have we have sufficient time to evaluate them before they're released in a way that operates over those horizons. Um and so therefore the model release cycle should slow down. If we should have more of a delay between releasing models." Uh and there's a flip side to that, which is, you know, what you said, which is that, "Okay, well, now you're creating more of a disparity between what is internal to the labs and what they're able to use, what we're able to use, and what the outside world is able to use. And that that is also um not an ideal situation, right? It's like I think I think math is actually a good illustration of this. I think in many ways like math is the first domain where we're going to see that we're seeing this pretty clearly, where we have a situation where we have a very powerful model internally that is currently not available to the outside world, that is able to solve incredible math problems. It's you know, and it's not just, you know, Millennium Prize problems. Like we there are many solutions to unsolved problems that um people have been able to get out of this model. And there is a question of like what do you do in that situation and we don't have a good answer. Like It is It is a a situation where like yeah, that's that's a that's an unfair advantage. And um there are trade-offs here. I don't have an answer for like how to weigh those trade-offs appropriately, but like there there are Yeah, there's there's there's a complexity on both sides of this.

If you enjoyed this clip, you can watch the full episode here and subscribe for more clips. Thanks.