hraness
Theme
Appearance

saved

They fixed AI man

by Less BitterLess Bitterpublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Less Bitter, who spent early 2026 calling AI coding a scam, says Claude Opus 5.5 on Ultra Code implemented Enjoy's 2,700-line multiplayer spec in about seven hours, including phone and browser pairing, a relay, and end-to-end encryption, without exhausting a $200 plan. He names the method evolutionary coding: tests and adversarial reviewers are the pressure, messy code can still work, and measurable production results beat line-by-line human review. He warns engineers who still block merges until a person understands every change.

ideas

  • A 2,700-line spec shipped in one session. Opus 5.5 on Ultra Code spent about seven hours pairing phone, browser, and desktop through a relay with end-to-end encryption, and Less Bitter says the result worked.
  • March is the failure baseline. Earlier agents claimed a shorter spec was finished after wrecking the repo; he treated that as the scam of AI coding until this run held the thread for hours.
  • Evolutionary coding keeps what passes tests. He treats messy agent code like a genome: the pressure is the test suite and adversarial reviewers, not a human reading every line.
  • Production numbers replace code review. Stability and customer satisfaction are the crash-rate analog; he says blocking merges until a person understands every change gives up the acceleration.

quotes

“You may call it vibe coding. I call it evolutionary coding.”

Less Bitter, naming tests as the pressure that decides which agent code survives.

“there's no future where you're manually reviewing every line of agent code.”

DHH, as quoted by Less Bitter, on why human line review cannot keep the acceleration.

“All you have to do is give up the code review.”

Less Bitter, stating the trade he thinks every team will take.

“The code quality does not matter.”

Less Bitter, arguing that production stability and customer satisfaction are the measurable results.

transcript

Let's say you had this spec here. This spec is 2700 lines. And imagine you gave this spec to a general purpose human being, a general purpose engineer who has some experience in your codebase, but you know, it's kind of like a new hire. How long do you think this would take him to build? I mean, first of all, how long would it take for him to read this? A couple days to wrap his head around it. Needless to say, this would take a human engineer, I don't know, three months to build this out, if even possible. I don't know how.

I mean, I don't even think a single person could handle this spec because first you'd have to read the spec, then you'd have to break it down into projects in parts. Now, the question is, could an AI pull off the implementation of this spec? And the last time I had an agent try to do a spec was back in like March of this year. And it failed catastrophically. It failed so much that I dedicated my life to tearing apart the scam of AI coding.

I'm like, dude, what is this [ __ ] Like they promise you that the agents could do all this and that you back in March you gave the agent a spec like this. It wasn't even this long or detailed. It was maybe like a 500 600 line spec and the agent would be like say less bro I got this and it would go and it would just drown in its own vomit and leave your codebase an absolute mess and it would just give up and it'd be like done and you'd be like did you finish? It's like, "Oh, I finished." All right.

You know, finished what the turn, right? It's like, I finished processing tokens. Did it finish the feature? No, not even close. Did it leave your codebase in a functional state? Not even close. And so, back then, smart engineers would face these little hurdles and be like, man, I got to design systems for these engineers so they can get through for these AI so they could get through these specs. Now, as you may have seen, Opus 5.5 in recent videos has turned me into somewhat of a believer in the coding um completion part of AI agentic engineering.

I've seen it able to do things it wasn't able to do back in March when I was when it was the last time I was working on a serious project. Now, I'm working on an app called Enjoy and I'm doing this fully agentically. You may call it vibe coding. I call it evolutionary coding. Okay, you you let the code evolve to satisfy environmental pressure, which is the tests and if it works, if it passes those pressures, then you let it proliferate. It's like biological genomes are an absolute mess, but they work. So enjoy.

This app lets you use AI coding agents that are on your computer, the terminal ones. It brings access to these terminal agents that were otherwise prior reserved only for like developers who knew how to use a terminal. Now anyone can access these. And as I was talking with a few people, a lot of them said, "Man, I have a like a non-technical co-founder that I'd love to share my project with. It would be cool if we could collaborate together." Anyway, long story short, I needed to go in with a whole new redesign of how the teams feature worked.

I was dreading it, man, because it was big. It was huge. It was like changing everything, right? It's like we built the app one way and now we had to change everything to make it multiplayer. And if this was any other era, if this was 2020, this would have been a 3, four month project where I would have just had to stop everything I was doing, go hide in a cave for three or four months, not shower for weeks, and just wallow and drown in technical misery.

And so, um, you know, I started with this huge voice memo, just thinking out loud in terms of how it should all work. You know, I just worked with it through Claude and it went ahead and it built this spec, right? It started with this spec. And I was like, "Oh, you built a spec, did you? Okay, well, what the [ __ ] do you want me to do with a spec?" Because the last time I gave you a spec, it didn't end well. And I opened the spec and it was this spec and I was like, "Oh my god, dude. Holy sh.

You're telling me, Claude, that you're going to be able to build out the whole feature according to a 2700 line spec? I was just like, "Oh, I'm so screwed. This project is screwed because I don't know how I'm going to build this. This needs to be built." And the last time I tried to do something like this, it failed. And I had to go in and take this spec and divide it into 10 little man maybe 20 little manageable pieces that an AI agent could do. And it was just no fun at all. So, I was like, you know what? Let's try. Let's try.

So, I put Claude on Ultra Code. Uh, effort level Ultra Code. If you don't know what Ultra Code is, it's like a nuclear the nuclear option of Claude. This thing will spin up 80 agents, man, to do your work. It'll spend an ungodly amount of time. And honestly, I'm still I'm on a $200 plan. So, so I just said build the whole thing end to end. I put it on Ultra Code and it said on it I'll build the full computers and devices model across the relay. So this was at 3:20 p.m. I guess it worked. Uh it ended up working about 10 hours.

And so about an hour later I was like have you done any work yet? Why do not I not see any changes or commits? So it was writing the spec, the 2600 line spec. An hour later I was like why still no work done? It says you know uh I it spawned four critics to review the spec.

uh 5:24 wrote the spec by 8:00 PM had done a couple commits by 8:15 I saw that it had 60 agents working and I'm like oh my god either this is some infinite recursive bug that spawn is spawning like these [ __ ] were hacking hugging face or some [ __ ] and it was like most of those are the reviews skeptic checks now I didn't tell Claude I didn't tell it how to do this um Ultra Code just spawned its own skeptics and adversarial ial reviewers to review the spec to review the implementation. Eight reviewers produced about 34 findings so far.

Before anything gets fixed, each finding is handed to two or three independent skeptics who try to disprove it. Three for high severity findings, two for the rest. Again, I didn't tell you to do any of this. This is just the way Ultra Code works. That adds roughly 80 short readonly agents on top of the eight viewers. 823 so far. 45 of 53 agents have finished. Okay, 9:55. It looks like it finished uh by 9:55, I guess. And the question is, so this was the whole spec. So, you know, from 3:00 p.m. to 9, let's say 10, that's 7 hours. 7 hours to build this spec.

And did it do it correctly? It did. It [ __ ] did, man. Dude, not only did it do it correctly, the end result was flawless. It's like everything just worked. And this is a very complicated flow here that involves pairing like a phone and a a web browser to the desktop client, coordinating between a relay server, end to end encrypting the traffic. So in 7 hours, Opus 5.5 built this whole spec and it did it correctly. Now you may be thinking if you are somewhat of a skeptic, okay, sure it works, but what about the code quality?

It's probably all slop or end to end encrypted. Yeah, okay. How could you ever trust that? Here's the thing about the current generation models that I see. If they tell you something, I it is pretty high confidence. They are pretty skeptical. They do not trust themselves. So, they always doublech checkck things, triple check things. And so if you ask it and you put it on a high enough effort level, does your implementation meet the spec, it will answer that accurately. This whole idea of hallucinations is is not something that you see anymore.

The models are pretty good at admitting that they're uncertain or unable to attain certainty. And so I'm sharing this not because I want to hype up AI, but I want to tell you that it they fixed AI, man. They fixed it. Like I'm amazed because the thing that diselusioned me from AI back in early of this year was its incapability of not only the inability to build out specs, but lying about it too and saying, "Yeah, I'll build out the spec." And just and just faltering completely.

And now I don't know what they did, but it's it's able to work for like 10 20 hours and not lose the thread and not devolve into incoherency. And I just didn't think that was possible. I didn't think like dude, look, I don't know what AI is. I don't know what if we're calling it intelligence, pseudo intelligence, next token predictor. I don't know what it is, but it is able to go through this entire spec and build it end to end with an insane amount of tests with an insane amount of fidelity to the spec requirements. Yeah, I don't know what to call that.

Um, if not like beyond intelligence, you know, I can't even assign a date to how long this would take a human being. Honestly, I just I don't even think this is possible for a human being to do in any reasonable amount of time. Not one human being, maybe a team of like 10 of them meeting up, breaking this thing apart, man. And dedicating their whole lives to it for like three months. That's what it would take. And Claude did it. It didn't even kill my 5 hour limit on the $200 plan. It didn't. Yeah. Look, this video is just to catch you up, man.

Is is to is to share with you the things that I'm seeing. Look, I I have been um I've I've oscillated a bit with my videos because I always loved AI in theory, but I was always disappointed in its execution. There were two videos I made back in March. One was a video about Cloudflare, uh vibe cloning, vibe forking Nex.js, and the other one was called like I was a 10x engineer, now I'm useless. In those videos, I talk about evolutionary coding. I'm like, wait a minute. Nature builds us somewhat blindly, right?

It tries things and if they pass the test, which is the sort of evolutionary pressure, it keeps that feature and it passes on. Meanwhile, DNA, the genome is junk, like it's filled with junk. It's it's or gunk, let's say. If if your criticism of AI building things is it's all slop, it's all junk, I got news for you, man. Well, you know, what do you think all this is? Now, I don't want to insult like your religious beliefs. I'm not here to say that obviously humans are marvels of the universe.

It's to say that the evolutionary method of building something is by far the best mechanism that nature has ever found for building something complex. And so there's talk about whether reading the code or not reading the code is irresponsible. DHH says it takes a while to build the confidence, but there's no future where you're manually reviewing every line of agent code. There's not much acceleration in that. You need adversarial agent reviews. You need automated testing. And maybe you spot check. That's it. From prompt from prop to production. I don't know how he says that. From prop to production.

I hate to say it, man, but I agree. Yeah. Look, you can't keep up with agent code. Everybody knows that. If you're reading a agent code, you're working at 1x. There's no gains there. There's no acceleration there. And the gains from aentic coding are too good to to give up. And so yeah, a lot of us and maybe even I was I think I I tweeted maybe three four months ago that not reviewing a dentic code is irresponsible and now I'm tweeting [ __ ] like does a CEO need to understand his employees code. Vibe coding is 2025.

Evolutionary coding is 2026 because this guy said I have a very strong feeling that these voices that say do not look at the code or code doesn't matter will fade away soon. They will not. This will only amplify. Shopify announced they're moving from React Native to native. How do you think they're doing that? It's because they don't have to read the code anymore. Otherwise, there's no speed up. You'll do a little spot check. Yes.

But otherwise, they're probably going to build the iOS app very carefully, let's say, and they'll just have an automated job that's whose job it is to sync the Android app with the iOS implementation. And they get that for free, more or less. Coinbase also announced they're moving from React Native to native. And honestly, if all we get from AI is the death of React Native, it'll all have been worth it. Man, have you tried the Twitch app on iOS? One of the worst apps I've ever used in my life. I cannot believe they ship that. It's React Native. Death to React Native.

There's no freaking way you could build, maintain, and operate products where you don't understand or can't reason about your own code. I I don't see like look at this point I think the AIs can reason well enough and report their results with a high level of confidence that you could be confident about it. We are there. We are there where if an AI tells me something now, it's high accuracy. That was not the case maybe even several months ago, but AIS are again have gotten better at reporting the true state of things. Maybe not always, but reliably enough.

Code quality isn't only about its visual appeal. It's also about understanding how the heck everything works. So that's why I said, "Does a CEO need to understand his employees code?" And so if you treat AI as an employee, what need is there to understand its code? If you still personally believe that understanding code is is important, then fine, that's fine. I'm not saying you're wrong, but I'm saying you will be wrong in a couple years. Yeah. Um it's just clear that that's where we're going. Like if you want the self-driving car, you got to let it drive, man.

And there's no there's no benefit to to to the technology if you're just going to take over. Look, you may say this says more about me than it does about AI, but I trust AI's work quality better than my own. I honestly have found that I tweeted this agent code quality may look worse to a human observer, but in my experience, agent build quality is way better. features have less bugs on launch and tend to just work, whereas human-driven builds always tended to be myopic and fail for embarrassing reasons.

I honestly trust agents more to build out features than than I would at this point trust a human being. And you may call this AI psychosis, but honestly, it's just empirical. Like, I didn't start this way. I I hope that you can appreciate the fact that I didn't start this way. I'm not some kind of AI hype guy. I follow the evidence empirically, firsthand only. I don't trust what anyone else says until I do it by myself. Yeah. Look, man. Uh, enjoy. It is I built a lot of software in my life and it is just the most bug-free product I've ever built.

Like before this, I built like a notes app. It was so simple, but the bugs were just never ending, man. Like I spent like eight years on that app, man. I just could not catch up. It was just so simple. And this app is way more complex. 10 times more complex, man. The things that it does and it just like there's there's nothing seriously wrong with it. It's because the tests are so good and the agents are so good at making CI stay green.

I think this whole idea of slop, this whole idea of of uh humans needing to read the code, these are um these are coping mechanisms, they are the future of AI is is is CI. That's the real bottleneck in all of this. It's the CI environment that has to run the the pressures against the agent code so it doesn't devolve.

If you were to offer me a choice between a human written product with human written tests and due to the fact that it was human written, the tests are are pretty small, let's say 500 tests, or you were to offer me an agent written product with 10,000 tests where the CI is always green, I would take the agent one.

I would take the agent one and yeah you would find that it would be a higher quality product higher quality code you know that whole question is is is meaningless it's irrelevant now how many of us software engineers have been burnt out from clean code like dude spent those eight years trying to write the cleanest code possible for the app that I built before this and it just it never ended man the code was just never clean enough was never stable enough like every product I've ever worked on every company I've ever worked worked for a man. Just constant hell, man. Constant hell.

And now all of a sudden, that's just not the case anymore. It's not hell anymore. It's far more enjoyable. But you have to let go a little bit and trust the agents. And I'm I just wanted to report that as as recently as Astra 6 and Opus 5.5, this this marks a change. This this this is a checkpoint in the future of AI coding. And if you're still on the side of holding on to code review and if you're advocating for that in your company, I just want to say like look, I know this is going to hurt to say, but or to hear.

I want to say be careful. Be careful being that guy. Be careful being the guy advocating for responsible code review and not letting pull requests get merged until a human has has had a very careful understanding of it. Because if you hang on to that too strictly because your idea of a gentic code is just negative, I think you might be in a bad place in your career. When you fast forward a couple years, the future is it is it it's going to be a gentic code. There's just no way. There's no way the benefits you get a thousand benefits.

All you have to do is give up the code review. Everyone's going to take that deal because what matters is the results. How stable is it in production? How satisfied are our customers? The code quality does not matter. It's about all the other numbers you can actually measure. And if those numbers are higher with aentic code and builds than they were with human builds, then that wins. The same with self-driving cars. You measure it not by the co the code quality of self-driving algorithms, but by the crash rate. How often do humans crash and how often do self-driving cars crash?

And if the self-driving is lower, well, there you have it. If you're an if you're still anti-AII in your software development company, look, you do what you want, but be very careful. I don't want you to lose your job over hanging on too much to the past. Even though on principle it may sound the right thing to do, there is serious reality happening here. There's a serious future when all this is going. And so, yep, I wanted to share that with you. Best of luck out there and thanks for watching.