hraness
Theme
Appearance

saved

Solving The AI Science Validation Bottleneck

by Mgoes (bio/acc 🤖💉)Xpublished

Hraness republishes this public post from a saved copy. The post is the author’s own words.

Mgoes (bio/acc 🤖💉) @m_goes_distance

Seeding Humanity 2.0 | Investor https://t.co/fCubX86kZn | Memes sometimes | Podcast host 👇

I've spent the past two years thinking about scientific superintelligence, and talking to founders. Dario's post yesterday, and the reactions to it, made me write it all up...

Solving The AI Science Validation Bottleneck

What matters in AI discovery is not what it seems.

Yesterday, Dario announced that Claude discovered a molecular machine that could represent a new gene editing mechanism. Its function, utility, and significance are, in his own words, not yet clear. It was nonetheless presented as evidence that AI for biology is on the same exponential trend as math.

It is also, almost to the letter, identical to the pitch I have heard from 30-or-so founders over the past two months.

Make a sleek manifesto outlining self-accelerating AI-agentic science loops and their potential to transform the world as we know it. List a bunch of highly credentialed science advisors. Promise that, if only we could remove human labor and bias from the scientific and engineering process, and throw in more intelligence and automation, surely we would unlock an exponential take-off in discoveries.

Let's stop for a minute and examine these claims critically.

Dario anticipates the obvious objection - that biology, unlike math, requires experiments - and waves it aside, because his team validated Claude's discovery in a few weeks. But what they validated is that the thing exists. That is the easy part. Not what it does, whether it works in a human body, whether it can be made at scale, or whether anyone will pay for it. He also concedes the loop will not speed up clinical trials, but argues it could greatly increase the number of promising candidates going into the pipeline.

I would argue that is exactly the problem.

To be clear, I am as bullish on the long-term potential of AI reshaping the physical world as the next tech bro. But I also believe it is paramount to be clear-eyed on how to get there, starting with identifying the right bottlenecks. Of which there are many. Though eliminating humans from the decision-making loop is nowhere near the top of the list. Nor is coming up with new hypotheses. Rather than testing their validity, and implementing them in the real world, against constraints. These are usually two sides of the same coin.

The recent Google DeepMind essay, as well as fresh research out of Imperial College London, shows that AI companies are finally catching up to the insight that ideation is not the bottleneck, but verification and execution are. The recent explosion in AI-generated hypotheses have made this obvious, and the closed-loop automated scientist companies are pouring in to fill the data-generation and verification void. The frontier labs themselves are building out their in-house science teams and wet labs to capture the upside, as evidenced by Dario's post.

Google DeepMind, "Conjecture Machines: AI agents and the new validation bottleneck in science", July 2026.

I will go one step further and argue that ideation has never been the bottleneck, and that most (though not all!) of these new AI-for-science efforts are still solving the wrong problem. And that gathering the right data to verify a given hypohesis can be a very expensive, lengthy, and messy endeavor, no matter the amount of intelligence and automation.

History is a great teacher. GLP-1s. Reusable rockets. VR headsets. Let's look at three well known examples of paradigm-shifting technologies, across three very different industries, and examine how much each could have been unlocked by scientific superintelligence with access to real world tools (the most generous version of the pitch), had it existed at their outset.

This is not an all-AI skeptic piece, and we will end on some actionable insights, I promise.

GLP-1s

Ozempic and its cousins appeared out of nowhere around 2021 and cancelled obesity, rewrote the playbook of pharma, and created the first trillion dollar drug company.

The only problem is - nothing about it came out of nowhere.

The idea that the gut releases a hormone telling the pancreas to make insulin was proposed shortly after Bayliss and Starling discovered the first hormone in 1902. The actual molecule was pinned down in the mid-1980s, when Joel Habener's lab at Mass General, with Svetlana Mojsov and Daniel Drucker, worked out that a fragment of proglucagon called GLP-1 was the bit that drove insulin release. By 1996 a Nature paper had shown the same hormone made animals eat less. Habener, Mojsov and Lotte Bjerre Knudsen finally won the Lasker Prize for it in 2024, roughly 38 years later.

"The Mechanism of Pancreatic Secretion", Journal of Physiology, 1902. The paper that discovered hormones.

So the idea that this hormone lowers blood sugar and suppresses appetite has been sitting in the literature since before most people reading this were born. The delay was not a shortage of insight.

Native GLP-1 survives about two minutes in the bloodstream before an enzyme chews it up. Every fix for that came from poking at the physical world, not from reasoning about it. In 1992 John Eng, a VA endocrinologist, found a long-lasting version of it in Gila monster venom, which became the first approved drug in the class, Byetta, in 2005. Novo Nordisk's Lotte Bjerre Knudsen made a version that sticks to albumin by hanging a fatty acid off it. She got there by making and testing analog after analog. That became Victoza in 2010, then Saxenda for obesity in 2014, then semaglutide in 2017, then Wegovy in 2021.

Thirty-five years from the molecule to the blockbuster, and most of it spent gathering resources for endless in-vitro work, escalating doses in humans one study at a time, running long-term animal toxicology through a rodent thyroid tumour scare, building peptide manufacturing at a price that worked, designing the pens, winning internal budget fights to fund the next program, recruiting patients for trial after trial, liaising with regulators, and then convincing payers to cover any of it.

The commercial reason for the delay is worse. Obesity drugs were radioactive. Fen-phen was pulled in 1997 after heart valve damage, rimonabant was withdrawn for psychiatric side effects, sibutramine went in 2010 for cardiovascular risk. No amount of intelligence tells you whether this one is different. Only a human trial does.¹

The Washington Post, September 1997.

And then the best part. The finding that turned these into the biggest franchise in pharma was not anybody's idea at all. After the rosiglitazone scare in 2007, the FDA began requiring every new diabetes drug to run a dedicated cardiovascular outcomes trial. Companies hated it, because it added thousands of patients and years of cost to every program. But those mandated trials are what turned up the cardiovascular benefit.

LEADER showed it for liraglutide in 2016. SELECT - 17,604 patients over roughly five years, showed semaglutide cut major cardiovascular events by 20% in people with obesity and no diabetes. That result, published in late 2023, is what made these drugs something insurers had to pay for.

Nobody deduced that. A regulation written to catch the last disaster forced an expensive and lengthy experiment nobody wanted to run, and the experiment found something worth more than the drug's original indication. The idea was free and a century old. What cost billions and decades was getting molecules into human beings, over and over, and convincing both resource allocators and regulators to keep going.²

Here's how much AI could have helped. The generous answer is still uncomfortable.

Give a frontier model the 1987 literature and it plausibly does the protein engineering faster. Designing a peptide analog that resists DPP-4 and binds albumin is exactly the kind of problem where the current tools are genuinely good. Retro got a 50x improvement³ on reprogramming factors this way in a couple of months against a decade of academic effort. Call it a real win and say AI compresses that stretch from eighteen years to five.

Now look at what it doesn't touch. AI can design an analog in weeks, but every one of them still has to be made, tested in animals, and dosed into people, and that part runs on biological time. Peptide manufacturing at scale means chemical plants, and Novo spent billions on them and still couldn't make enough Wegovy in 2022. Dose escalation runs at the speed of dosing humans and watching what happens to them, and you cannot parallelize your way past a body. SELECT took 17,604 people and about five years, and no model shortens that, because the thing being measured is how many of them end up having heart attacks. And the decade when nobody would fund obesity research wasn't an analytical error a smarter system would have corrected. It was a bet, made on incomplete data after fen-phen, and only a trial could settle it.

So a superintelligence handed the whole problem in 1987 generously gets you the drug in maybe 2010 instead of 2021. That's over a decade saved, and it's worth having. It is, however, not a compressed century, and everything it saved came out of the one step that happens in software.

Reusable rockets

Now SpaceX. Everyone knows the story. Elon decided rockets should land. Everyone said it was impossible. He did it anyways.

None of that explains what truly happened.

The idea of a reusable rocket is about a century old. The imperial Russian aerospace engineer Konstantin Tsiolkovsky wrote about it in 1920s. The chief Nazi (later turned US) rocket engineer Von Braun drew it in 1952. NASA built an entire vehicle around it: the Space Shuttle flew in 1981 and was sold to Congress on exactly that promise - cheap reusable access to orbit. It ended up costing more per flight than the expendable rockets it was meant to replace.

Collier's, March 22, 1952. Von Braun's winged upper stage was designed to glide back to Earth and fly again.

Between 1993 and 1996, the now-defunct McDonnell Douglas corporation and NASA flew the DC-X Delta Clipper, a rocket that went up, hovered, and landed vertically on its own legs. The exact maneuver everyone would later call impossible. The program ended in 1996 when a landing strut failed, the vehicle tipped over and burned. Rotary Rocket tried it. Kistler tried it and went bankrupt. Blue Origin was founded on the same premise in 2000, two years before SpaceX existed.

So by the time SpaceX showed up, the idea was free, the physics was known, and the landing had already been demonstrated. What was missing was neither insight nor intelligence. It was the ability to fail at it repeatedly, for years, whilst burning billions of dollars, without being terminated or going bust.

That is the thing SpaceX actually built.

No less impressive. But look at how they did it. Landing attempts were bolted onto missions customers had already paid for. The booster was going into the ocean anyway, so trying to land it instead cost SpaceX almost nothing. They missed the droneship during landing in January 2015, then missed it again in April. But it barely mattered, because the customer's satellite was already in orbit. In December one landed. In March 2017, 15 years after SpaceX founding, they flew one a second time.

Underneath all that sat vertical integration, because you can only afford to destroy hardware you build yourself at cost. Funded by NASA's COTS and CRS contracts - roughly $278 million, and then $1.6 billion. Plus Elon's ability to raise $1.5b from vision-aligned investors, which meant a decade-and-a-half of losses, without a shareholder revolt. And, of course, his ability to keep the growing team motivated, in face of extreme uncertainty and repeat failures - including a dozen or so rocket explosions.

And then, once it all worked, SpaceX still had to solve the hard problem of finding a repeatable demand engine to continue funding its ongoing R&D that would later produce Falcon Heavy and the Starship - as all of US government's missions combined did not add up to enough. Enter Starlink - another 8 years and $10b burnt, before it started generating cashflows.

The missing ingredient was never the idea. That was sitting there for almost a century, known to anyone who looked. It was coordination and execution - mustering vast resources in terms of humans, time, and capital - and successfully pointing them at a single objective, over the course of many iterations, in an environment where many other ideas relentlessly compete for the same resources.

The fact that this project would yield superior results where others failed could only be established by running the experiment. And to continue running it, for years. The success of SpaceX could not have been derived from first principles, by thinking about it deeply enough, then spitting out the result. It was an experiment that had to be run in a complex environment with way too many variables to calculate.⁴ Against the constraints of physics, the market, regulators, politicians, competitors, and interest rates.

The DC-X Delta Clipper attempting landing at White Sands, New Mexico, in the 1990s. It took off, hovered, and landed vertically on its own legs.

Would superhuman artificial intelligence have made SpaceX obvious, after all the previously failed attempts? Would it have allocated the necessary capital to Elon, instead of Bezos, who was first and had a much more established track record at that time, having created the world's largest e-commerce company?

Would AI have simply replaced Elon and all his engineers and built the whole thing autonomously? Even that would have merely reduced their overhead expenses (though not eliminated them entirely - imagine the Anthropic Claude bill!). Even if we generously assume zero overhead expense, the main costs of physical hardware, machines, fuel, factories, launch pads, would all have remained largely intact.

Indeed, Elon is already famous for having run some of the most autonomous production lines in the world - which was still not sufficient to reduce the cost or timelines beyond these figures. Neither does it 10x the speed of fundraising, obtaining permits to build launch sites and access telco bandwith, or sales cycles to NASA and the DOW.

And, of course, an AI-SpaceX would still have had to aggressively compete for resources with every other AI company's promising idea, which there never was a shortage of, in the first place.

VR Headsets

This one is the counterexample, and it's the one that should worry you most.

Head-mounted displays are not new either. Ivan Sutherland built one at Harvard in 1968, nicknamed the Sword of Damocles because it hung from the ceiling. Virtuality put VR arcades in malls in 1991. Nintendo shipped the Virtual Boy in 1995 and killed it within a year. Google Glass arrived in 2013 and became a punchline. Palmer Luckey's Oculus Kickstarter in 2012 restarted the whole thing, and Facebook bought it in 2014 for about $2 billion.

Ivan Sutherland's Sword of Damocles headset, 1968.

Then came the largest concentrated bet in consumer technology in history. Zuckerberg renamed the company after it in 2021, and Reality Labs has now lost almost $90 billion. Apple spent years and shipped Vision Pro in 2024 at $3,499, and by early 2026 had cut production and pivoted toward glasses instead.

Now compare that to SpaceX. Patient capital, founder control, a decade of losses absorbed without flinching or public failure, relentless iteration. Identical playbook. An order of magnitude more money. The opposite result.

The difference was not intelligence, or conviction, or capital. It was that the experiment came back negative. The engineering worked, the market did not. People who tried it mostly decided they did not want a bulky computer strapped to their face. And notice which version did work: cheap camera glasses with no display - Meta's Ray-Bans, a fraction of the ambition and a fraction of the budget. Nobody predicted that either. It came out of shipping things and watching what people bought.

Would a superintelligence have known? It could have read every failed headset post-mortem back to 1968 and every ergonomics study on head-mounted weight, and I'd bet it still tells you to build it, because the aggregate of past failures is the same input Zuck had, and the case is genuinely arguable. What it cannot do is tell you, in 2014, whether people in 2024 will want this. That fact did not exist yet. It had to be manufactured by spending billions on R&D over a decade of iterations, then shipping a product and observing the data.

Could AI have compressed the cycle in terms of cost and time? Surely to an extent. Not by an order of magnitude.

In conclusion

Intelligence was never the bottleneck in science and engineering.

The cro-magnon was no less intelligent or full of ideas than humans today. But he lacked the millennia of trial-and-error progress we are now standing on the shoulders of.

In each one of our examples, the bottleneck was running empirical real-world experiments and convincing others to muster the resources necessary to do so, over extended periods of time. And, as the three different trajectories show - it was not predictable in advance whether the experiment will succeed (SpaceX), fail (Meta's Oculus), or yield an entirely different thing than initially predicted (GLP-1s). Decades and billions were burnt repeatedly iterating towards a moving goalpost, in each scenario.

I think this situation is very familiar to the tens of millions of scientists globally, scrambling to secure grant money for their research - which, naturally, they consider the most important research out there, willing to stake their lives and careers on - if anybody else could just see it. Or to the comparable number of founders attempting to raise venture capital for their startup. 'Ideas are worthless, execution is everything' has been a mantra in this space for decades.

The most generous bull-case for scientific AI is to claim that this build-measure-learn loop⁵ can be accelerated. And it plausibly can, as my scenarios show. But, in each one of our examples - expensive physical world hardware, manufacturing, human biology, regulatory, and market readiness bottlenecks dominated - with no clear path towards being accelerated by an order of magnitude, 'compressing a century of discoveries into a single decade',⁶ if only the founders had unlimited access to Mythos.

AI will absolutely permeate scientific discovery and engineering. It's already an integral part of the process. But this is an automation story - which has started to unfold in the 18th century and is nowhere near finished yet - not an intelligence story. This can be still worth a lot and will compound, over time. Automation can help with validation of ideas. But most truly breakthrough ideas are too big to be validated cheaply, quickly, and effectively, even if the human-in-the-loop is fully eliminated.⁷

What's the right bottleneck to focus on then, for anyone wanting to accelerate discovery with AI?

The right premise is not "How do we eliminate humans from the process entirely?" but "Which test in my loop is the slowest or most expensive, and can I make that one fast and cheap?"

Targeting anything else than that specific test will result in marginal gains, at best. Or, in more generated candidates than resources to validate them, which has been the case in pharma for decades now, and is a well known fact across the industry. I have to repeat and restate this point to drive it home. Automating and scaling anything else than the slowest and most expensive part of the validation step will only clog the existing bottlenecks further.⁸

So - which step is that?

The answer differs by domain. In math, it is almost nothing - checking a proof is nearly free - which is exactly why AI went from high-school level to the top open problems in three years. In industrial hardware, it is the cost of prototyping, assembly, and testing - across the entire supply chain, down to materials. This is why both Tesla and SpaceX have vertically integrated. In drugs it is human trials, which may be the hardest thing on this list to move - as we are talking about observing the effects in live human bodies, across large and diverse population samples, some of which may show up only years down the line. In consumer products it is the customer discovery loop - which Meta, ironically, is the best in the world at and still got the answer wrong.

METR, "Task-Completion Time Horizons of Frontier AI Models", May 2026. AI's progress is exponential where checking the answer is cheap. Empirical sciences like biology don't give you that cheap check.

My intuition is we will end up with dozens of varied-shaped AI-enabled industrial giants, bootstrapped by being aimed at dozens of hyper-specific bottlenecks, not one single scientific superintelligence. Just like the internet has platformed thousands of companies uniquely interfacing with different parts of meatspace, instead of one superapp. Which means the winning move is the one it has always been in startups - pick an extremely narrow problem and go all the way through it.

A good example is AlphaFold: protein folding only, one verifiable readout, and decades of accumulated labeled data to learn from. This is a Nobel-prize breakthrough. But no ASI. Most of the others will require more than specialized models or data center buildout. Some will own and power the hardest physical hardware bottlenecks - just like SpaceX does with its vertically-integrated supply chain for building and launching rockets. Others will dominate getting through clinical trials.

At the end of the day, what reasoning and automation gets pointed at matters as much, if not more, than the raw intelligence you can harness.

Claude may well have found the next CRISPR. We'll know in a decade or two - once someone has paid for and run the experiments, at scale, all the way to humans.

In the meantime, the value will accrue to whomever makes the most expensive reality-test cheap.

---

My focus is on investing in life sciences, namely frontier bio that interfaces with the human body or mind. Therefore - over the next couple of months, I will be publishing a deeply-researched series of deep-dives on the past 20 years of lessons of applying computing, intelligence, and machine learning to hard problems in biotech - illustrated by some of the most famous successes and failures in the space. There is so much to learn from, for companies on the mission towards building an 'AI scientist'.

Follow me here, or subscribe to get them in your email.

May this be your guide. Ad astra simul!

---

¹ Roughly 90% of drugs that look promising in-vitro and enter human trials never reach approval (Wong, Siah & Lo, Biostatistics, 2019). The trials, not the discovery, account for most of the cost of developing a drug.

² For why drug development was never constrained by ideas, see Ruxandra Teslo, "Intelligence is not the main bottleneck"

³ OpenAI and Retro Biosciences, "Accelerating life sciences research" (August 2025): a custom model, GPT-4b micro, designed Yamanaka factor variants with about 50x higher expression of stem cell reprogramming markers in vitro.

⁴ Whether the outcomes of individual projects or companies are truly computationally irreducible is hard to prove - though the fact that even simple computer programs can be i rreducible points in this direction (see Stephen Wolfram, A New Kind of Science, 2002). For the purposes of this essay, however, this doesn't matter, as current AI systems are far from such capabilities, at the level of complexity required here.

⁵ Eric Ries, The Lean Startup (2011)

⁶ From Dario Amodei's Machines of Loving Grace (October 2024)

⁷ The cheap-and-fast tests for these huge projects are usually poor proxies (does the rocket work in a simulation, does the protein bind in-vitro, are people saying they are excited about the potential of VR)

⁸ Roughly 23,000 drug candidates are in development worldwide, while the FDA approves around 50 new drugs a year (Citeline, 2026; FDA, 2024). And more candidates don't help: improving the predictive validity of a test beats screening 10 to 100 times more of them (Scannell & Bosley, "When Quality Beats Quantity", PLOS ONE, 2016).

Dario Amodei @DarioAmodei: Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud. Show more. Anthropic @AnthropicAI: Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme's gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don't yet understand what this system does, but only a handful of ...

July 2026. Conjecture Machines: AI agents and the new validation bottleneck in science. By Don Wallace, Conor Griffin, Sean O'Neill, Thang Luong, Owen Larter. Download PDF.

[Reprinted from the Journal of Physiology, Vol. XXVIII. No. 5, 1902.] THE MECHANISM OF PANCREATIC SECRETION. By W. M. BAYLISS AND E. H. STARLING. (Seventeen Figures in Text.) (From the Physiological Laboratory of University College, London.) CONTENTS. I. Historical. II. Experimental methods. III. The effect of the injection of acid into the duodenum and jejunum.

2 DIET DRUGS ARE PULLED OFF MARKET. HEALTH CONCERNS GROW AFTER FDA LINKS PILLS TO RARE HEART PROBLEM. September 15, 1997. More than 29 years ago.

Collier's. March 22, 1952. Fifteen Cents. Man Will Conquer Space Soon. TOP SCIENTISTS TELL HOW IN 15 STARTLING PAGES.

DC-X Delta Clipper photograph: no readable text.

Ivan Sutherland Sword of Damocles photograph: no readable text.

METR chart. Title: Length of software tasks that different LLMs can complete 50% of the time. Logo: METR. Note: Measurements above 16 hrs are unreliable with our current task suite. Y-axis: Task duration (for humans), where logistic regression of our data predicts the AI has a 50% chance of succeeding. Ticks: 16 hours, Optimally reduce the size of a language model; 4 hours, Train adversarially robust image model; 1 hour, Train classifier; 6 min, Find fact on web; 36 sec, Count words in passage; 4 sec, Answer question. X-axis: LLM release date, 2020, 2021, 2022, 2023, 2024, 2025, 2026. Labeled points: GPT-2; GPT-3; GPT-3.5; GPT-4; GPT-4o; o1-preview; o1; o3; Claude Opus 4.6; Claude Mythos Preview (early). Controls: Time Horizon 1.1 (Current); Log Scale; Linear Scale; 50% Success; 80% Success.

Article cover, blue-tinted laboratory photograph: no readable text.

Screenshot of Dario Amodei's post quoting Anthropic on a Claude-found enzyme system beside CRISPR-like DNA repeats.Google DeepMind page header for Conjecture Machines, July 2026, by Don Wallace, Conor Griffin, Sean O'Neill, Thang Luong, and Owen Larter.Title page of Bayliss and Starling, The Mechanism of Pancreatic Secretion, reprinted from the Journal of Physiology, 1902.Washington Post headline from September 15, 1997: two diet drugs pulled off the market after FDA heart concerns.Collier's magazine cover, March 22, 1952, headlined Man Will Conquer Space Soon, with a winged rocket above Earth.Grainy photograph of the DC-X Delta Clipper on four engine plumes through haze above a pale horizon.Black-and-white photograph of Ivan Sutherland's 1968 Sword of Damocles head-mounted display in a lab.METR chart of software-task time horizons frontier models complete half the time, from GPT-2 through Claude Mythos Preview.