Hraness
Theme
Appearance

saved

Meet the Ex-OpenAI & DeepMind Leads Creating an AI Scientist

by Mario Gabriele, Liam Fedus and Dogus CubukThe Generalistpublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Mario Gabriele interviews Periodic Labs co-founders Liam Fedus and Dogus Cubuk about building synthesis superintelligence: AI that closes the loop between hypothesis, simulation, and real lab experiments. They argue LLMs trained only on papers cannot invent high-temperature superconductors, so Periodic is building high-throughput labs, telemetry, and training data from the scientific process itself. Fedus recounts ChatGPT’s surprise launch; Cubuk explains why synthesis, not discovery alone, is the materials bottleneck.

ideas

  • Thinkism is not enough. A model that only reasons over existing literature cannot invent new materials; Periodic insists on end-to-end loops against physical experiments.
  • Synthesis is the bottleneck. Cubuk argues making a material at lab and industrial scale is usually harder than proposing it, citing cuprates versus nickelates decades apart.
  • Train on the scientific process, not the paper. Fedus wants fully stitched experiment lineages (hypotheses, predictions, runs, and reruns) so models learn how science is done.
  • Contamination-free RL snapshots. Private lab data lets Periodic ask what a scientist should do next without answers already sitting in pretraining corpora.
  • Own the gaps the frontier leaves. They buy off-the-shelf gear when it works, and invest in hardware, simulation, and smaller models where throughput or cost breaks.

quotes

“Our goal is to build a system that can intentionally synthesize matter — we’re calling this synthesis superintelligence.”

Liam Fedus, stating Periodic’s synthesis-superintelligence goal.

“The data simply doesn’t exist otherwise.”

Liam Fedus, on why physical iteration is required for new knowledge.

“For example, we expect that if we’re successful with Periodic, it will actually increase the number of scientists in the world.”

Dogus Cubuk, on AI expanding rather than replacing scientific work.

“But honestly, outside of math, almost everything interesting is like this — the full context is never available.”

Dogus Cubuk, contrasting physics with instantly verifiable domains.

transcript

AI has gotten very good at writing software. That’s no accident. After all, code has an answer key — a computer can check it in a fraction of a second. The program either runs or it doesn’t. Physics has an answer key too. The physical world too, but consulting it means running actual experiments, something an AI model can’t do on its own.

Liam Fedus and Dogus Cubuk are building a company to change that. They are the founders of Periodic Labs, a startup building infrastructure to create what they call synthetic superintelligence. Liam previously led post-training at OpenAI and was one of the creators of ChatGPT, while Dogus was a research scientist at Google DeepMind.

They’ve brought their experience together to build what might be described as an AI scientist, capable of conducting thousands of experiments a day and learning from each one of them. In this episode, we discuss the limits of LLMs, why science is so hard, and what Liam learned from launching ChatGPT.

I’m Mario, and this is The Generalist.

I’m really excited to have you both here — I’ve been really looking forward to this conversation. I think you’re building such a fascinating company, so maybe we just start there. Tell me a little bit about Periodic. What are you trying to do?

Our goal is to build a system that can intentionally synthesize matter — we’re calling this synthesis superintelligence. Our core belief is that in order to do this, you need to make connection with the real world. A system like this isn’t going to emerge just from reading textbooks and papers.

It’s really necessary to use simulations, model the physical world, but then actually carry out the experiments. Thinking alone is not enough. It’s necessary to get new information from the world, conjecture how it might be, get the experimental results back, and do this end-to-end loop. By training end-to-end against the physical world, new types of intelligence and AI systems will emerge.

So rather than a system that’s just been trained on the body of final scientific literature, it’s trained on the scientific process itself. Rather than knowing the state of science, it knows how to do science.

I’ve heard one way you’ve maybe described the mission in the past is almost to build an AI scientist, and decomposing the different pieces of that puzzle through that lens. How do you think about, fundamentally, what a scientist does, and what you need to replicate?

The traditional way to look at science is that you come up with a hypothesis, and based on the hypothesis you run experiments, and the experiments either validate that hypothesis or invalidate it, and then you update your hypothesis. What’s really critical for us here is the observation that science always tries to push the boundary of our understanding — which by definition means it can’t just rely on what we already know, because science is trying to figure out what we don’t yet know.

That’s almost at odds with the traditional description of machine learning, where a model trains on a training set and does well on a test set drawn from the same distribution. There’s a clear tension between what we call machine learning and what we call science.

That’s where we’re basically trying to make progress. How have humans gotten around this problem? Because at the end of the day, humans are also machine learning algorithms, and they’ve often figured out things they didn’t know before — they had to iterate, run experiments, write down a theory, get yelled at by reviewers, and then update.

Einstein had to do this. Bardeen had to do this. So we feel like AI will have to do it too. We’ve been spending a lot of time and focus on building labs where an LLM and a simulation can iterate together. That’s our goal.

Liam, you mentioned this idea of “thinking” — for folks who maybe don’t know what that really means, can you tell us a little bit about it and parse the differentiation a bit?

I think it’s this concept that an AI that just thinks for a very long time — if you extrapolate the inference tokens well beyond anything we’ve seen — can just do anything. One of our deep scientific goals is to discover new high-temperature superconductors, and we don’t think a system trained the way models are trained today will think for a billion or a trillion tokens and suddenly arrive at a new high-temp superconductor.

We think that’s just not in line with how we’ve pushed knowledge and science forward historically. That’s kind of a core belief of ours: you need to connect it to the physical world and iterate. The data simply doesn’t exist otherwise. That’s how you actually build new knowledge, not just rehash existing knowledge.

High-temperature superconductors — how did you land on that as the right first step for this company?

There are a couple of reasons — let me talk through a few of them. As Liam said, our goal is synthesis superintelligence, and if we can synthesize anything in the lab, we might as well synthesize the most impactful thing, which is a superconductor. That’s one reason.

Another is: think of any of the sci-fi futures you want to imagine. They often involve superconductivity — quantum computing, fusion, lossless transmission of energy. These all have superconductors somewhere in there. And there’s a sense in which superconductivity is the most macroscale quantum-mechanical phenomenon there is. There’s something very fascinating about it.

If you think about the history of condensed matter physics, it goes through these roughly seven-year cycles of what’s popular, and then it falls out of fashion — topological insulators, heavy fermion physics. It kind of matches the length of a PhD; these topics come and go.

Superconductivity is one topic that’s never gone out of style. Since its discovery by Onnes back in 1910, people have been fascinated by it, and we’re still fascinated by it today.

And in the pursuit of this, you have to build up such sophisticated simulation infrastructure and AI infrastructure — and that has a huge amount of generality.

So it gives you the right stair-steps, and it’s also massively impactful.

Exactly. In the pursuit of this super-ambitious goal, you build up a huge amount of incredibly impactful technology.

Dogus, I’ve heard you mention Onnes before. For folks who don’t know that history — what he achieved — can you tell us a little bit about why he’s an inspiration?

Yeah, absolutely — he’s an inspiration to us in many ways. In the 1800s, there was a big rush to liquefy as many gases as possible. Around 1850 or so, people liquefied oxygen, for example, and that ended up playing a big role in industry. They were trying to go down the periodic table and liquefy as many of the gases as possible.

As you can imagine, you had to lower the temperature more and more. Helium was the hardest one, because it’s a noble gas — it doesn’t interact with other atoms very strongly, which means its liquefaction temperature is very low. And helium, especially among the noble gases, is the smallest one.

So that was the last one to be tackled, and Onnes built a lab to go after it. One of his interesting ideas was that science should be done at an industrial scale, by professionals — he hired professional technicians and engineers to work in a science lab, which people at the time found a little strange.

And he succeeded. After a lot of hard work, with a large team, he was able to liquefy helium. Interestingly, he got the Nobel Prize for liquefying helium — not for discovering superconductivity. He’d basically created a new capability: lowering the temperature to below four Kelvin. Zero Kelvin is absolute zero — you can’t go below that — and four Kelvin is really cold. Once he reached four Kelvin, he could push a bit further, to 1.8 Kelvin.

After that, he knew people were wondering what would happen to metals if they got really cold. Some really smart people, like Lord Kelvin, thought that at such low temperatures, electrons would basically freeze — which is intuitive, in a way; it’s so cold, electrons don’t want to move anymore. So Onnes wanted to try it, and one of the scientists in his lab tried putting different metals into liquefied helium.

One of the things they tried was mercury. It turns out mercury is a superconductor at such low temperatures. Once they observed that, it was a big deal — and that opened up this whole field around superconductivity.

Thank you for that history, that’s very interesting. You’ve both used this phrase “synthesis superintelligence.” I’m sure folks have an understanding of what synthesis is, but why that word? Why is that the right wrapper, the right description? I imagine there’s something quite deep there in why you’ve chosen to frame it that way.

There are historical reasons and scientific reasons. One historical reason is that the US used to be really good at synthesis, especially during the Bell Labs era. After Bell Labs stopped doing research at that scale, the US kind of lost its strength in synthesis, and it moved to some other countries.

When you say Bell Labs was great at synthesis, what’s an example of that?

Good question. Say there’s a new compound, a new material that’s interesting, and I come to you and say, “Mario, can you make this?” That can be very hard. If it’s been made before, you can just follow the previous recipe — kind of like baking. But if it’s never been done before, it’s actually hard to figure out, because at a high level you’re trying to get the atoms into the right shape.

And that’s difficult, because you’re not actually seeing the atoms — you can’t move them individually. Today, US academia has focused a lot more on functional properties, like finding a new battery or a new solar cell. But there’s this fundamental capability that’s still very important: being able to synthesize things.

Here’s another reason: in materials discovery, you might have heard of this roughly twenty-year bottleneck, where you discover a material in the lab, but it usually takes ten to twenty years before it impacts technology. One of the main reasons is that you can make it in the lab, but figuring out how to scale up manufacturing — so you can make it at industrial scale and sell it — is usually a huge issue. A lot of what you see in the lab doesn’t transfer to large scale.

And to be honest, people often think of materials discovery as being mainly about discovering the material. But if you go into the lab, usually the hardest part is making the thing you’ve discovered — it’s relatively easy to figure out what could be an interesting material. What’s usually hard is making it.

Let me give you a quick example. When Alex Müller wanted to study transition metal oxides for new superconductors, he succeeded, around 1985, and discovered cuprates, which were the first unconventional superconductors — he got a Nobel Prize for it. These were the high-temperature superconductors.

He’d actually originally been trying nickelates instead of cuprates — if you look at the periodic table, nickel is right next to copper. So instead of copper oxide, he tried nickel oxide. It didn’t work, so he gave up and moved to cuprates. But then, forty years later, our neighbors at Stanford — Harold Hwang — discovered that nickelates are also superconductors, once he figured out how to make them.

I guess what we’re trying to say is: even when discovering a material is the main goal, being able to synthesize it efficiently is the biggest bottleneck. So we’re spending a lot of our time figuring out how to make the things we want to make.

For someone who last conducted a scientific experiment probably in tenth grade, and nothing serious — when you say it’s hard to make things, what are the ways that’s difficult, precisely? Is it a question of precision? Consistency? Pick any example, I suppose — what are the most common ways that process fails?

Even if you know how you want to configure the atoms — okay, if they were arranged this particular way, it would be stable, it would have the properties we care about — there’s a whole set of processes you need to go through to get there. What elemental precursors do you use, in what proportions, what temperature profile, what pressures? What are all the processing conditions?

There’s a huge amount of challenge in all of that, and it’s necessary to actually produce a new material. It’s insufficient to just bring the precursors together at ambient pressure and temperature — you have to push into new regimes. That’s extremely difficult. But once you have a system that can do this, it can intentionally and efficiently start designing matter, which we think is extremely foundational.

Another naive question — once you have the recipe, so to speak, for one of these things, how easily can you make it consistently? It sounds like once you have the idea, it’s hard to get it right the first time — but once you’ve dialed that in, is it quite replicable? Or is it actually surprisingly hard?

It really depends. Let me give you some examples of why it’s hard to replicate. Take furnaces. If you think of a very simple synthesis — solid-state synthesis — you mix precursor powders, put them in an oven or furnace for some time at some temperature, and take them out. That’s the simplest thing you can do, and even there, there are a lot of issues.

One issue is that precursors might have different amounts of impurities. You buy a precursor that’s, say, 99.1% pure, but there are other impurities, and those impurities can affect the results. Another issue is that furnaces change their temperature stability over time. One thing we realized in our own lab — a real pain point — was that a furnace was working well for the first three runs, but after that its temperature stability went haywire. Things like that can make replicability difficult.

But humans have definitely gotten better at this with experience. Another example, from cuprates: when they were first discovered to be superconductors, some people wondered if it was even real, because people had trouble replicating it in the lab — it was a difficult synthesis procedure. Today, people have dialed it in so much that it’s in textbooks — if you open a solid-state chemistry textbook, some of the synthesis examples will include cuprates like YBCO. So it’s one of those things that seems unreplicable and hard to figure out at first, but as we gain experience, we get better at it.

And one key piece of what we do at Periodic is that we think about the hardware engineering side of all of this too. Something that used to be a hidden variable — like an oven degrading slightly over time, or asymmetries within the furnace — if you can get better telemetry on these things, you can pipe that information back to the system, and it can build a better sense of it.

One of the things I think about when we talk about the state of modern science and modern knowledge is the replication crisis — it’s hard to reproduce results we think we know, so there are things we don’t necessarily know are true, or don’t know with full confidence. It sounds like because of that telemetry, and the way you can engineer this system with a different level of fidelity, you can hopefully start to figure out why these things fail, and get better replication and reproducibility.

Yeah, and we also run replicates. We have to understand the variance of different systems, and really teach the system to make better decisions under uncertainty. For example, if there’s some aberrant noise in one pattern, we don’t want it to over-index on that and say, “No, this should just be rerun — this looks like noise.” Building that kind of scientific intuition into these systems is what we want to do.

I want to go so much deeper into what you’ve built, but maybe let’s take a step back in time first. I think you’ve both said before that you met at the gym, flipping a tire. How did you both get to that point? Maybe we could start with you — you were telling me before we started that you grew up in western Turkey, I believe.

Yeah, that’s right. Growing up, I always knew I wanted to be a physicist, but what I didn’t realize at the time — though it makes more sense looking back — is that I always enjoyed simulating physics and math. If I say “QBasic,” do you know what that means?

I will not.

Okay, some of the listeners will. QBasic was this programming language available when I was a kid, and as soon as I got my hands on it and figured out how to program, I was doing simulations of applied math and physics — though at the time it didn’t quite click for me. Back then I was really just thinking about theoretical physics; my heroes were people like Einstein and Feynman. In high school I learned more languages, like Java, and again it was all simulation-related.

When I started my PhD, I realized something: physicists were using tools from the sixties and seventies to analyze data, but the amount of data had increased exponentially, because of Moore’s Law — people in the nineties had exponentially more compute than people in the seventies. So physicists were producing enormous amounts of data but analyzing it by hand, or with simple analytic models. This was around 2010, and I felt like we should see how machine learning could do physics.

There was a lot of pushback at the time, because machine learning is probabilistic, it’s biased, it’s not always calibrated — and physicists, especially the Platonist ones, tend to think there’s one reality, one beautiful theory. So there was pushback. But in 2012, when the ImageNet moment happened and deep learning started gaining prominence, by the time I finished my PhD, machine learning was everywhere. Physics departments actually had to put machine learning somewhere in a grant proposal — especially in condensed matter physics — just to get funding. It was a huge change in about five years.

Around that time, Google Brain had this residency program — I think it was partly Ilya’s idea, which was brilliant — to bring in people with some machine learning background or interest, but expertise in other fields, to learn deep learning and contribute to it at Google Brain. That’s how I joined, as part of the residency program. Liam joined in a similar way — he can explain that himself. So we were both doing deep learning research, and I was at the gym trying to keep up with a workout. It was difficult, so I picked him as the strongest person I could find.

Yeah — you saw the arms, right? “This guy’s going to help me.”

And it succeeded.

You know, we managed.

So I kept doing regular deep learning research — this was an exciting time, I’d just published at NeurIPS, ICLR, ICML — deep learning was advancing through these conference publications, and there was a lot of fun in that. I was also still pursuing my earlier passion, simulating solid-state physics.

At some point, Liam and our other friend Barret Zoph left to start ChatGPT, and I ended up staying and really focusing on physics. By the time I left, I was running a large team of chemists and materials scientists at Google DeepMind, and we were doing a lot of simulation and machine learning work. Liam and I stayed connected, often talking about physics and other things, and we felt like the technology had reached a place where we could really impact the market and the industry. So we left our respective roles to start the company.

Amazing. Okay — what’s your origin story?

I grew up in Maine, and I also had an early love for physics. In college I was torn between physics, chemistry, and economics, and I was part of a dark matter detection program. A key piece of that was reconstructing particle tracks, which became a physics model — really a physics-informed machine learning model. So I started doing machine learning in physics at a very early stage, in undergrad, and then again in graduate school, reconstructing particles, this time at CERN.

Through that work, we kept pushing the amount of computation we put into it, and the algorithms for machine learning. But after a while it became clear that if I really wanted to push machine learning forward, I should be doing it in industry. So I went to Google, on the Google Accelerated Science team, and later Google Brain — that’s where Dogus and I met. I worked on a mix of things in the science program, then switched to generative models and reinforcement learning, and was really in a scaling mentality. It wasn’t necessarily applied to physics, but I saw how impactful language models were becoming, and worked on some of the architecture needed to scale them up — sparsity, Switch Transformers, mixture-of-experts. That became an important architecture for serving today’s models. We continued scaling that up, did some of the first trillion-parameter models.

At some point it became clear there was a huge possible impact for this in the broader world — in a product. That’s what led me and a few others to OpenAI. We felt it was a great time to start building products with some of these things. When we arrived, GPT-3.5 was out as an API, GPT-4 was internally available, and we had some time to develop it. At some point we decided to do a quick sprint using GPT-3.5 — a two-week sprint — and that turned into ChatGPT.

And did that product do well?

It exceeded some of our expectations. We had internal polls about how many people would use it — people guessed numbers like 10,000, 50,000. A high bar was a million. No one guessed it would be what it became. It was genuinely a low-key research preview.

Your mental model must have then totally shifted, right? I mean, that happened to literally all of us, but you were —

Well, the context before that too is that Meta had just released Galactica and had to withdraw it within a few days of release.

I didn’t remember that.

That was only a couple of weeks before, and there had been no successful chatbot up to that point. So the prior that this was going to be valuable technology was quite low. It was further surprising because we had GPT-4 internally, and it wasn’t this hugely viral product within the company. So putting out a weaker model to the public was really just part of iterative deployment. And obviously it worked better than our predictions.

What was the period after release like for you? Was it just totally frenetic — chaos?

Yeah, at times. The servers couldn’t keep up, and you’re just trying to rapidly improve it. The early model couldn’t do math, couldn’t write very well, had many deep flaws — very different from what we have today. But it was just a series of getting feedback, improving the task distribution, improving data quality — all these kind of boring details — that led to the system we have today.

Very few people will ever have an experience like that, I suspect, but I imagine it serves you well as a founder today. What do you think it taught you? What did it give you?

Question everything. It’s really interesting — people would come in years after the initial creation, and it’s very easy to anchor on some random decision made a year or a year and a half earlier, not necessarily an optimal decision, but you take it as one of the assumptions you build on. But I can recall, no — that was just one conversation, one decision someone made — feel free to change these things.

As we were building that up, and eventually I was running the full post-training team, which sits between research and product, we made incredible progress in the digital world. But Dogus and I were talking, and we felt the impact of these systems on the physical world was always going to be constrained unless we actually connected them to the real world. Reading text, reading textbooks and papers, deriving questions from that — that’s great foundational technology, but insufficient. Ultimately, the main objective for artificial intelligence that we’re really excited about is new scientific discoveries, new technology — something tangible you can hold in your hand. That was the motivation leading up to Periodic.

You’ve both worked with some of the very best builders and researchers in this space — why did you pick each other? What was it about Dogus that made you think, this is the guy I want to spend the next chapter of my career with?

Just an absolutely brilliant thinker — there’s no one better in the world to do this problem with. Dogus has been at this intersection for fifteen years, more.

Yeah, fifteen years.

Right. And we’ve known each other for a decade, so it was just obvious there’s no one better in the world to do it with.

What about you, Dogus — why Liam?

Many reasons. One is that with a co-founder, the most important thing is trust — you’re trusting your own career, but also the people you bring into the company. There are people at our company I’ve worked with for many years, and Liam has too, and I wouldn’t do this with someone I couldn’t trust fully.

Another reason is that Liam has been extremely successful in ML, but has always been very passionate about physics. Whenever we caught up, he’d get excited talking about quantum physics, quantum computing. I think it’s really important for the person to genuinely care about these things, and I could tell Liam really does — even with so many exciting things he could be doing on the ML side, this is what he wanted to do. And of course, we needed deep expertise in RL and LLMs, and Liam has done RL research and contributed to the frontier for so long. So it was a very obvious, easy choice.

You mentioned bringing people into the company — I think one of the things you can see from the outside, but that also came up talking to some of your investors in preparation, is the talent density you’ve created, which is extremely impressive, especially for a company this young, still in its very early innings. How have you convinced some of the very best people in this industry — who are being fought over ferociously — to pick this as their mission?

I think it’s mission, quite simply. People see the ambition of what we’re trying to achieve, and the possible impact — something that can actually design the physical world, something that understands physics and chemistry, which is even more foundational than a system that understands language. I think that really appeals to people.

I think it also appeals to people in terms of leverage — how much of a difference can you make in this position? You might be a very successful contributor to a coding system, but there’s always a counterfactual: if I’m not the five-hundredth contributor to this coding agent, what’s the delta? People look at that and think, my skills could be better used elsewhere, where I can have a larger impact on the world. I think that appeals to a lot of people — as well as scientists seeing how rapidly AI is advancing in the digital realm, in math, theoretical physics, and so on, and wondering what that could mean for the physical sciences. I think that’s a huge draw.

I imagine it doesn’t hurt that when you look at the best researchers in AI, so many of them have physics backgrounds — it seems to have drawn a lot of that talent. I’d imagine for a lot of these folks it’s a chance to bridge their childhood love with where they’ve built their careers. I don’t know if that’s how you’ve seen it.

Definitely. I think there are two aspects. As you said, a lot of these researchers deeply care about science — that might be a personal-interest thing. But on the other hand, they really want AI to impact society in a positive way, and I think science is such a good direction for AI because there’s no end to it. For example, we expect that if we’re successful with Periodic, it will actually increase the number of scientists in the world. Whereas for certain roles, if AI replaces the employee, it’s kind of a zero-sum game — for science, it’s the opposite. The more we can show that science can impact the world — especially solid-state physics, chemistry — the more people will go into that field. There’s no end to science, no end to engineering. We can discover more, we can visit more planets.

Let’s get really granular, to the extent you’re able to share it — what does one experimental loop look like at Periodic? What are the different machines, the different processes involved?

Say we’re doing powder synthesis, trying to make a superconductor. The first decision is what precursors to mix, because the thing you’re trying to make probably doesn’t exist on your shelf yet — you have to get the ingredients. Those can usually be binary powders — meaning two kinds of atoms — or elemental precursors, a single element. Basically, you figure out what precursors to combine for the thing you’re trying to make.

Is that you figuring it out, or the AI figuring it out?

Great question. Traditionally it would be a human figuring it out. For us, the AI can really speed up this process — especially down the line, if you want to run thousands of experiments a day, you can’t have a human do it one by one.

So when it’s speeding it up, is it maybe saying, “Here are thirty good options for you” — narrowing the scope of what you might look at, and introducing new ideas?

That can happen, or it can actually say, “For this goal, here’s what you should mix” — and we try that. So you mix the precursors, and now you need to react them, because at room temperature these materials won’t react with each other — they’ll just stay as precursors. There are different options, but a very common one is what we talked about: you put it in a furnace with a temperature profile. Liam has often pointed out that the temperature profile looks like a learning-rate decay profile, like in machine learning.

Sorry, I’m just so interested in each of these steps — putting it into the furnace, is that a human doing it? A robot? Is that already automated?

That’s already been automated by others before we started Periodic — there have already been attempts at getting a robot arm to take the mixed precursor and put it in a furnace. We can do either: have a robot do it, or have a human do it. But as you want to scale up to thousands a day, then you definitely need robots to do it.

Yes, okay — so it’s in the furnace. Then what comes?

The very traditional path is you wait a predetermined amount of time — predetermined by us or by the AI — and then you take it out, hoping that during that time and temperature profile, the reaction you hoped would happen did happen, and the reactions you didn’t want to happen, didn’t.

So now you have it — you cool it down a bit so you don’t burn your hand — and at some point you put it into a characterization instrument. For inorganic crystals, the method we use most is XRD, X-ray diffraction. It’s the same kind of X-ray you’d get at the dentist — it just happens that the X-ray wavelength is comparable to the distance between atoms. So if you pass X-rays through a solid, you can see what the atoms are doing, how they’re positioned relative to each other, and that tells you whether your experiment succeeded.

We find that AI can play a big role here too, because once we get the X-ray diffraction pattern, it’s not always easy for a human to tell what you made, but the AI can help with that and tell you whether it succeeded or not, and send it back for a new experiment if needed. And if it succeeded, maybe you want to check if it’s a superconductor — so you put it into another instrument that measures the diamagnetic response, which tells you if it’s a superconductor at a given temperature.

Yeah, I think the XRD example is really interesting too, because the materials don’t come out labeled.

Yes.

Right — “well, here’s what you’ve made.”

Exactly.

That’s right. So AI is allowing us to basically do the labeling, and figure out what we’re making, at a much faster pace. Anyone could go into a lab, mix a bunch of powders, and say they ran a lot of experiments — but if you’re not seeing what you’re doing, there’s no error-correction loop, no intentionality to it. So those are some of the first AI systems we had to build, and they let us scale up the throughput of the lab significantly.

Are you gathering lots of ancillary data along the way? Is that useful? I don’t even really know what that would be, but I imagine a mechanized version of this process permits that in a way a purely human version doesn’t.

Better telemetry is one aspect, but whether it’s robots or humans, I think stitching the data together is one of the most important pieces. This is what I was alluding to earlier — most AI has been trained on the final artifacts of science: the final paper, the final recounting in a textbook. That isn’t how the scientific process actually unfolded; it’s a retelling of the story.

What we do is be very careful in collecting the intentions — what were the hypotheses, what were the computational predictions, what was the execution through all the machinery — and this fully stitched-together lineage is very unique. That’s something we’ve spent a lot of time putting together. In training, the system is learning not just what the final outcome was, but what the process was, what judgments were made, what reruns were done as new evidence came in — basically, what the process of doing science looked like.

In a way — maybe this is a bad analogy, you’d know better than I would — it feels like a chain of thought for science, where you’re starting to see —

Yeah, yeah. — interacting with the —

real world. Yes, exactly. As you’ve now started running some of these experiments, what do you notice that having that trace actually gives you? Is it better suggestions for the next experiment than you’d have gotten otherwise? Is the next whole set of experiments somehow leveled up compared to starting cold?

One thing is just being able to pay attention to detail when there’s so much detail. We’ve seen the LLM detect major mistakes and debug them. One thing that happened in our lab: it turned out that samples are placed in a circular sample holder, and we made a mistake — they were shifted by two positions, so all the samples were mislabeled. We had no idea what happened, but the LLM was able to go through it and realize, “You just shifted counterclockwise by two — if you do the inverse rotation, now they’re all labeled correctly.” A human could do that, but it takes a huge amount of attention to detail and time. Now imagine doing that across thousands of experiments at once.

The other thing LLMs are really good at, which I’m excited by, is that they can operate at the level of different expertise and different PhDs at once. For example, our LLM can do simulations at the level of a PhD in simulation, but can also do experimental data analysis at the level of a PhD in experiment. That’s unique, because most people get a PhD in one or the other, and things fall through the cracks in between. So having a platform that’s an expert across all these different modalities, and can reason across them at the same time — that’s really valuable.

One thing that happened for us: a result came up that wasn’t what we were trying to make, and wasn’t a material in any database, so we couldn’t tell what we’d made. But the LLM ran new simulations and figured out what it must have been, and we were able to validate it with the next experiment. I often say the people who should run simulations and the people who know how to run simulations are usually disjoint sets of people.

Yeah.

I think LLMs will bring these together. And it’s fractal — I’m talking about one interface, between simulation experts and non-simulation experts, but it’s also true between, say, inorganic chemists and organic chemists, or chemists and physicists. Human expertise is kind of like a fractal, and there are interfaces throughout it. Being able to speak the language of all these different groups is going to be a qualitative change.

Another interesting thing to add: as we build up this database of everything that’s been run, all the decisions made, you can start constructing new kinds of machine learning tasks — you take a snapshot of the state of the world and ask, given the computational and experimental state at that point, what did the scientist do next, and what was the outcome of that? That’s a very rich source of training data for reinforcement learning.

One analogy people have talked about is: let’s train an LLM up to general relativity, and see if the system can produce something like that on its own. We’re doing an analog of that at our scale, in our own experimental and computational data. This is a kind of data that really confers a huge advantage, because it just doesn’t exist anywhere else.

You mentioned this idea of a snapshot in time, which is such an interesting idea to go a level deeper on — what might a snapshot at Periodic look like, and what would you get out of looking at it?

Say, hypothetically, you had 100,000 experimental data points. You could say: what was the state of experimental evidence as of some date? These were our computational predictions, these were the experiments we’d run up to that point, here are some notes on what we were thinking. From that point, that becomes your environment state, and you ask: given this basket of data, what did the scientist do next, and given what they did, what was the outcome — what were the new empirical measurements?

That becomes an incredibly interesting source of reinforcement learning tasks and data that’s incredibly hard to construct outside of Periodic. And it’s easier for us to do this kind of work because, as we accrue more evidence and data that doesn’t exist in the public literature, we know our model isn’t contaminated. There’s a risk when your pretrained model already knows the answer, and you do reinforcement learning against an answer that was already in the pretraining corpus — the system can basically fake its way to the right answer, and you reinforce that strategy, which isn’t going to generalize or be particularly interesting. We’re able to circumvent that by having a basket of data we know doesn’t exist in the literature.

When I think about what you’re doing, there are obviously so many different pieces to pull together. What doesn’t make sense for you to do? Do you need to develop every piece of hardware yourselves, or can you use someone else’s robot? Will you ultimately create your own models, or can you fine-tune existing ones for these different pieces of the process? What do you feel like you need to uniquely own?

For hardware, if there’s an off-the-shelf component we can just buy and use, we’re very happy to do that — we don’t need to reinvent the wheel. It turns out that for certain scientific instruments, the throughput or the quality isn’t there, so for those we invest in hardware engineering to improve throughput and so on.

What’s an example of something that just doesn’t hit the scale you need, that you have to reinvent?

One that already matches our needs is XRD — there’s already so much work in high-throughput XRD that we have instruments that can run forty samples at once. But for things that are more materials-property-oriented, where measurement is slower and might take an hour per sample, that could be a big blocker for us — so we invest in new engineering to make it work. Same for simulations: for some things there’s already a really good codebase, so we don’t need to improve it, but for others it’s not there yet, or not scaled enough, and we invest in software engineering to make it happen. Same on the LLM side.

Right now we spend very little time trying to improve software engineering itself — we use these systems to build up our codebases and help with these analyses, and we’re absolute beneficiaries of how good Codex and other tools have become. What we’re targeting are areas where the frontier just isn’t sufficient, or where we can do it vastly more cheaply — that’s where we focus our efforts.

We deeply believe, and I think we’ve seen good evidence, that when you actually train against this data, you can make reasoning more efficient — you add new kinds of capabilities and go beyond what inference-time thinking alone can do. Compressing this knowledge, compressing these strategies into the weights, is fundamentally valuable, and that’s why all the frontier labs keep doing this — we didn’t stop at GPT-4 and just do inference-time reasoning from there. Compression really needs to happen, and you get different kinds of patterns that emerge.

Take tool use, for example: there’s declarative use of a tool — here’s how you use it, here are the specs, the inputs, the expected outputs — versus procedural, intuitive knowledge, like knowing when you should use this tool in this situation. A friend of mine used the example of reading a manual on how to play saxophone versus actually learning to play it. You want to build that intuition into the weights.

We try not to duplicate things where the frontier is already sufficiently capable or cheap enough. We’ve had a few instances where, because we’re pushing the lab’s throughput so much, some analyses would have been prohibitively expensive at one point. One of our engineers told us that at our max lab throughput, a certain analysis would cost roughly thirty million dollars a year. So we looked at how to compress that into smaller, more efficient models — taking it from thirty million dollars down to a sub-million-dollar cost. That lets us reinvest that capital elsewhere, into other machines, training, or simulation. That’s a type of work we’ve been doing.

Speaking of capital — you just raised a new round. What does that let you do now? Is it predominantly compute? Hiring? All of the above?

Compute and infrastructure — those are the two key pieces for Periodic. Compute has become so central to pushing the frontier of synthesis superintelligence, and also our simulation workstreams, so we spend a huge amount of compute on both. But key to our hypothesis and belief is that we have to continue building out more labs and acquiring more infrastructure. That’s how we’re creating systems that wouldn’t otherwise exist.

How have you gone about securing the compute you need, when there’s clearly a massive supply crunch and everyone’s trying to get their hands on it? What does that look like for you?

We go through our own sets of relationships with providers we’ve already partnered with, and we understand the constraints for others — but we are able to find compute. It’s certainly a different market than it was a year ago; the frontier labs are monetizing compute at a very high rate, so they have this flywheel of soaking up all the additional capacity. But we still find there’s sufficient compute for what we need to push forward.

Another key piece for Periodic is that by having the physical infrastructure — the labs, the data no one else has — that confers a compute efficiency advantage. That’s really key, because you don’t have to match, one-for-one, the compute of a much larger company, since you have sources of data no one else has, and those efficiencies make it a lot more tractable.

Over the past year, what’s the rough order of magnitude of experiments you’ve run, and what do you hope that looks like a year from now? Are you trying to raise it by an order of magnitude? Doubling? Tripling? What does that look like?

We started with zero experiments, obviously, and we’ve been ramping up quickly — we really wanted to get experimental results as early as possible, so we built the lab and started getting results fast. Our goal on roughly a one-year timeline was about a thousand experiments a day, and that’s what we’re targeting.

One thing worth pointing out: just running experiments without getting anything out of them is less difficult, but also less desirable. So we’re scaling up the volume, but also making sure it comes with high-quality telemetry. Our hypothesis is that LLMs won’t just arbitrarily generalize — if you believe that, you have to build labs for physics, but you also have to build many labs, because one lab you build generalizes to that kind of lab, and maybe similar labs, but there are many other labs that won’t generalize from it. So we’re very invested in building many high-quality, high-throughput, modern labs, which is very exciting, honestly.

So are you at a thousand experiments a day now, or is that the goal a year from now?

We’ll get there soon.

Okay — well, that’s wild.

And we’re also going to go beyond that, adding more modalities, different kinds of experimental directions.

We realized we could conduct very high-throughput experiments that are also fairly low-latency. In a lot of areas of science, there’s lower signal-to-noise, and things can take days, weeks, sometimes months. That’s been a huge motivation for some of our starting points.

Wow.

That lets us build up this huge quantity of data while also maintaining quality and diversity, and that becomes an incredible asset.

The diversity piece is something I’m interested in — maybe because I don’t have enough understanding of superconductors — but how do you make sure you’re pursuing different modalities, getting data that’s sufficiently diverse and not overly specialized?

There’s definitely a trade-off to consider. If you make it as diverse as possible, these experiments don’t learn from each other — you haven’t really gone deep enough in any one modality. If you make it really narrow, there’s a big problem: if that particular modality doesn’t lead to a new superconductor, you’re stuck, no matter how well you execute.

So we’re trying to walk that trade-off correctly, and one of our guiding principles is actually pretty technical: we want to make sure the different modalities we consider benefit each other, because there are scientific experiments — different synthesis methods, say — where getting good at one also helps you get good at the other, and vice versa. So we build our LLM training, our simulations, and our hardware in a way that lets them benefit from each other. That’s been really effective — it’s how we build experimental diversity while still synergistically benefiting from each experiment.

To use ML parlance, we want positive generalization.

Does that come from your own intuition and knowledge, or is it also AI-assisted? Where does the human come in there?

Originally it was mostly human intuition, but we’ve already seen indications of the LLM contributing to this. Sometimes it will say, “You’ve done this experiment, this synthesis method, but if you were to add this, it’s a new modality that could help” — not a massive hop in the synthesis-methodology space, but a small, useful difference. It usually has a pretty good understanding of the literature — what kinds of synthesis attempts have been made for this chemistry — and it can guide us: “Maybe you should try this next.”

What’s the right way to monetize this kind of product? Does it make sense to develop these discoveries yourselves and sell them? Does more contract work make sense? What’s the right model to start with, and what does the right model look like at maturity?

Maybe an interesting analogy is what we’re seeing in software engineering. When we first shipped ChatGPT and these early models, we didn’t promise a full-fledged software product — we weren’t saying, “We’re going to produce a Kubernetes replacement for you and monetize that directly.” It was a copilot to accelerate a software engineer and get them to solutions more quickly. Fast forward a few years, and we’re starting to talk about monetizing outcomes — a newer concept in AI.

I think something similar is going to unfold here. The tools we’ve been building up in our own discovery loops are hugely valuable to a lot of different players in the industry, and we’re able to build custom systems for partners and customers to help them get to their solutions much more quickly. But in parallel, we’re also talking about outcome-based pricing, where we retain the IP and the upside of the final output of our systems.

There’s another interesting piece here too: if the software has, in its weights, the recipe for a room-temperature superconductor, you don’t want to monetize that through a software deal — you want to retain that IP. So those are the kinds of trade-offs we have to balance.

Are there analogies from other industries you find yourselves referring to? I imagine some part of the current scientific apparatus is useful to learn from, but I could also imagine different forms of manufacturing having taken this much further — whether that’s building a computer, an iPhone, a car, whatever it might be.

One thing I think is very relevant to us is drug discovery. There was a time when discovering a drug didn’t let you make much money off the IP, but that changed — today you can discover a drug, go through some clinical trials, and get paid billions of dollars for the IP alone. The field changed, and the capabilities that support that market changed too. I think materials is at a similar point now: you can make a lot of money selling materials, but materials IP doesn’t really capture the value it deserves yet. It could go through a similar transition to drug discovery.

We feel that as predictions get better, IP will have more value — that’s why we’re trying to close the loop all the way from discovery to synthesis to scale-up of synthesis. That’s one of the analogies, I think.

Before we hit record, you were mentioning the difference in performance between mathematics and science — maybe you can tell us a bit about that.

This is something we’re very passionate about, and it’s one of the reasons we started the company. If you talk to LLMs, they’re incredibly intelligent — you’re blown away. On certain things, their performance is insane: coding, math, theoretical computer science. But you haven’t seen that kind of impact on the physical world. Why is that?

We think one reason is that with math, theoretical computer science, and coding, the reward is instantly verifiable — you write code, you can run the unit test; you have a mathematical theorem, you can validate it. Another reason is that all the context is available to the LLM — if you’re writing code, the model can see all the code; if you’re doing math, it can see all the axioms and corollaries. Physics will never be like that. When I run an experiment, there are more atoms involved than Avogadro’s number — no computer can store the positions of all of them. So we have to accept that in real life, LLMs will never have full context.

But honestly, outside of math, almost everything interesting is like this — the full context is never available. Think about politics, finance, literature — what makes life interesting is that you don’t have all the context, but you still have to make decisions under uncertainty. Physics is very much like that. So if LLMs are going to play a big role in the real world, in the physical space, they have to figure out how to deal with missing context, rewards that aren’t instantly verifiable, and noise. We feel there’s a very exciting future to be part of there, and that’s why we’re working on this.

I love that — that’s so interesting.

And the strategies that are optimal for math, or for software development, may not be the ones optimal for decision-making under uncertainty. That’s the vision we’re pushing forward here.

To what extent is Periodic’s fate tethered to that of the foundation models? Obviously this isn’t going to be the case, but if we didn’t get another model from this point forward, how far do you think you could push things?

From the software engineering side, I think we could literally stop today and still find every remaining Nobel Prize, at least in the physical sciences — a lot of physical science relies on fairly simple analysis, it’s not competitive programming. So I think from that perspective we’re in very good shape, and we’re huge beneficiaries of that. I hope they continue to improve rapidly, and I expect they will, but I don’t view that as a blocker for us anymore.

Absolutely. There’s so much to do on experimental analysis, simulation, scale-up, and data analysis, and the current models — and our own models — are already good enough. Let’s say AI is very popular today, but suppose tomorrow it isn’t, for some reason — there are certain things that have happened in physics that will never go back. One of them is force fields. Have you heard that term before?

No.

The atoms we deal with are governed by fairly simple quantum mechanics — you can run a quantum mechanical calculation to get the forces on each atom, or you can use an empirical model that approximates that force, which is called a force field. It’s a very old topic; we use it all the time to model materials and molecules. Force fields have changed so much because of graph neural networks that I don’t think it will ever go back — now we can build very general models across the whole periodic table. In the past you could only really do three to five elements at a time. It’s a tough one to walk back. I feel like machine learning has already made such an impact that it won’t reverse — and as Liam said, today’s capability is already enough to meaningfully advance science and engineering. Of course, we expect it to keep getting better, and we’ll keep benefiting from that.

Clearly we all agree these models are going to keep getting better. Planning with that in mind, what changes for you? What do you expect to get unlocked that isn’t unlocked yet?

Some things we’re really excited about: much more advanced software engineering. We do a lot of very advanced simulations, and we’re starting to build more advanced simulations where we control every bit of the code — things that would have taken a team a year or two, and now it’s like an engineer’s weekend project. So we’re extremely excited about that, especially in areas where the AI system can write software and simulations that are maximally consistent with experimental evidence.

It still feels like what you’re doing is very differentiated in this market — right now this is at the “max mimicry” stage in a lot of AI, where someone has one good idea and five other companies follow, which is probably healthy in some ways. How long before you think that starts happening to you, and who enters the fray? Do the big foundation model companies start saying, “Actually, this is invaluable data — we need to be doing this ourselves”? Where does the competition come from, in your view?

I’d suspect that in the next couple of years, a system that can literally engineer matter is going to be of key importance to everybody. I think that’s the motivation for moving as quickly as possible — building out a huge network of labs and infrastructure, building simulations and AI systems that can control and orchestrate all of that. That’s our belief. I’m curious too.

I do think we’re pretty unique in the team we’ve gathered and the approach we’re taking, but taking a step back, there are actually a lot of people working on this now — almost every week I hear about a new company in the AI-for-science space, which is great. If you really believe science is endless, you want more and more smart, dedicated people working on this. Looking back, I think it makes sense — people are seeing how much LLMs are changing the digital world, and wondering how much they could change the physical world. So it’s not a surprise to us that a lot of people are getting into this field.

You two are co-CEOs, right?

Yeah.

How do you divide and conquer — what are the swim lanes you naturally take?

Very naturally, based on our expertise — I spend a ton of time on the AI systems, that’s been a core piece of my day-to-day thinking, as well as thinking about how we monetize these AI systems in the real world.

Yeah, there’s a natural division there. I always like to wrap up with some more abstract questions — curious what both of you would say. Given what you do, if you had unlimited resources and no operational constraints, what’s an experiment you’d really like to run?

I would love to do reaction calorimetry — there’s an experiment that tells you the formation enthalpy of a crystal. What I mean is: how much energy would it take to take this crystal apart, remove the bonds and connections? What’s nice about that is it’s a single number you can get — the formation enthalpy at some temperature — that you can directly compare against simulation. It’s usually hard to find an exact comparison between simulation and experiment; it’s usually convoluted, narrative-based. But here it’s direct — it’s just a very complicated, expensive experiment.

If we had infinite money, I’d love to do thousands of them. There’s some data from NIST, the Standards Institute in the US, who’ve done these experiments for hundreds of materials, and it’s been invaluable for the solid-state chemistry field. If we had infinite resources, I’d love to really scale that up and get enough formation-enthalpy data to compare extensively with simulation and machine learning.

Amazing. What about you?

I’d love to see, if we had access to a full fab, how quickly you could ramp up a new generation of chips — how quickly you could overcome the materials engineering problems in logic, memory, processing, and say, “Here’s the new generation” — how you could take a yield ramp that would normally take many months and accelerate it significantly. I think those would be incredibly interesting experiments.

Wow, fascinating. Okay, last one — if you could recommend a book to everyone on Earth, what would you want to recommend?

I’d say — probably a popular one in Silicon Valley — The Beginning of Infinity, by David Deutsch.

That is a popular choice. On this question, I still have not read it — I don’t know what my mental resistance is to it.

You really mean it?

Yeah, really. I think it has, very much, a spirit of optimism — it gets at this idea of the unboundedness of science. I really enjoyed it.

Yeah, I think I read part of it at some point.

Anyway, I really liked this book called Subtle Is the Lord — it’s a biography of Albert Einstein by one of his friends, who was also an exceptional physicist, Abraham Pais. What makes it unique is that most biographies are written by people who aren’t experts in the subject’s field, so there’s usually a mismatch between what they understand and what they should be saying. This book is incredible because it goes through Einstein’s life with very precise technical detail about what he actually did. It’s an incredible friendship book, because they were friends, and this guy wrote this incredible book for his friend — but it’s also an incredible physics book. So if you’re interested in friendship or physics, it’s a great combination.

I love that. Amazing. Liam, thank you so much — this was a real pleasure.

Thank you so much for having us — it’s been a great time, super fun.

That’s it — thank you for listening to this episode of The Generalist Podcast. Please subscribe on Apple Podcasts, Spotify, or your preferred podcast app. Ratings and reviews help others discover these discussions, so if you enjoyed the conversation, I’d be grateful if you could take a moment to leave one.

For all past episodes and more, visit us at The Generalist Substack. See you next time, as we continue to explore the future.