saved
How Datadog built a universal machine tool for Claude Code
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Sesh Nalla, VP of Engineering at Datadog, says about 90 percent of Datadog used AI coding tools on production code in four months. Claude Code drove at least two thirds of that, about 3,000 engineers. A hand-built queue, courier, took a year. A Kafka-compatible system he calls Helix took about three days with Claude Code doing most of the construction. Temper is the machine tool: agents write a state-machine spec, a verifier checks it, and the runtime can hot-reload it.
ideas
- About 90 percent of Datadog used AI coding tools on production code in four months. He says that is roughly 3,000 engineers, and Claude Code drove at least two thirds.
- Courier took a year by hand. Helix, a Kafka-compatible system, took about three days. Claude Code did most of the construction, and he says production still needs mileage.
- Temper makes the state machine explicit. Agents write states, actions, and triggers. The compiler sits outside the model, and a spec change can hot-reload.
- The verifier gates a live load. He describes checks on reachable states, fault schedules, and property tests, plus policy on which agent may mutate a rollout.
quotes
“In the last four months, about 90% of data dog used AI coding tools for production code.”
“It took us one year to build courier.”
“So you're no longer writing the code.”
“So the idea is Tempere makes that state machine explicit.”
transcript
V.P. of Engineering of Data Dog, Sesh Mala. I know it is 4 p.m. How's the energy going right now?
So far. How are you liking the conference? Great. All right, let's see some blood flowing.
So, picture of hands if you have heard of machine tools before. What's on the slide?
Have you used any? I was expecting zero hands anyways.
So, okay, so that's the talk anyways. So, these are tools like jigs, fixtures, gauges and melts
that you see in manufacturing to produce precise and repeatable machine pods.
The kind that you assemble them into larger machines, like engines, aircrafts, nuclear reactors, lunar landing
modules that you saw this morning. So, machine tools were a breakthrough
during the industrialization period that enabled the scale due to the interchangeability
with standardization and precision. So, that was the inspiration for my talk
and what we built at Data Dog. And I will share with you how all this fits into building
software with clock code. So, what you're seeing on the slide there is my view
of the last 18 months. I think there is a different view of this graph
at multiple sessions today. These non-scientifics, I don't please consider that my eval,
is purely personal, but is the case of ambition with the models.
For most of 2025, the models were useful to me, but within very narrow boundaries,
like local changes, small functions, tests, glue code, throwaway prototypes for me to learn a lot,
kind of learned a lot through them. And then, around late 2025, I think you all have seen this
and noticed it when the slope started to change, like this exponential.
I started trusting Claude with larger and more ambiguous systems-coped work.
Before I talk about machine tools, I wanna go through the lineage of some ideas
that led me here. So, this is 2024.
It's a write-up. We were building a distributed queuing system
called courier from scratch. All of this was before agents, all by hand.
Can we believe it? Like, still, software was built by hand then.
The hard part was, for any distributed system that is human-built or agent-built,
it's not just building the parts, but it's building the pieces and making the interaction
between them observable, testable, and verifiable. So, we were rigorous with formal modeling and simulations,
so you see various techniques in this post. All of this is classical systems work,
where you identify the parts, where mistakes would be expensive or hard-prevers,
and you raise the rigor for those parts that you don't want them to slip to production.
The next idea was, around September 2025, we called it bits evolve.
It's a closed-loop evolutionary optimization harness, inspired by Alpha Evolve from DeepMind.
The idea is that you let parts of code improve themselves on a narrow controlled hardness.
I think we've seen some announcements today at the keynote about dreaming and various loops.
So, this was us trying in September 2025. It's an ensemble or a council of models, big and small,
and they generate variants of code, whatever you designate, you want them to improve.
In our case, it's heart functions, blocks of code. And then, you have a cascade that you see
on the right-hand side of the screen, with benchmarks, tests, and production observability,
decide what survives. It's like natural selection.
So, this was the first glimpse for me, that parts of software, maybe, could be cultivated,
like living organisms, like plants or microbes, and growing through variation with feedback and adaptation.
However, the insight was that this kind of evolution is only as good as the environment it adapts within.
If you have bad benchmarks, they produce bad evolutions. Weak observability, your optimizations are shallow.
So then, during this period of leap, or the exponential that started,
we were building whole distributed systems with Cloud Code. I mentioned the Opus 4.5 inflection point,
where I raised my ambition. It wasn't like sudden, like I didn't wake up one day
and then started trusting Cloud with whole systems, or even my databases, we have a lot of them.
It happened progressively, like I tried how to task, then larger subsystems, failed a lot,
learned a lot through a lot of experiments, and then started seeing successes around this period.
So, what we decided to be more ambitious, could we build something as big as Kafka?
Show of hands if you've heard of Kafka. Okay, more hands now.
So Kafka is a streaming service. So we kind of attempted, okay, can we build this from scratch?
It's like repeating the same playbook that we had for courier. It's a queuing system, pretty close, right?
Same rigor. So you will see the Kafka on the right-hand side of the slide.
But Cloud Code doing most of the construction with one human building it.
In a few days, to our disbelief, we had a full functional Kafka compatible system working.
And we called it Helix. So the source code methodology and the details are all linked
in this post, feel free to check them out later. But taking Helix to production now requires building mileage
and that has been challenging for us. So the next natural move for us was,
I spoke about bits of valve earlier. Could bits evolve evolve parts of Helix
the same way it evolved hard functions. The ambition was like too big,
from small functions to like, can the whole component evolve? Can it provide the mileage we needed so we could, you know,
get to production? The answer was not quite.
The surface area was too large. Even with the verification cascade that I showed you
in the previous slide, quite rigorous. There were too many places where the human had to interpret
and correct, it's too multi-ton, too interactive. So you're like, okay, can we look for a narrower surface?
Can we dial back the ambition a little bit? So this post was about that.
So we chose our metrics aggregation server. We are data docs, we have lots of metrics.
Could we improve the materialization logic live? Not offline, like what we did with Bitsov all before.
Can we optimize them poor customer with a proof carrying path around the change?
So a human doesn't have to review every candidate that is being generated.
So if you look at the flow across these four projects, if you observe the pattern,
each project exposed the bottleneck for the next. There were many talks and comments today
about the bottleneck being moved. So with courier, the bottleneck was construction
with humans building systems through careful design, modeling and verification.
It took us one year to build courier. And then if you compare it, it took like three days
to build Kafka in like what, 12 months. So with Bitsovolve, the bottleneck moved to the feedback loop.
Like the model produces variation, but the harness decides what survives.
And with Helix, the Kafka system that we built, the bottleneck moved again,
where agents could build large parts of the system. So we have seen it.
But then the human have to coordinate to ship the work to production through tools and mechanisms
built for humans. So that's the Amdos law that Dario was talking about
earlier today. So that's the jump we are making.
If you look at this slide at the center from mechanization to industrialization.
So mechanization means agents are doing more of the work now. And industrialization, if it were to borrow the metaphor,
and that's why I kind of introduced machine tools, means work becomes repeatable, verifiable, controllable,
and scalable. So the idea of a machine tool is, I know it sounds cool,
but why do we need them now for agent built software? Are in these four just lunar modules to land?
Because of the complexity and the ambiguity growing. Right?
So each time we are trying to like increase our ambition levels, started from targeted changes on existing systems.
In the last four months, about 90% of data dog used AI coding tools for production code.
That's roughly 3000 engineers. And cloud code drove at least two thirds of that.
And most of the work as I described about helix was still single human driving,
like one engineer steering one or more agent sessions. And the work was moving across this map.
More complex to generate on one axis and more ambiguous to verify on the other axis.
Here are a few concrete examples. I'm not going to enumerate all of these,
but feel free to take a picture if useful. And if any of these resonate with you,
I'm happy to chat about more afterward later on. But the main point I want to underline here is that
these are generating personalized flows in our software development lifecycle.
Because one human could do a lot, a lot more than what they were used to do before.
So for engineers, in the software delivery lifecycle I was talking about, if I were to use the word flow,
it used to mean direct relationship between intent and code. You understood the problem, you wrote the code,
you tested it, you reviewed it, you shipped it, you operated it, you repeated it again and over and over again.
But with agents, that abstraction level is changing rapidly. I haven't seen code, I mean, I haven't seen code in a while.
I've made my piece with it because I've been a manager for a while, but many people, like they're still trying
to go through their seven stages of grief, of not seeing code on a day to day basis.
So you're no longer writing the code. You're shaping the work.
You're deciding what the agent should see. We saw the keynotes today, like outcomes, right?
What tools it should have, what success means, how failure should be detected.
All of this is powerful, it's like everyone's promoted three levels up into management chain,
which they weren't signed up for. Engineers, right?
Because it's a huge leap, and it's also disorienting, it's like you pushed against gravity very fast.
It can feel sickening, right? We haven't acclimated to this altitude of working,
specifically engineers, who love looking at code. So before this jump in model capabilities,
the human team was the factory. Tools were designed around human attention, our judgment,
and the operational memory of what's actually happening in production is in this human organization,
our collective brains, our minds, connecting back to courier, like I said,
only 12 months ago, this was the world we were in, and only four months ago, it continued and started to change.
And that's the inflection point with Cloud Code and Opus 4.5. We are starting to see one lead human,
coordinate multiple interactive sessions. I've seen like many screenshots of like parallel sessions,
like cranking out stuff, it's disorienting for me to personally watch.
Three, four, five agents working on different parts. I heard Jarred today saying he's doing 10 things at a time.
Like the stuff she was showing was only 10% of his, whatever time being spent.
So these tools were still human-shaped. The agents were two orders of magnitude faster,
and this tool chain isn't built for their speed. So what happened, the human became the bridge between
agent execution and the human-shaped systems. And now all this operational knowledge,
like you wake up at night, something broken, it's just in that person's head,
and probably in some markdown files that agents work on, and in between them.
So that's for the ramplified with Cloud-managed agents. So agents are doing a lot more background work,
it's compute, and they start taking judgment bearing rolls. Meaning they are not just following your instructions anymore,
they're making their own decisions, and they're running longer for hours,
sometimes overnight, sometimes for days. I don't know the longest task,
someone like benchmarked 20 hours, 28 hours. So they construct their own tools in these sessions,
they write their own code. And the mismatch is still there,
like each agent invents its own tools, its own glue, its own conventions,
and that system becomes really hard to share and operate. You can see the blur between what the agent sessions produced
as like intermediate tools, and what is your product doing, like your code.
You get a lot of output fast, a lot of it is useful, but some of it can look like false progress,
and most of the tool construction that only makes sense in that local session.
And you start to see that blur. So this is where I felt we need something more structuralateful.
If agents are going to build an large parts of our systems of databases,
which are mission critical, they need of this machine tool concept
that I am trying to introduce. instead of agent inventing disconnected tools for every local need,
it produces precise specifications of the intent and problem domain. A jig or a CNC machine
have you seen some of those that computer aided hot machines where you give them specifications of what your screw threading needs to be,
what this needs to be, it is extremely repeatable, you can run them,
and you can build aircrafts and things like that with them, right? So in this case, the agent does notDate
and provides the final mechanism each time. It produces a precise description,
and iterates with tamper or tamper like mechanism to make something work first,
and then later turn that into something repeatable, checkable and reusable.
So you could actually build a software factory around your code base. So that's the concept of a dark factory.
Simon Wilson of SimonVilison.net. Pretty amazing blog.
I think he is one of the most influential AI voices right now in teaching how to work with agents and build software
has been popularizing this phrase called dark factory. I think it's a pretty good encapsulation of a software process
where the agents keep working without the humans on the virtual-factory floor.
You can turn off the lights. So the human role now becomes like designing the factory
and the constraints and the outcomes and the verification loop. So this thing can run for hours and days and流s
producing what you wanted to produce. So something like tamper can fill in this role
of a machine tool to build such factories. Let's look at a dark factory concretely with Helix as a target.
What I shared about Hel It's a Kafka-like streaming, streaming service.
It is probably one of the five expensive services we are running at Data.gov, like impatience.
So we have been shadowing our production workloads with Helix. In some cases, we actually believe it can be materially cheaper
than our current production solution. Can you believe it?
We took a week to build it and we started shadowing it and we saw like two to five X opportunities
that it can be cheaper. But getting from this promising state
to production still takes to work on the mileage the system needs to earn that it can run
and multiple people can operate it not just a person who built it.
So we created a bunch of synthetic workloads that models our production shapes
and we constructed a factory. So software factory for Helix uses
Tempere in three distinct ways. First, as an agent control plan to cloud managed agents
where the sessions, roles, work, use and operational life cycle are more precisely managed.
Second, as a way for agents to build their own tools with small Tempe, orsenal apps bridging the STLC tooling
like Git CI deployment. And third, as a Helix control API,
. The interface that and the life cycle surface around the Helix data plan to exercise this workload.
So that was a surprise to me. The surprise was it started to feel more general
than agent infrastructure. A lot of software if you squint closely
is just control logic around database, APIs around state, policies around mutation
life cycle transitions, integrations with external systems. So Tempere could be this universal in a sense
that it can be applied to any software that has the shape I described.
Before I go deeper into how Tempere works, you might be wondering at this point,
why is this different from asking Cloud code to build a CRUD app like in TypeScript or Python?
Cloud can do that very well. We have seen like lots of PRs and lots of code going.
However, in normal CRUD apps, the control logic is spread across routes,
database constraints, service code, background jobs and documentation.
It may all have good tests and coverage, but the operational model,
which is generally a state machine shipped, is mostly implicit in the code base.
So the idea is Tempere makes that state machine explicit. This is in particularly novel.
We had this with runtimes like Erlang OTP for decades, actor runtimes, and more recently with workflow engines
and durable execution runtimes like Tempere, they all have like popularized this precise runtimes,
so you can like run1016 applications. So let's lookreenshots some Tem
would arrive at this precise declarative artifacts. And then you have a compiler equivalentilippre,
verify it, and it can hard deploy into a runtime. There are other runtimes who does this, like Ar,
like Erlang Beam, does that if you have heard of it? Anybody heard of Erlang Beam?
There you go. So the run path feels the same as any other CRUD shaped API.
You wouldn't notice the difference. It's important to note that the agent is not generating
arbitrary application code directly. It's, wehillary the abstraction.
It's generating a structured description that compiles into a runtimeship.
The compilation step is outside the other run. It's same like you write Rust code
and you give it to Rust compiler, right? So Tempere turns this blueprint
into something called formal state transitions. It's very common in functional programming
and actor runtimes if you have used them. This is the most important technical detail.
Formal here doesn't mean every possible property is proved. It isn't theoretical.
It means the basic shape of the application is represented as a precise transition system.
And when you have that, it is a much better reasoning for both humans and agents than arbitrary code.
Pardon myustainable design of this slide with no syntax highlighting. But that's an agent written spec.
For a helix rollout, that spec looks like this. States, actions and triggers.
I actually got an idea this morning when I saw the announcement on Cloud managed triggers.
Maybe this could be just a pass through. You could declare a trigger here and then have Cloud run them.
But the idea is you define your state's actions and triggers or the agent does it, Cloud does it.
And in this case, like the deployment description starts rolling is only valid when planned or entity moves to rolling
and that triggers a side effect, like going and patching the Kubernetes, which is not idempotent.
And then a callback that comes back, whether to mark the progress or fail. What Tempere does, so that artifact is declarative, so what Tempere does,
it takes this and generates that spec into this transition table, a concept where the critical control logic is data-like.
It is not just spaghetti, imperatively encoded in code.
It's just data-like thatPictangible, uncheckable, and it is not hidden in any improvised chain of service methods.
This is easier for agents to work with. And they can change this dynamically with safety.
That's the promise. I have seen people like writing,
orling scripts and hard deploying into Beam. I think that's a pretty creative way of using runtimes that are
extremely hardened over higher-shroomed systems for agents to work with behav So in this case, if a rollout needs a new state or a rollback path,
the agent can make a targeted spec change and it can hot reload it. So the iteration speed is even faster.
You don'tasures to go through CI and then deploy and everything. Sooccupied when you're leaving agents overnight,
I think they can come up with some pretty good progress overnight without compromising on safety, PR reviews and code reviews and things like that.
And we have policy gates. Because the transition table is data-like,
Tempere evaluates the state transitions and the policy decisions together, such as who can mutate a deployment rollout,
which actions can an operator agent, like the team of agents that are working on the Dart factory, you can say this operator agents can only do so much,
or these actions are forbidden for a builder agent. And you can ask, or you can specify independently as a human,
can a completed rollout be rolled again, or can a failed tool result continue on its rollout path, et cetera.
And we have the side effects, an effect system, also very popular in typed runtimeselfare like TypeScript.
Effects are deliberately small typed operations here in Tempere. Keeping them small prevents the state machine from becoming a back door for
arbitrary application behavior. But if you need arbitrary application behavior,
you could package them as WASM modules. Familiar with WASM?
Come on. Mahjong. So this is where arbitrary code generated by an LLM can leave.
It's a very narrow over place. So what that means is it makes troubleshooting easier.
On the last building block for Tempere is the verifier. Verifier is, verification is basically the bottleneck right now for pretty much everything.
We have seen so many announcements and discussions around it. So in Tempere, the verifier is a gate before the transition table is loaded into the runtimes.
That's what allows it to say this is safe. You can put this and load it live.AMD is multiple levels like a Swiss cheese pattern.
Not all levels need to find everything exhaustively. The agent will clot is generally very good at making judgment calls.
Do I need to run all levels or just some levels which same like it uses its compiler output. So we have to start learning.
Level 2, model checks the reachable state graph. Can any path reach a bad state?
And level 3 runs schedules, injects, falls. Timing failure conditions and et cetera.
And level 4 uses property tests, which I highly encourage that we Bulgarge that we have to start learning. All of this is an exhaustive one day one.
Every discovery gap compounds the verifier. I also heard Boris mentioning the compounding effect of you find tests and you keep the
blocks fixing the gaps. A missing condition revealed in production or simulation can get added to the model or the
test suite. And it gets compounds.
So where is all this going? I don't know.
I can't predict much aboututral where this could go. But the idea is if each artifact of a temporal app or a bundle is concise, fits in your head,
I spend a lot of time running machine critical infrastructure for data. Thousands of times past three or four years.
You won't be able to keep your complex machine critical logic in your head when you want to Ho operate it.
For a complex business domain like banking or financial systems, you still should be able to use it forBronze.
"},". However,tnc the cost of Oilers was too.
It was too highiscons softwarePope.theless. The machine built with machine tools made parts composable and
inspectable and replaceable that we could build our machines. And my claim is that for software agent built for a scale
scale from here on, if you were to build databases and put them in production. I learned whereas if agent skin builds software autonomously inside with such discipline and rigor,
maybe, maybe, we don't need to stop at dark factories. I don't know, there is a word dark in there.
It kind of sounds sad. So the whole software built this way can feel like an organism that we can grow,
cultivate and evolve through feedback, selection and adaptation. And it looks like, I don't know, agriculture or,
or a directed evolution where you choose the direction in which your software needs to go. You can choose Kafka to be a queue because there are more customers using it as a
FIFA queue where it says no, we need more buffers. We don't need all these brokers.
Or maybe just dreaming, right? We have dreams now in cloud.
That's all I have. Thank you so much for listening.