Hraness
Theme
Appearance

saved

My code is disposable now

by carsonfarmerXpublished

Hraness republishes this public post from a saved copy. The post is the author’s own words.

carsonfarmer @carsonfarmer

Engineering and R&D | Cofounder & CTO at @recalllabs_ | AI, Machine Learning, and Distributed Systems | Former Prof. | 🇨🇦/acc

My code is disposable now

I started a new project recently and wrote about it. Because I spent the time up front to get the design right, building it took very few prompts. It ended up as iroh-acp-go, ~178 lines of Go that lets an editor use a coding agent on another machine without any intermediary.

A program that small is cheap enough to write again, so I decided to embrace that and treat the code as disposable. The part worth keeping is all the thinking that came before it: what the program should do, what it should leave out, and why. Most of that lived in my head and in a chat history. If I wrote it down well enough, an agent could rebuild the code from it whenever I needed.

Regenerative software

Chad Fowler calls this regenerative software. If an agent can write the code again whenever you need it, the code stops being the thing you protect. You protect whatever lets you write it again and check that the result is right. In a follow-up, The Phoenix Primitives, he lists what that takes: a description of what the software must do, tests that can judge any version of it, a limit on what the agent gets to see, and a record of how each version was made.

He's also written that the spec can come later. You build something to find out what you want. Then you pull a spec out of it and rebuild from the spec to see whether you wrote everything down. That's the order I went in. I didn't find anything on how to lay all this out in a real repository, though, so I made that part up.

Writing it down

Everything lives in a.regenerate/ folder. The spec says what the program must do, down to the command-line flags and the log lines. A decision log says why. A prompt tells a fresh agent what to do with the two of them, and a ledger records every rebuild, including the ones that fail.

The folder and the code do different jobs. The code, and the tests I wrote for it, are the product. A rebuild replaces both. The folder is the infrastructure around them. It holds the behavior I pulled out of the code, and a way to check any version of the program against it.

A second set of tests, which I call the spec suite, checks any version of the program against the spec. It only looks at the code from the outside: the functions the library exports, the flags each command takes, and how the server behaves on the network. Fowler would call my own tests ephemeral and the spec suite a durable evaluation. The suite runs against my code on every push, so if the code and the spec drift apart, the build fails.

Rebuilding it blind

A spec is only worth keeping if someone who has never seen the code can build the program from it. To find out whether mine was good enough, I gave a fresh agent the spec, the decision log and the prompt, and nothing else. It couldn't see my code or the spec suite, and it had to write its own tests. I held the suite back on purpose. An agent that can see the tests will write code to pass them and skim the spec, and then I learn nothing about the spec. When the agent finished, I ran its tests and the spec suite against what it built.

I did this twice, with Sonnet 5.5 running as a subagent in Claude Code. Both rebuilds passed everything, and both stayed under the line limits, etc I'd set for the project. (I also did this for the Rust version!).

What the rebuilds found

Chad has a post called The Implementation Remembers. It's about how working code picks up knowledge that nobody wrote down, and how easy that knowledge is to lose when you rewrite the code. The rebuilds were how I found out "what my code knew".

My spec was bound to have gaps, and an agent building from it would have to fill them. I told the agent not to stop and ask. It should pick something, keep going, and at the end hand me a list of every choice it made on its own. Both rebuilds passed the tests, so those lists are where I found out what the spec was missing.

Most of the gaps were small choices my code had already made. My "original" client ignores any arguments after the ticket, for example, and the spec never said so. The agent that wrote the original code made calls like that as it went, and nobody wrote them down. Other gaps were places where the spec contradicted itself, like two sections that disagreed about what the server prints when you start it with missing arguments. The spec now covers all of these.

The first agent also found a quirk in my code. When the server rejects a client whose ID isn't on its allowlist, it sends back "not allowed". My library has two ways to connect, and only one of them, Dial, passes that reason along. The other, ConnectAgent, goes through acp-go, which reports "context canceled" instead. My command-line client uses the first, so I'd never noticed. The spec now covers this as well.

The rebuilds also caught a documentation error. My README said that if an editor kills the client, the agent on the server shuts down within 10 seconds. Both rebuilds measured 10 to 15 seconds. When I timed my own code, it did the same. The spec and the README now say 10 to 15.

How much to write down

Every gap I fixed made the spec longer, and that can only go so far. Gabriella Gonzalez argues that a sufficiently detailed spec is code. Mario Zechner makes the same point in a talk about building pi, and warns about the opposite problem too. If you leave blanks in a spec, a model fills them with whatever it learned from code on the internet, and a lot of that code is bad.

I want something in between. Code has to settle every detail, whether it matters or not. A spec gets to choose. Mine should pin down what matters and leave the rest open, where any reasonable choice will do. The agents' lists helped me decide which was which.

The first agent had to guess what my server should return when you stop it. My code returns a particular error, but nothing should depend on which one, so the spec now leaves that open. The agent also had to guess whether stopping the server should close its network endpoint. The spec now says no, because the program that created the endpoint may still be using it.

Any settings you pass the server replace its defaults, except the protocols it accepts, which are added to its own. My spec left out that exception, so the second agent built what the spec said rather than what I intended. Now the spec is explicit about this behavior.

What's next

The code is still the same 178 lines. The thinking behind them is written down now, and two agents have rebuilt the code from it and shown me where it fell short. When a dependency ships a new release, I can have an agent rewrite the code against it and let the spec suite decide whether the result is good enough to keep. The next entry in the ledger will be a third blind run against the fixed spec, and I should probably run that one in a sandbox so the agent can't stumble on my code.

Update

I also did all of this for a Rust version, because, why not?

https://github.com/carsonfarmer/iroh-acp-rs

Cover graphic: a glowing orange phoenix rising through translucent Go source code on a dark blue field; no readable title text.