saved
Ephemeral testing
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Daniel Lemire proposes ephemeral testing, where an AI agent builds and tests throwaway software on top of your code to judge that code. Lemire calls it a form of integration testing in which the software on top is discarded afterward. A library with a clean API, stable invariants, and useful errors lets the agent build something that works quickly, while hidden state, surprising defaults, or incomplete docs produce patches and failures. He says rebuilding everything with AI whenever needed is not practical, because software needs some stability.
ideas
- Agent failures are evidence about the foundation. Lemire says a library with hidden state, surprising defaults, or incomplete docs produces a pile of patches and failures, and those failures reflect on the library rather than the agent.
- Building the upper layers replaces guessing what they need. Instead of designing the core while anticipating other layers, the agent simulates those layers by building them and then throws them away.
- The test can be repeated with different agents and tasks on the same foundation. Lemire applies it when considering a new feature, asking his AI to quickly prototype what he might later build on it.
quotes
“You do not assess the original work directly. You assess how good the software built on top of it is.”
“The failures are evidence about your code, not about the agent.”
“you just simulate the other layers by actually building them.”