saved
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Zhening Li et al. introduce JAZ, whose invoke primitive is a minimal agent loop that treats LLM-written code and recursive subagents as the harness. Each invoke call has its body generated at runtime; inputs, the prompt, and REPL history are ordinary variables. With prompting alone, invoke reaches 70% on StuLife far-recall versus Letta at 62% and about half the cost, and 74% on AppWorld continual self-improvement versus ACE at 70%. Code is at github.com/jaz-lang/jaz. The HTML capture is partial after a rewrite safety limit and one redacted credential-shaped fragment.
ideas
- invoke is a language primitive, not a toolkit stack. The LLM writes the body of each call; recursive invoke is the default subagent path, generalizing CodeAct and recursive language models.
- Everything the model sees is a REPL variable. Named inputs, the user prompt, and interaction history are programmable objects, so agents can search prior history instead of only reading a message list.
- StuLife far-recall favors prompt-only invoke over specialized memory. With GPT-5.4 nano, invoke averages about 70% pass on far-recall tasks versus Letta at about 62% and roughly half Letta's dollar cost.
- AppWorld self-improvement also stays inside the loop. A top-level GPT-5.4 invoke with nano sub-invokes averages about 74% TGC, ahead of ACE on CodeAct at about 70%, while editing prompts and skills from test feedback.
- Hooks and dynamic scope keep the core thin. Observability, budgets, validation, and ConfigOverride attach locally or via scope without replacing the invoke loop.
quotes
“We study whether a minimal harness that is little more than the agent loop itself can exhibit these capabilities.”
“everything visible to the model is also a variable in the code environment”
“With GPT-5.4 nano, it achieves 70% on the recall-heavy subset of StuLife”
“it achieves 74%, outperforming CodeAct (68%), CodeAct+subagents (71%), and a specialized self-improvement harness, ACE”