hraness
Theme
Appearance

saved

Agent traces are the new oil

by Sam Z LiuXpublished

Hraness republishes this public post from a saved copy. The post is the author’s own words.

Sam Z Liu @samzliu

Building Stash @joinstashai - learning from agent traces Prev. Stanford PhD in AI Alignment + Decision Making (Dropout) | BCG | Harvard

Agent traces are the new oil

Agent traces are quickly becoming the new currency that AI companies across the stack use to transact over. Allegations of Chinese open source distillation attacks are an instance of the immense value that traces hold: acquire them at any cost: beg, borrow, or steal be damned. However, most companies today aren’t properly leveraging this critical asset.

In the chatGPT era, agent conversation transcripts only involved back and forth messages between the LLM and the user. Then with the rise of agents, we added tool calls and reasoning to the mix. The record of user messages, agent internal reasoning, tool calls, and agent responses is an agent trace. These are critically different from their predecessors of conversation transcripts. As long horizon agents come into play, more and more of the trace is dominated by the agent's actions rather than user messages. We have seen a huge growth in attention over agent traces due to three compounding factors that have taken shape over the last year:

  • RL means that current training is largely done over on-policy roll-outs (i.e. traces) instead of pre-training data
  • Long horizon agents with more tokens and actions per run mean traces have become more valuable
  • Agents are being deployed more broadly across organizations

The under-appreciated asset

The right traces in sufficient quantities make up the backbone of most AI companies today. They contain the plethora of information needed to move AI forward: reasoning, tool calls, retrieved information, errors, outputs, and human feedback. For companies that are building AI, these traces fuel the arms race of ever improving models and agents.

  • For data vendors: The right traces annotated by human experts or generated against RL environments can make you a small fortune if you sell them to frontier labs.
  • For frontier labs: A competitor's traces can help your model catch-up while properly annotated ones in the right domains push your model to new frontiers.
  • For the application layer: The traces from production systems can help identify bugs or user experience issues while the traces produced by eval systems help you hill-climb against north start metrics.

For users and enterprises deploying agents, however, traces are largely ignored and under-appreciated. They are seen as logs, to be filed away for monitoring and compliance if they are stored at all. (Here's a quick reminder that Claude Code automatically deletes your sessions after 30 days unless you set it otherwise!) However, as agents do more and more work within enterprises, these traces are no longer disposable logs. They become the operational work of the company itself.

What's a trace mean to you?

If you’re deploying agents, there are three main ways in which traces remain useful beyond the original task the agent is assigned:

Operational excellence -

  • Understand token spend and latency
  • Measure ROI of workflows
  • Diagnose why an agent failed
  • Discover uses that customers invented but the product team did not anticipate.

Safety Guardrails -

  • The Hugging Face incident is a warning signal for what can happen without robust guardrails.
  • As multi-agent systems scale, we will need better ways to keep a tab on the sheer volume of data. OpenAI reportedly used 10k agents running in parallel on their solution to the Navier-Stokes problem.
  • As agents are given more power (calling APIs, interacting with customers directly, spending money, etc.), it's important to limit downside exposure.

Continuous improvement

  • Successful traces contain reusable playbooks and procedures that should spread across a team or organization. Even failed traces contain human feedback, correct work, and other learnings that can be put to use.
  • Today, the improvement of agents and models given traces is relegated to companies at the cutting edge that can afford a post-training team
  • Tomorrow, it will be important for every company to own their own intelligence to avoid vendor lock-in and maintain their competitive positioning against other companies with the same base models.

Why extracting value from traces is hard

The relative simple nature of traces is deceptive. They seem like just a log of what your agent has done, but actually making use of them effectively in production proves challenging:

  • The data is unstructured. There is usually a mix of prose, JSON schemas, tool calls, user messages, code, and even images. Making sense of how everything is connected is a non-trivial task.
  • There is a huge volume of data. Agents can produce tokens much faster than humans can, and ingesting all of those tokens with yet another LLM is prohibitively expensive, usually representing a meaningful fraction of the cost it took to produce those tokens in the first place. Classification-focused models like Jev may change the equation here however.
  • The feedback is messy. Success and failure may be hard to determine from the trace alone. Users may not give any feedback at all or give implicit feedback such as not continuing the session.
  • Long horizon credit assignment is hard. A task may have failed but contain mostly correct decisions and traces except for at the end or have succeeded but with a lot of failed intermediary steps that shouldn't have been taken. It's unclear what decision in a long trace led to success or failure.
  • The judgements are qualitative. Most insights we would want to pull from a trace are judgement calls which require more "feel" than typical verifiable signals can (e.g. is this safe to run? is this what the user intended?)
  • Interpretability could get worse. Neuralase, recurrent transformers, and closed source models refusing to provide reasoning traces all point to an emerging world where the reasoning of models that power agents become more opaque over time.

The new oil

There’s an analogy that’s especially pertinent here. In the 2010s, companies were scrambling to build cloud infrastructure, data lakes, and ontologies. The hope was that amassing this wealth of data in one place would enable the organization to more effectively leverage it to optimize operations, drive customer value, dissolve silos, and support advanced analytics.

Today, we are seeing a new trend around agent traces and context graphs that rhymes with the past. Companies increasingly want to own their own intelligence and avoid the reverse information paradox. In a world where agents automate most of a business and everyone has access to the same models, your competitive advantage is only as good as the ability for you to build your own proprietary layer of intelligence that improves over time from experience

This is similar to the institutional knowledge that lives in employees’ minds, except the minds of tomorrow are electronic. The medium for that knowledge is the agent traces. This will contain the history of a company’s operations, what good looks like, and how it makes decisions, becoming a new strategic layer for businesses. And like big data and oil that came before, it only gains value through extraction, refinement, and distribution.

If you are thinking through how to make use of agent traces, please reach out!

Weekly agent token volume on Openrouter. Y-axis labels 0, 10T, 20T, and 30T. A point near Dec 2024 is labeled 0.4T. X-axis labels Dec 2024, Jun 2025, Dec 2025, and Jun 2026.

Length of software tasks that different LLMs can complete 50% of the time. Measurements above 16 hrs are unreliable with our current task suite. METR. Task labels: 16 hours, Optimally reduce the size of a language model; 4 hours, Train adversarially robust image model; 1 hour, Train classifier; 6 min, Find fact on web; 36 sec, Count words in passage; 4 sec, Answer question. Model labels include GPT-2, GPT-3, GPT-3.5, GPT-4, GPT-4o, o1-preview, o1, o3, Claude Opus 4.6, and Claude Mythos Preview (early). X-axis: LLM release date, 2020 through 2026.

Cover illustration: no readable text.

Line chart titled Weekly agent token volume on Openrouter, rising from about 0.4T in December 2024 to above 30T by June 2026.METR chart of the length of software tasks different LLMs complete 50 percent of the time, from GPT-2 at a few seconds in 2020 to Claude Opus 4.6 and Claude Mythos Preview near 16 hours in 2026.Illustration of a pipeline pouring glowing traces like oil across a tiled floor, with derricks in the distance. No readable text.