Measured across 2,000 instrumented agent runs.
The model call is the cheapest part of the run. Averth reconstructs the fully loaded cost of every agent run: retries, tool spend, context growth, escalation, human review. Then it finds the budget and routing policies that would have prevented the waste.
Per accepted outcome · measured from an instrumented 2,000-run test workload
One accepted outcome, measured end to end. The token dashboard stops at the model call. The ledger keeps going through everything the run actually consumed.
Two hundred runs, ranked by cost. Most of them are nearly free. Then the curve bends, and a handful of runs burn more than the rest combined.
In our 2,000-run test workload, the most expensive 5% of runs consumed 61% of total spend. Mean cost per attempt tells you nothing. The shape of the tail tells you everything.
Token dashboards meter one layer. The other four are where your budget actually goes.
The state accumulation tax. Every step re-reads the last, so input context can compound rapidly across long runs. We meter it per step: 4.7× growth from first to last step in the test workload.
External calls, sandboxes, retrieval, browser time. None of it appears on a provider invoice. $159.64 across the test workload, attributed per run and per tool.
The share of tokens that produced the outcome. The rest burned on retries, backtracking, and failed runs. Test workload yield: 6%.
Our test workload showed a pronounced heavy tail: a few percent of runs burning most of the budget. We report p50 / p95 / max per attempt so the tail stays visible: $0.02 / $5.68 / $9.78.
Every model in the mix, attributed separately, plus review and escalation at a loaded labor rate. In the test workload, humans were ≈63% of the fully loaded cost.
No integration, no call. You run the exporter locally and send a small batch of sanitized runs; we return the economics.
Run the exporter locally. Prompts and customer content stay in your environment; Averth receives only sanitized economic event metadata.
Every run attributed across the five layers: fully loaded cost per accepted outcome, the expensive tail, reopen and escalation cost, cost-per-outcome drift.
The statement, plus a replay of the budget and routing policies that would have changed those runs. Findings first, then the policies that fix them.
Averth starts by reconstructing the economics of historical runs. The same policy layer can then continuously identify when a run should be stopped, rerouted, or escalated.
Meter every run across all five cost layers.
Trace cost to model, tool, retry path, human.
Price each outcome against what it earned.
Identify when a run should be stopped, rerouted, or capped.
Policies we already replay offline: stop after $X per run, route to cheaper models, escalate when continuation cost exceeds human cost, cap retries, flag cost-per-outcome drift after model or prompt updates.
See the policy module →The instrumentation library is MIT-licensed. Wrap your agent loop, run your workload, print the ledger. The same code that produced every number on this page.
pip install git+https://github.com/Wesley3141/Averth.git
from averth import Tracker from averth.report import report_text # wrap your agent loop t = Tracker("support-agent", budget_per_success=6.00) t.start_attempt(case_id="case-1042") # ... your agent runs ... t.log_tool_call("zendesk.search", cost=0.004) t.log_escalation(1.6, "low confidence") # human review t.end_attempt(success=True) print(report_text(t.pnl(), token_dashboard_per_success=0.20))
No. Token dashboards answer what the model cost. Averth answers what the outcome cost: retries, failed runs, tool and API spend, context growth, escalation, and human review at a loaded rate. In our test workload, the dashboard number was $0.20 and the answer was $0.82.
A small sanitized batch of runs: timestamps, model calls, tool calls, retry counts, escalation flags, resolution labels. No prompts, no customer content, no production access. The pilot runs read-only inside your environment.
Fully loaded cost per accepted outcome, the expensive tail (which runs, and why), reopen and escalation cost, and a replay of the budget and routing policies that would have changed those runs. Data first, then the policies that fix the waste.
The diagnostic uses retrospective policy replay: we show exactly which historical runs would have been stopped or rerouted under each policy. Production enforcement is not enabled as part of the diagnostic.
One sanitized batch of runs. Back comes your fully loaded cost per accepted outcome, the tail, and the policies that would have prevented the waste.
No integration, raw customer content, or call needed first.