$ Averth Economic control for AI agents Get the diagnostic

The invoice said $0.20.
The outcome cost $0.82.

Measured across 2,000 instrumented agent runs.

The model call is the cheapest part of the run. Averth reconstructs the fully loaded cost of every agent run: retries, tool spend, context growth, escalation, human review. Then it finds the budget and routing policies that would have prevented the waste.

Per accepted outcome · measured from an instrumented 2,000-run test workload

$0.82Fully loaded cost per accepted outcome
$0.20What the token dashboard reports
61%Of spend from the priciest 5% of runs
6%Of tokens on the terminal path
The ledger

Four times the invoice.
Here is where it goes.

One accepted outcome, measured end to end. The token dashboard stops at the model call. The ledger keeps going through everything the run actually consumed.

Cost of one accepted outcome

MEASURED · 2,000-RUN WORKLOAD
NOTEComponents are rounded and measured from the test workload. Human review is priced at a loaded labor rate; your review cost per run will differ, which is exactly what the diagnostic prices for you.

Statement of cost

PER ACCEPTED OUTCOME
Terminal path
$0.01
Retry & failed runs
$0.19
Across the workload
$159.64
Attribution
Per run, per tool, per attempt
Share of fully loaded cost
≈ 63%
Priced at
Loaded labor rate
The tail

Averages lie.
The tail bills you.

Two hundred runs, ranked by cost. Most of them are nearly free. Then the curve bends, and a handful of runs burn more than the rest combined.

In our 2,000-run test workload, the most expensive 5% of runs consumed 61% of total spend. Mean cost per attempt tells you nothing. The shape of the tail tells you everything.

RUNS RANKED BY COST →TOP 5% → 61% OF SPEND · MEASURED · SHAPE ILLUSTRATIVE
What we meter

Five layers the invoice
never sees.

Token dashboards meter one layer. The other four are where your budget actually goes.

01

Context compounding

The state accumulation tax. Every step re-reads the last, so input context can compound rapidly across long runs. We meter it per step: 4.7× growth from first to last step in the test workload.

4.7x
02

Tool & API spend

External calls, sandboxes, retrieval, browser time. None of it appears on a provider invoice. $159.64 across the test workload, attributed per run and per tool.

03

Terminal-path yield

The share of tokens that produced the outcome. The rest burned on retries, backtracking, and failed runs. Test workload yield: 6%.

6%
04

Heavy-tail outliers

Our test workload showed a pronounced heavy tail: a few percent of runs burning most of the budget. We report p50 / p95 / max per attempt so the tail stays visible: $0.02 / $5.68 / $9.78.

05

Multi-model + human labor

Every model in the mix, attributed separately, plus review and escalation at a loaded labor rate. In the test workload, humans were ≈63% of the fully loaded cost.

63%
The method

A diagnostic, not a dashboard.

No integration, no call. You run the exporter locally and send a small batch of sanitized runs; we return the economics.

01

Send a batch

Run the exporter locally. Prompts and customer content stay in your environment; Averth receives only sanitized economic event metadata.

02

We reconstruct the ledger

Every run attributed across the five layers: fully loaded cost per accepted outcome, the expensive tail, reopen and escalation cost, cost-per-outcome drift.

03

You get the P&L

The statement, plus a replay of the budget and routing policies that would have changed those runs. Findings first, then the policies that fix them.

Beyond the diagnostic

From diagnosis to control.

Averth starts by reconstructing the economics of historical runs. The same policy layer can then continuously identify when a run should be stopped, rerouted, or escalated.

Observe

Meter every run across all five cost layers.

→
Attribute

Trace cost to model, tool, retry path, human.

→
Value

Price each outcome against what it earned.

→
Govern

Identify when a run should be stopped, rerouted, or capped.

Policies we already replay offline: stop after $X per run, route to cheaper models, escalate when continuation cost exceeds human cost, cap retries, flag cost-per-outcome drift after model or prompt updates.

See the policy module →
Open source

Meter it yourself.

The instrumentation library is MIT-licensed. Wrap your agent loop, run your workload, print the ledger. The same code that produced every number on this page.

pip install git+https://github.com/Wesley3141/Averth.git

quickstart.py
from averth import Tracker
from averth.report import report_text

# wrap your agent loop
t = Tracker("support-agent", budget_per_success=6.00)
t.start_attempt(case_id="case-1042")
# ... your agent runs ...
t.log_tool_call("zendesk.search", cost=0.004)
t.log_escalation(1.6, "low confidence")  # human review
t.end_attempt(success=True)

print(report_text(t.pnl(),
      token_dashboard_per_success=0.20))
Questions

Asked often.

Is this another token-cost dashboard?+

No. Token dashboards answer what the model cost. Averth answers what the outcome cost: retries, failed runs, tool and API spend, context growth, escalation, and human review at a loaded rate. In our test workload, the dashboard number was $0.20 and the answer was $0.82.

What do you need from us?+

A small sanitized batch of runs: timestamps, model calls, tool calls, retry counts, escalation flags, resolution labels. No prompts, no customer content, no production access. The pilot runs read-only inside your environment.

What comes back?+

Fully loaded cost per accepted outcome, the expensive tail (which runs, and why), reopen and escalation cost, and a replay of the budget and routing policies that would have changed those runs. Data first, then the policies that fix the waste.

Do you enforce budgets in production?+

The diagnostic uses retrospective policy replay: we show exactly which historical runs would have been stopped or rerouted under each policy. Production enforcement is not enabled as part of the diagnostic.

Send a batch.
Get the P&L.

One sanitized batch of runs. Back comes your fully loaded cost per accepted outcome, the tail, and the policies that would have prevented the waste.

No integration, raw customer content, or call needed first.