lgoyal6 / agentbench

agentbench.

An evaluation harness for agents that treats token economics as a first-class metric. A model that scores five points higher and burns fifty times the tokens is not obviously the right choice, so every number here arrives next to its cost.

This is the first reference run, and it is one model over four suites. Reported with the interval a run this size actually supports, which is the thing an accuracy figure alone hides.

tool use
15/15 and a 0.80 lower bound
dearest correct answer
$0.0011 math reasoning
cheapest
$0.0004 summarization
runs published
4 of 5 one is void
Figure 1

Accuracy, and the interval it actually supports

Each suite is between ten and twenty tasks. At ten, one task is ten points, so the Wilson interval is drawn next to the score rather than left out of it. A run that reports 1.00 has not shown you a perfect model; it has shown you no failures in ten tries.

Loading
Suite
accuracy
-
95% interval
-
tasks
-
per correct answer
-
latency p50 / p95
-

Why the interval belongs on the chart

Summarization scored 10 out of 10. The 95% interval on that runs from 0.72 to 1.00. Tool use scored 15 out of 15 and its lower bound is 0.80. Both are good results and neither one distinguishes a model that is right 99% of the time from one that is right three times in four.

Nothing about that is a criticism of the harness. It is what a first reference run costs, and printing the bar makes the size of the run impossible to forget.

what 1.00 means here
No failures in ten tries.Lower bound 0.72.
what would fix it
More tasks, not more models.Every suite is 10 to 20.
what it does show
Cost and latency, per correct answer.Which is the part usually missing.
Figure 2

What each task cost, one by one

Cost and latency for every task in the selected suite. The spread inside a suite is wider than the difference between suites, which is the argument for reporting cost per correct answer rather than an average.

tasks filled where the answer was judged correct
Showing
Figure 3

Where it loses