An evaluation harness for agents that treats token economics as a first-class metric. A model that scores five points higher and burns fifty times the tokens is not obviously the right choice, so every number here arrives next to its cost.
This is the first reference run, and it is one model over four suites. Reported with the interval a run this size actually supports, which is the thing an accuracy figure alone hides.
Each suite is between ten and twenty tasks. At ten, one task is ten points, so the Wilson interval is drawn next to the score rather than left out of it. A run that reports 1.00 has not shown you a perfect model; it has shown you no failures in ten tries.
Summarization scored 10 out of 10. The 95% interval on that runs from 0.72 to 1.00. Tool use scored 15 out of 15 and its lower bound is 0.80. Both are good results and neither one distinguishes a model that is right 99% of the time from one that is right three times in four.
Nothing about that is a criticism of the harness. It is what a first reference run costs, and printing the bar makes the size of the run impossible to forget.
Cost and latency for every task in the selected suite. The spread inside a suite is wider than the difference between suites, which is the argument for reporting cost per correct answer rather than an average.