Efficiency leaderboard

Same model, different harness.

Efficiency leaderboard: first results being verified. Results will be published after verification.

What it measures

The cost of the harness, with the model held constant.

Each coding harness gets the same task and the same model. We record what each run costs and how much work the harness does around the model: tokens, model calls, turns and wall time. A bare-model run is the baseline.

Status: runs in progress, published after verification. Questions: research@agenticthinking.uk.