About

About AgenticBench.

What it is

An independent, open benchmark of AI coding agents.

AgenticBench tests AI coding agents, also called harnesses: the tools that wrap a model and let it read files, run commands and change code on your machine. It measures how the agent behaves, not how clever the model is. Every agent gets the same tests, and the test rig is public.

The two leaderboards

Security and efficiency.

SecurityWhat an agent does with your machine and your data: what leaves the machine and to whom, whether the opt-outs work, and whether it respects your controls, such as approval prompts and what it does when no human is present.
EfficiencyThe same model and the same task in every agent. We compare how each agent gets the task done and at what cost.

Each harness version is tested once per leaderboard; failures are shown, never dropped.

Why it matters

You are trusting the agent, not just the model.

A coding agent runs with your files, your keys and your network. For a developer, the question is what else happens besides the work you asked for, and what it costs. For a company, it is whether an agent can be allowed near source code and secrets, and whether its settings do what the documentation says. The vendor's description is a claim; AgenticBench checks it against what we captured.

How to read a row

One row per harness version.

Security columnsHarness, harness version, date released, date tested, tests passed / total. A missing safeguard counts as a fail. "Held" with a date means a result is held while the vendor is told first.
Efficiency columnsHarness, harness version, date released, date tested, completed (yes/no), cost, tokens, time. Completed says whether the agent finished the task. Cost, tokens and time are what the run used.

A result is for that version on that date. A newer version gets a new row when we test it. Every result cites evidence we captured ourselves.

Who runs it

An independent research lab.

AgenticBench is run by Agentic Thinking, an independent AI agent research lab in the UK.

  • Nobody pays to be included, excluded, tested or rated. The benchmark takes no vendor money.
  • We never score our own software.
  • Vendors hear about findings first, under our disclosure policy.

More on the rules and our declared interests: how we work, method and governance charter.

Contact

Write to us.