How AgenticBench is run.
Who runs the benchmark, how it stays independent, how agents are tested, how vendors hear about findings, how results are published, and the interests we declare. The detailed rules live on the main page; this page links to them.
An independent research lab.
AgenticBench is run by Agentic Thinking, an independent AI agent research lab in the UK. The lab's work has three parts.
No vendor money.
- Nobody pays to be included, excluded, tested or rated.
- There are no paid badges, seals or certifications, and there never will be.
- We never score our own software.
- We do not do bug bounty work and do not take bug bounty payouts.
The full rules, including disputes, re-tests and corrections, are in the governance charter.
Same tests, clean machine, every connection recorded.
- Each agent is installed fresh, at its latest version, from the vendor's own channel, in a throwaway container.
- Only fake secrets are planted, on our own machines and accounts. Every network connection is recorded.
- Every agent gets the same published tests.
- Every published result cites evidence we captured ourselves, and we publish that evidence. While a disclosure is open, affected results show only “Held” and the date, with no detail published, unless the early-publication exception below applies.
- Agents whose terms forbid automated testing or benchmarking are excluded, with the reason. We do not work around those terms.
The details are in the method and the threat model.
Vendors hear first.
- We disclose by email to the vendor's published security or privacy contact. We do not resubmit through web forms or third-party platforms. An automated reply redirecting us elsewhere counts as receipt and does not change our timeline.
- About 30 days for behaviour that differs from the vendor's documentation, privacy statement or settings. 90 days for security vulnerabilities. We extend the window when a fix is in progress, re-test fixed versions and publish both results.
- We publish sooner only if users are at immediate risk or a flaw is being exploited, and only after telling the vendor.
- The vendor's reply is published with the result, word for word, unless they ask us not to quote them.
- A vendor can dispute a result. We re-test with any documented configuration they ask for, publish the outcome, and mark the result as disputed until it is resolved.
- We do not sign agreements or accept platform terms that would restrict what we publish, and we do not accept payment from vendors in connection with a finding.
The full policy is in the disclosure policy.
A rolling leaderboard, dated corrections.
- Results are two rolling leaderboards, one row per harness version, with the date it was released and the date we tested it. Each harness version is tested once per leaderboard; failures are shown, never dropped.
- Clear errors are corrected on the page where it appeared, with a dated entry that shows the old value, the new value and why it changed.
- Changes to the method are dated.
- Results stay labelled preliminary until someone else has reproduced them independently.
We would like you to check our work: the public rig has what you need.
One researcher, with AI agent teams.
The lab is one researcher directing teams of AI agents. The agents help build the test rig, run the tests and review the results.
The researcher checks every published result and every outbound message, including every disclosure to a vendor, before it goes out. AI review is used in addition to human review, not instead of it.
Verified access for security research. Our security testing is carried out under Anthropic's Cyber Verification Program and OpenAI's Daybreak Blue trusted access. These verify our identity and our defensive use case.
What we own, so you can weigh our work.
- Agentic Thinking stewards the AgentHook standard and maintains HookBus. Both are Apache 2.0.
- Agentic Thinking has developed AI agent governance and workflow software: AgentProtect, AgenticStudio and HookBus Agent. None of it is sold or currently offered, and none of it is scored on AgenticBench. Because of this, tamper evidence is reported but never scored.
- AgentAuditor, the recorder used in our incident research, is ours and is not yet published.
- Agentic Thinking holds UK patent applications GB2608069.7 (HookBus) and GB2604445.3 (CRE).
Write to us.
- AgenticBench: research@agenticthinking.uk
- The lab: leo@agenticthinking.uk
- Security issues in our open-source code: security@agenticthinking.uk
The lab's own trust page is at agenticthinking.uk/security.html.