Governance charter
Governance charter.
Governance charter
How the benchmark is run.
Independence.Agentic Thinking is a research lab. It has developed AI agent governance and workflow software, none of which is for sale, and we will never score our own software on AgenticBench. The benchmark takes no vendor money. Because the lab has developed software in this space, tamper evidence is reported but never scored.
How tests are setCategories, tests and thresholds are published with their reasons before results, versioned, and applied to every agent at once. Where a vendor documents a behaviour, we score against the documentation.
No vendor moneyNobody pays to be included, excluded, tested or rated. No sponsorship or advertising from vendors of tested agents.
No paid badges, everThere will never be a paid certification, seal or "verified" badge. Results are free and the same for everyone.
Vendor replies verbatimFindings go to the vendor first. Their reply is published word for word with the result, unless they ask us not to quote them.
Disputes and re-testsA vendor can dispute a result. We re-test with any documented configuration they ask for, publish the outcome, and mark the result as disputed until it is resolved.
Corrections in the openMistakes are corrected on the page where they appeared, with the date and what changed. Owned errors in method are recorded.
Evidence firstEvery pass or fail cites something we captured ourselves. Results stay labelled preliminary until they have been reproduced independently.