Honest reporting.
A second test series is being built around one question: when an agent says it has done something, did it actually do it? The check is not against the agent's own summary but against the record of what ran on the machine. Nothing here is published yet, and there are no results on this page.
A confident summary is not evidence.
An agent's reply is the main thing most users read. If it says the tests pass, that files were updated or that a step was skipped, the user usually takes that at face value. Where the summary and the machine disagree, the summary is the part that gets believed and the machine is the part that is true.
This matters more in unattended use, where nobody is watching the commands run. It also matters for the record: an agent that misreports its own work is harder to audit later, because the only account of the session is the one the session gave about itself.
Claims checked against what ran.
Per version, as a rate and as cases.
Results will be per version, dated, and re-run when a version changes. The tests will be published with their prompts, their pass and fail rules and their evidence rules before any result, and scored the same way for every agent.
Not published yet.
The task set, the claim extractor and the checking method are being written and reviewed. Nothing on this page is a result, and no vendor is named. When the series is ready, this page becomes the index for it, with the prompts, the rules and the captures published alongside the results.
Want to comment on the method before it is fixed, or suggest a claim worth testing? Write to research@agenticthinking.uk.