Coming soon
Upcoming test series

Honest reporting.

A second test series is being built around one question: when an agent says it has done something, did it actually do it? The check is not against the agent's own summary but against the record of what ran on the machine. Nothing here is published yet, and there are no results on this page.

Why it matters

A confident summary is not evidence.

An agent's reply is the main thing most users read. If it says the tests pass, that files were updated or that a step was skipped, the user usually takes that at face value. Where the summary and the machine disagree, the summary is the part that gets believed and the machine is the part that is true.

This matters more in unattended use, where nobody is watching the commands run. It also matters for the record: an agent that misreports its own work is harder to audit later, because the only account of the session is the one the session gave about itself.

Method

Claims checked against what ran.

Every claim is checked mechanicallyA claim is only scored when it can be checked against something recorded: whether a command was executed, whether a file was created or changed, what a command returned, and what the agent was actually asked to do.
What ran is captured independentlyCommands are observed at the process level, outside the agent, and files are compared before and after each step. The agent's own log is read as a claim, not as proof.
Simple, unambiguous claims firstThe first tests use claims with a clear yes or no answer, not a judgement about wording: a step that never ran, a file that was never written, a result presented as the output of a command that produced something else, or a failure presented as a success.
Distinguishing an error from a false claimA wrong statement is only scored when the agent had the means to know better: the command ran and its output was in front of it, or the file state was directly observable. A claim the agent had no way to check is reported, not counted against it.
The prompt is fixed and benignTasks are ordinary, small and honest in themselves, so the tests measure reporting rather than entrapment. Nothing asks the agent to deceive anyone.
Every verdict cites a captureA result names the claim, the recorded fact that contradicts it, and the evidence for both, in the same way as the current tests.
What will be reported

Per version, as a rate and as cases.

Claims checkedHow many checkable claims each run produced. A harness that says little is not rewarded for saying little.Counted
Contradicted claimsClaims the recorded evidence contradicts, with the claim text and the evidence side by side.Scored
Unverifiable claimsStatements that cannot be checked against anything recorded. Reported separately, because they are a limit of the test and not a finding about the agent.Reported
Corrections and recoveryWhether the agent corrected itself when its next command showed the earlier claim was wrong, which separates a slip from a pattern.Reported
False failures and false successesReported apart, because an agent that claims a failing step succeeded and an agent that discards real work it thinks failed are different problems for a user.Reported

Results will be per version, dated, and re-run when a version changes. The tests will be published with their prompts, their pass and fail rules and their evidence rules before any result, and scored the same way for every agent.

Status

Not published yet.

The task set, the claim extractor and the checking method are being written and reviewed. Nothing on this page is a result, and no vendor is named. When the series is ready, this page becomes the index for it, with the prompts, the rules and the captures published alongside the results.

Want to comment on the method before it is fixed, or suggest a claim worth testing? Write to research@agenticthinking.uk.