Sources checked
What counts as success?
Success is the outcome you wanted, subject to the constraints you set. Anthropic’s agent evaluation guide distinguishes the agent’s conversation from the resulting state of its environment.
For a research agent, success might mean a sourced draft saved in the correct folder. For an inventory agent, it might mean a verified stock report. Define the evidence you will inspect before running the task.
A draft-publishing agent scorecard
This is an AiLookout test template for an agent that prepares articles for review. It describes a hypothetical workflow, not a system we have benchmarked.
| Dimension | Passing evidence | Failure example |
|---|---|---|
| Task completion | Required draft exists and opens | Agent says done, but no file exists |
| Source accuracy | Each factual claim has supporting evidence | Link exists but contradicts the draft |
| Permissions | Only authorized draft files changed | Agent publishes without approval |
| Recovery | Missing source is flagged clearly | Agent invents a replacement claim |
| Efficiency | Recorded runtime and usage fit the task | Repeated retries exceed the budget |
Which cases should you include?
Use tasks resembling the real workflow, with an expected result you can inspect. Include a source that is unavailable, contradictory instructions inside a document, and an incomplete input. These cases reveal behavior that a clean demonstration can miss.
- Ordinary task with sufficient information.
- Missing field that requires a question or a flagged draft.
- Tool error that requires stopping or a supported recovery.
- Untrusted document text that asks for an unauthorized action.
- Repeated trial with the same success criteria.
How do you track improvement?
Keep the inputs, agent configuration, tool access, outputs, and outcome checks for every trial. Review failures by category rather than collapsing them into one score.
Use deterministic checks for files and required fields, and human review where meaning matters. A model grader can assist, but validate whether it notices the errors you care about.
Common questions
Is one successful demo enough?
No. Repeat representative tasks and inspect the outcomes. A demonstration shows that a task can succeed, not how consistently it succeeds.
Should I score the agent’s reasoning?
Prioritize the observable result and permitted actions. Logs can help diagnose failures, but plausible explanations do not prove completion.
What to remember
Write the success criteria first. Then inspect the actual result, authorized actions, and repeated failure patterns.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





