What counts as success?

Success is the outcome you wanted, subject to the constraints you set. Anthropic’s agent evaluation guide distinguishes the agent’s conversation from the resulting state of its environment.

For a research agent, success might mean a sourced draft saved in the correct folder. For an inventory agent, it might mean a verified stock report. Define the evidence you will inspect before running the task.

A draft-publishing agent scorecard

This is an AiLookout test template for an agent that prepares articles for review. It describes a hypothetical workflow, not a system we have benchmarked.

Evaluate a draft-producing agent
DimensionPassing evidenceFailure example
Task completionRequired draft exists and opensAgent says done, but no file exists
Source accuracyEach factual claim has supporting evidenceLink exists but contradicts the draft
PermissionsOnly authorized draft files changedAgent publishes without approval
RecoveryMissing source is flagged clearlyAgent invents a replacement claim
EfficiencyRecorded runtime and usage fit the taskRepeated retries exceed the budget

Which cases should you include?

Use tasks resembling the real workflow, with an expected result you can inspect. Include a source that is unavailable, contradictory instructions inside a document, and an incomplete input. These cases reveal behavior that a clean demonstration can miss.

  • Ordinary task with sufficient information.
  • Missing field that requires a question or a flagged draft.
  • Tool error that requires stopping or a supported recovery.
  • Untrusted document text that asks for an unauthorized action.
  • Repeated trial with the same success criteria.

How do you track improvement?

Keep the inputs, agent configuration, tool access, outputs, and outcome checks for every trial. Review failures by category rather than collapsing them into one score.

Use deterministic checks for files and required fields, and human review where meaning matters. A model grader can assist, but validate whether it notices the errors you care about.

Common questions

Is one successful demo enough?

No. Repeat representative tasks and inspect the outcomes. A demonstration shows that a task can succeed, not how consistently it succeeds.

Should I score the agent’s reasoning?

Prioritize the observable result and permitted actions. Logs can help diagnose failures, but plausible explanations do not prove completion.

THE TAKEAWAY

What to remember

Write the success criteria first. Then inspect the actual result, authorized actions, and repeated failure patterns.

Sources & further reading

  1. Anthropic: demystifying evals for AI agents ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories