Writing · Applied AI

Evaluate Agents Against Historical Work, Not Demo Prompts

How to turn historical work into leakage-aware agent evaluations that test decisions, tool use, escalation and recovery as well as the final answer.

Historical-work evaluation

An evaluation reconstructed from representative work. It preserves the information, tools, permissions and decision sequence available at the time while withholding the eventual answer.

Most agent demonstrations use a clean prompt, a named objective and a tool that happens to return the needed field. They usually stop before anyone has to operate the result. By that point, the difficult work has already been removed.

This is useful for showing an interface. It is weak evidence that an agent can perform real work.

Operational work contains ambiguity, missing state, conflicting evidence, permission boundaries and consequences that appear several steps after the first response. A capable agent must decide what to inspect, what not to assume, when to ask, when to act and when to stop. Those decisions are precisely what a polished prompt tends to erase.

If the purpose of evaluation is deployment judgment, the test material should begin with the work as it was—not with a cleaned-up description written after the answer was known.

Demo prompts flatten the work

A resolved case creates hindsight. Once the cause is known, every useful clue looks more obvious than it did in sequence. The final explanation compresses false starts, missing information and judgment calls into a coherent story.

Turning that story into a prompt introduces three distortions:

  1. Outcome leakage: the wording points toward the known resolution.
  2. Context completion: missing or inconvenient information is added for clarity.
  3. Trajectory removal: the evaluator grades the answer but not how the system arrived there.

These distortions make an agent look more capable than the operating environment will allow. They also make the evaluation less useful to the product team. A failure on a clean prompt says little about whether the problem is retrieval, investigation, tool design, permissions, observability or escalation.

The first principle is therefore simple: preserve the shape of the work before preserving the wording of the record.

The unit of evaluation is a trajectory

For multi-step work, the final answer is only one artifact. The more important unit is the trajectory: the sequence of observations, questions, hypotheses, tool calls, decisions and state changes that produced it.

Two systems can reach the same conclusion through very different operating behavior. One may use the available evidence, make a bounded observation and escalate at the correct boundary. Another may guess correctly, access information outside the case scope and take a risky action that happens to succeed.

Outcome-only grading treats those systems as equivalent. They are not equivalent in production.

What a trajectory-level evaluation preserves
LayerWhat to captureWhy it matters
InformationWhat was available at each decision pointPrevents later evidence from leaking into earlier judgment.
Reasoning stateClaims, uncertainty and active alternativesShows whether the system distinguishes evidence from inference.
InteractionQuestions, tool calls and responsesReveals burden, information value and tool discipline.
AuthorityPermissions and state-changing boundariesSeparates diagnostic capability from authorization to act.
HandoffEscalation trigger and transferred contextTests whether a person can continue without rebuilding the case.

The evaluation does not need access to a model’s private chain of thought. It needs an observable operating record: what the system claimed, what evidence it cited, which action it requested, what the tool returned and how the next decision changed.

Reconstruct the decision-time horizon

The central discipline is to recreate what could have been known at the time.

Imagine a hypothetical support case in which a connected device stops reporting. The later record may contain a technician’s inspection, a replacement event and a final diagnosis. Showing any of those facts in the opening prompt turns an investigation into retrieval.

Instead, define an event horizon for each step:

  • the initial report and its original ambiguity;
  • product state that was genuinely observable then;
  • tools and documentation available to the role;
  • answers that become available only after a targeted question;
  • later evidence released only when the trajectory earns it.

The phrase “earns it” matters. An agent should not receive a decisive observation merely because the script reached step three. The observation should follow from a relevant question or permitted tool action. Otherwise the evaluation still rewards waiting rather than investigating.

Replaying the final record is therefore not a historical evaluation. The test has to reconstruct the information boundary that existed before the outcome was known.

Build four evaluation artifacts

A reusable case is easier to review when it is separated into four artifacts.

1. The case envelope

The case envelope defines the task, intended outcome, role, starting information, tools, permissions and stopping conditions. It also records why the case belongs in the evaluation set: routine work, ambiguity, a known boundary, an escalation decision or a particular failure mode.

This is where scope becomes explicit. “Resolve the case” is rarely precise enough. A system may be expected to diagnose, recommend, execute a reversible change, prepare a handoff or some combination of those—but not necessarily all of them.

2. The event ledger

The event ledger is the ordered source of truth for observations that can become available. Each event needs provenance, time context, release conditions and any known limitations.

The ledger prevents the evaluation harness from improvising facts. It also exposes gaps in the surrounding product. If a decisive state cannot be represented with provenance, the organization may have an observability problem before it has an agent problem.

3. The reference analysis

The reference analysis records acceptable paths, important alternatives, unsafe moves and the evidence that changes the decision. It is not a single ideal transcript.

Real work often permits several competent sequences. The reference should therefore define decision invariants—what must be noticed, what must not be assumed and where authority changes—while allowing different competent paths.

4. The grading contract

The grading contract defines the dimensions, severity weights, automatic checks, human review rules and deployment thresholds before a model is tested.

This prevents criteria from drifting toward the behavior of whichever system is currently under discussion. It also makes disagreement useful: reviewers can challenge the contract rather than arguing from impressions after a demo.

Use real work without publishing private work

Historical work is valuable because it contains operating reality. It can also contain personal data, customer records, security detail, contractual information and recognizable incidents. Evaluation design must begin with permission and minimization, not with exporting everything into a new tool.

A safe preparation process should answer:

  • Is this record permitted for this evaluation purpose?
  • Which fields are necessary to preserve the decision problem?
  • Can identifiers and irrelevant details be removed without changing the work?
  • Could the remaining sequence still identify a person, customer or incident?
  • Where will source material, model inputs and outputs be stored?
  • Who may inspect failed runs?
  • When will the material be deleted or refreshed?

An anonymized record is not automatically safe, and a synthetic rewrite is not automatically representative. The aim is to preserve the operating constraints while removing information the evaluation does not need.

Public writing should go further. It can explain the method, schema and failure classes without reproducing any source record. A hypothetical example can illustrate the framework; it should never be presented as a disguised customer story.

Sample representative work

An evaluation set made only from memorable failures will exaggerate the tail. A set made only from clean, successfully resolved work will understate it. Expert-curated examples can also become a test of expert taste rather than normal operating performance.

Build the set in strata:

  • routine cases where consistency and efficiency matter;
  • ambiguous cases where question quality matters;
  • boundary cases where escalation or authorization matters;
  • rare but consequential cases where severity dominates frequency;
  • cases with incomplete or contradictory evidence;
  • cases that were unresolved, reopened or handed between roles.

Record the sampling logic. The evaluation score is not meaningful without knowing what the set represents.

The same applies to human comparison. Match the case, information and role as closely as possible. Comparing an agent with a retrospective expert answer is not the same as comparing it with the people who perform the work under normal constraints.

Grade operating behavior by consequence

Average answer quality is too coarse for an agent that can ask questions, use tools or change state. I prefer six evaluation dimensions:

A trajectory-level grading contract
DimensionReview questionExample failure
Diagnostic progressDid each step reduce relevant uncertainty?Repeated activity without separating active hypotheses.
Evidence disciplineWere report, observation and inference kept distinct?A plausible assumption was presented as observed state.
Question qualityWas the question answerable and worth the burden?A broad request replaced a targeted discriminating question.
Tool safetyDid scope, permission and effect remain controlled?A diagnostic path crossed into an unauthorized change.
Boundary judgmentDid the system stop or escalate at the right point?High-consequence uncertainty was hidden by confident language.
Handoff qualityCould the receiving person continue efficiently?The transcript was transferred without a decision-relevant case state.

These dimensions should not be averaged blindly. A severe permission failure can outweigh many well-phrased routine resolutions. Deployment thresholds should reflect consequence, detectability, recoverability, propagation and frequency.

Separate model, system and workflow failures

When an agent fails, “the model got it wrong” is often an incomplete diagnosis.

The necessary evidence may not exist. Retrieval may surface an obsolete document. A tool may collapse three physical states into one field. Permissions may be too broad. The prompt may reward premature certainty. The human handoff may have no receiving owner.

Tag failures at the layer where a corrective action can be taken:

  • source knowledge;
  • product observability;
  • retrieval and context assembly;
  • model judgment;
  • tool contract;
  • authorization;
  • interaction design;
  • escalation workflow;
  • evaluation harness.

This classification turns evaluation into product development. Without it, the team either changes the prompt for every failure or concludes that the model is not ready, while the same system weakness remains.

Version the evaluation with the product

An evaluation set expires. Policies change. Tools gain authority. Documentation improves. Products add states. Models and orchestration change. A case that was representative last year may reward obsolete behavior now.

Each evaluation release should record:

  • the workflow and product boundary it represents;
  • source-data eligibility and review date;
  • tool and permission versions;
  • model and system configuration;
  • case-set changes and the reason for them;
  • known coverage gaps;
  • the trigger for the next review.

Keep a stable holdout for comparability, but do not protect a benchmark at the expense of relevance. The benchmark exists to support a sound deployment decision in a changing operating system.

Evaluate the work you intend to delegate

Historical-work evaluation is slow to prepare because it refuses the clarity of hindsight. That inconvenience is the point.

It shows whether the agent can operate with the information that actually exists, whether its tools expose the right evidence, whether permissions match consequence and whether a person can take over without starting again. It also reveals where the current human workflow depends on unwritten knowledge or compensates for weak product observability.

Keep demo prompts for communication. Base release decisions on reconstructed work that preserves the event horizon and exposes the full trajectory. Only then is there evidence for how much authority the system has earned.

Sources and further reading

  1. NIST, AI RMF Core, includes outcomes for representative data selection, contextual validation, documented human oversight and ongoing measurement.
  2. OpenAI, “Measuring the performance of our models on real-world tasks”, describes GDPval’s use of work-like tasks, expert-authored rubrics and blinded expert comparison, while noting that one-shot evaluation does not capture interactive workflows.
  3. Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, studies the interleaving of reasoning and actions in tasks that require information from external environments.