Build evaluation evidence

Create realistic, reproducible evidence for the exact behavior and release decision—without turning a passing score into false authority.

Self-directed lesson

Field question
Which cases, graders, thresholds, and limitations can support this exact release decision?
Working artifact
Evaluation set + report
Effort
2–3 hours

You will learn

  • Why reference authority, adjudication, and contamination controls matter.
  • How behavioral, deterministic, safety, cost, and latency evidence remain separate lanes.
  • Why evaluation supports but never authorizes release or customer acceptance.

Field practice

  1. Collect routine, boundary, adversarial, and recovery cases from the real workflow.
  2. Define independent grading, thresholds, failure promotion, and cost/latency budgets.
  3. Run the exact candidate and report both passing evidence and limitations.

Before you begin

Bring the exact behavior boundary, representative workflow cases, reference authority, and a candidate configuration.

Learning objectives

  • Build routine, boundary, adversarial, and recovery cases from the real workflow.
  • Separate deterministic, behavioral, safety, latency, cost, and human-review evidence.
  • Bind results and limitations to the exact candidate without turning a score into authority.

Core lesson

What you need to know

01

Design evidence for a decision

An evaluation exists to reduce uncertainty for a specific release or product decision. Define the decision, protected effects, candidate binding, reference authority, adjudication path, and thresholds before running it. Use realistic cases that cover routine work, boundaries, known exceptions, adversarial inputs, partial failure, and recovery.

Expected results need an evidence basis and review date. Do not treat anonymous labels or model-generated answers as ground truth. When experts disagree, preserve the disagreement and adjudication rather than averaging it away.

02

Keep evidence lanes separate

Schema validity, citations, permissions, duplicate safety, and readback are deterministic contracts. Judgment quality may need expert or model-assisted grading. Security tests exercise denial and containment. Latency and cost require measurement. Human review capacity and adoption require workflow evidence. A single aggregate score should not allow strength in one lane to hide a protected failure in another.

Use negative controls and false-green tests. Mutate a citation, stale an approval, duplicate an effect, cross a tenant, or break rollback and verify the suite fails. Otherwise the evaluation may be measuring fixture familiarity rather than the intended contract.

03

Report limitations and change triggers

Bind the report to model route, prompt or policy, tools, context, data revision, code build, configuration, and grader revision. State known blind spots, confidence, cost, and the changes that require rerun. Passing supports a release decision; it does not make that decision.

Worked field case

Correction recommendation release eval

The team must determine whether candidate rec-0.4.2 supplies enough evaluation evidence to proceed toward release review.

Evidence available

  • Two missing-note cases emit unsupported rationales, and five disputed corrections remain unadjudicated.
  • Duplicate recommendation fails; timeout recovery is inconclusive; rollback and queued-work behavior are unexercised.
  • Prompt-injection and cross-tenant tests pass within their stated limitations, but strength in those lanes cannot waive protected failures.

Reasoning path

  1. Run deterministic contract checks before behavioral grading.
  2. Preserve disputed labels and missing adjudication rather than averaging them into a quality score.
  3. Record which failures block rec-0.4.2 and which corrected evidence a successor candidate would need.
Result

The evaluation lane rejects rec-0.4.2 for pilot consideration. Mission 7 completes the security-boundary lane; only Mission 8 can make a release decision for a corrected exact candidate.

Practice exercise

Build the evaluation lane

Create a compact evaluation set for the vertical slice from Mission 5. This evidence informs—but does not complete—the release argument assembled in Mission 8.

  1. Define release decision, candidate binding, protected effects, reference authority, and thresholds.
  2. Add routine, boundary, adversarial, and recovery cases.
  3. Separate deterministic, behavioral, safety, cost, latency, and human-review results.
  4. Add at least two false-green mutations and document rerun triggers.
Keep

An evaluation set and report that a reviewer can reproduce and challenge.

Review your work

Field rubric

  • Cases represent the target workflow rather than generic benchmarks.
  • Protected failures cannot be averaged away.
  • Reference authority and adjudication are explicit.
  • Results bind to an exact candidate and state limitations.

Complete when

  • The suite fails when a protected contract is intentionally broken.
  • Cost, latency, variance, and human-review burden are reported.
  • No evaluation output claims release or customer authority.

Check your understanding

Would the evaluation fail if citations, authorization, idempotency, or recovery were broken?

Working template

Download the mission artifact, complete it with source evidence, and review it against the rubric above.

Download Evaluation set + report template

Guided study and progress

Open this mission in the guided school to confirm the exercise and rubric, save completion in this browser, and resume the ten-mission path.

Open guided mission

Compare and go deeper

This lesson is self-contained. Compare your work with the calibrated answer, then use the attributed references when you need governed detail or additional implementation practice.

Annotated answerFDE Guide reference ↗Attributed technical depth ↗