Working template
Evaluation Set and Report
Evaluation Set and Report, a learner-facing working resource from the free and open-source Lumyn FDE School.
For self-directed implementation practice, use the fictional claims-denial case and optional technical labs.
Evidence boundary
- Evidence mode: live field evidence / sanitized retrospective / synthetic practice:
- Evidence classes used: assertion / observation / policy / measured operations / exact-build test / exercised capability:
- Intended use or decision:
- Scope and applicability:
- Current limitations and prohibited claims:
Release decision
- Decision this evaluation informs:
- Exact candidate binding: code, configuration, data, model route, prompt or policy, tools, graders:
- Protected effects:
- Reference authority and independent approver:
- Adjudication path:
- Thresholds defined before execution:
Case registry
| Case ID | Class | Workflow slice | Source revision | Expected result basis | Protected failure | Review status |
|---|---|---|---|---|---|---|
| routine / boundary / adversarial / recovery |
Evidence lanes
| Lane | Metric or contract | Threshold | Result | Evidence | Limitation |
|---|---|---|---|---|---|
| Deterministic contracts | |||||
| Behavioral quality | |||||
| Safety and authorization | |||||
| Latency | |||||
| Full cost | |||||
| Human review and adoption |
Do not average protected failures into an aggregate pass.
False-green checks
- Broken or mismatched citation:
- Stale approval:
- Duplicate effect:
- Cross-tenant source:
- Failed readback or rollback:
- Fixture or grader contamination:
Record whether each mutation makes the suite fail as intended.
Result
- Candidate result: supports / does not support / inconclusive:
- Supported scope:
- Failed and borderline cases:
- Variance and repeat behavior:
- Cost and latency:
- Known limitations:
- Changes that require rerun:
- Reviewer and date:
Evaluation evidence supports a decision; it never grants release, customer acceptance, or effect authority.