Operate the AI service

Review the live service as a business and technical system: outcome, adoption, reliability, safety, cost, change, and retirement.

Self-directed lesson

Field question
Is the service still valuable, safe, affordable, supportable, and owned?
Working artifact
Production service review
Effort
90 minutes

You will learn

  • How business outcome, adoption, model behavior, system reliability, safety, and cost form one service review.
  • How incidents, drift, capability changes, and model upgrades trigger controlled decisions.
  • How continuation, constraint, scaling, and retirement remain explicit choices.

Field practice

  1. Review outcome, adoption, reliability, safety, cost, and unresolved change together.
  2. Trace one incident or regression through detection, containment, recovery, and learning.
  3. Record continue, constrain, scale, or retire with owner and next review date.

Before you begin

Bring current outcome, adoption, reliability, safety, cost, incident, change, and ownership evidence for one service period, with its evidence mode and limitations explicit.

Learning objectives

  • Review the AI service as a business and technical system rather than a model dashboard.
  • Turn incidents, drift, cost, and adoption evidence into continue, constrain, scale, or retire decisions.
  • Keep model, software, workflow, and business signals separately observable.

Core lesson

What you need to know

01

Operate the outcome, not just the model

A production service review begins with the accepted outcome, eligible population, baseline, target, adoption indicator, operating owner, metric owner, verifier, and attribution limits. Pair them with reliability, safety, latency, cost, human-review load, exception rate, and change history. Good model quality cannot compensate for non-adoption or negative economics.

Define service objectives that reflect user and business impact as well as infrastructure. Monitor the decision route: source availability, context quality, model behavior, tool behavior, human review, effect result, and source-of-truth readback. This makes it possible to localize failure without treating every incident as model drift.

02

Turn change into a controlled decision

Model upgrades, prompt changes, new data sources, policy revisions, tool versions, and workflow changes can invalidate evidence. Use the typed dependency graph to identify affected evaluations, artifacts, and operating controls, then rerun the relevant checks. The graph routes review; owners and gates still decide.

Review full cost and capacity, including review queues, support, incidents, regressions, and change work. Define cost ceilings and saturation signals before growth. A service that saves task time but creates excessive review or support burden may need constraint or retirement.

03

Preserve negative and retirement evidence

Record continue, constrain, scale, pause, or retire with evidence, owner, limitations, and next review. Preserve stopped experiments and retired behavior so the organization does not repeat the same failed assumption after staff or model changes.

Worked field case

Pilot service review after four weeks

The rec-0.4.3 NorthLake pilot has useful citation behavior, but adoption, cycle time, and operating evidence miss continuation conditions.

Evidence available

  • Citation acceptance is 92% after adjudication; the small cohort does not establish a correction-error trend.
  • Only 46% of eligible cases use recommendations; specialists report duplicate evidence entry.
  • Median time improves 14.3%, below the 15% continuation floor, while unresolved support cost keeps full economics incomplete.

Reasoning path

  1. Separate model quality from workflow adoption and integration burden.
  2. Trace duplicate entry to the workbench integration rather than retraining the model.
  3. Record constrain: hold volume, fix evidence handoff, remeasure adoption and cycle time in two weeks.
Result

The team constrains the pilot rather than turning promising citation evidence into a broad quality or value claim, then targets the evidence-handoff bottleneck.

Practice exercise

Run a service review

Use one period of live field evidence, a sanitized retrospective, or synthetic practice and review outcome, adoption, reliability, safety, cost, change, and ownership together.

  1. Compare accepted outcome and guardrails with current measured evidence.
  2. Trace one incident or regression across model, software, workflow, and human layers.
  3. Assess cost, queue capacity, support burden, and recent changes.
  4. Record continue, constrain, scale, pause, or retire with next review conditions.
Keep

A production service review that supports an operating decision rather than a status report.

Review your work

Field rubric

  • Outcome, adoption, reliability, safety, and cost are reviewed together.
  • Model behavior is not used as a proxy for workflow value.
  • Change impact is traced to affected evidence and controls.
  • The decision has an owner, trigger, and next review date.

Complete when

  • One operating decision follows from current evidence.
  • The highest-impact incident or constraint has a traced root cause.
  • Continuation and retirement are both viable outcomes.

Check your understanding

Could a good model metric hide poor adoption, unsafe effects, or negative operating value?

Working template

Download the mission artifact, complete it with source evidence, and review it against the rubric above.

Download Production service review template

Guided study and progress

Open this mission in the guided school to confirm the exercise and rubric, save completion in this browser, and resume the ten-mission path.

Open guided mission

Compare and go deeper

This lesson is self-contained. Compare your work with the calibrated answer, then use the attributed references when you need governed detail or additional implementation practice.

Annotated answerFDE Guide reference ↗Attributed technical depth ↗