Skip to content
Context Windowby Alex Janjic
Menu
contextwindow.us/posts/the-ai-operating-model-in-practice-part-4Article

The AI Operating Model in Practice - Treat every agent behavior change as a release - Part 4

A prompt change can send more refunds to specialists without changing code. Decide its release using tested actions, qualified reviewer time and someone who can stop new admissions.

AI systems
Source links stay with the relevant section
Two groups of blank refund cards approach one review tray with a yellow stop tab at its entrance.

A better score can still delay refund decisions

In a fictional refund service, a prompt edit sends more cases to specialist review. The assistant may avoid weak automatic decisions. Its quality score could rise while customers wait longer for a decision.

Treat that prompt edit as an agent behavior release: a change to what an assistant does in production. The code can stay the same. The release decision must include the people and services that handle the changed work.

A tool, permission, model setting or retrieval setting can also change an action. Retrieval means information the assistant looks up before acting. Record these parts together even when separate systems deliver them.

One release set, several people responsible

For the fictional refund assistant, name the old set Refund Baseline Birch and the proposed set Refund Candidate Cedar. Record exact references for application code, prompt, provider and model settings, tools, access policy and retrieval configuration. Include skills, rules and examples when they affect behavior.

Version the case set used for testing and the result for each release separately. A retrieval setting is not the same as the business records it reads. Record which policy snapshot a case used and check live order and payment records before acting. A changing provider or business record limits exact replay.

Nila, the proposal author, states the intended routing change. Sol accepts or rejects the evidence. Dev decides exposure for domain engineering. Mara can stop new admissions when review coverage fails. The work should have named owners without creating a new committee for every prompt edit.

Decide exposure from tests and actual coverage

Name release setNila records intent and revisions
Compare casesSol accepts action and outcome evidence
Check coverageMara checks reviewer slots and deadlines
Admit covered casesDev sets bounded exposure
Watch outcomesRoutes, queue, writes and spending
Hold new casesMara stops intake; Ivo resolves writes
  1. Name release settestCompare cases
  2. Compare casesacceptCheck coverage
  3. Check coveragecoveredAdmit covered cases
  4. Admit covered casesobserveWatch outcomes
  5. Watch outcomeslimit failsHold new cases
Read this diagram as text

Nila records the intended change and the release set. Sol compares affected and useful-work cases, then accepts the evidence. Mara checks eligible reviewer slots and case deadlines. Dev admits only a bounded group. The team watches routes, queues, writes and spending. If a limit fails, Mara stops new candidate admissions and Ivo resolves uncertain writes. Started cases retain their release identity and get an explicit safety decision.

A failed limit blocks new candidate admissions. Started cases keep their release identity and get an explicit safety decision.

Build, test, deploy and monitor the same set

Build starts with the product outcome. Refund staff say which cases need specialist judgment. Product and engineering choose the prompt, tool definitions, permissions and policy sources that produce that outcome. The shared platform can supply common runtime and observation tools.

Test compares baseline and candidate on affected cases against stated acceptance rules. Keep an allowed-useful-work case that must finish, so blocking every action cannot count as success. Sol checks proposed actions, attempted service calls and stored refund state before accepting results.

Part 3, Turn agent failures into release-blocking tests, shows this distinction using synthetic scripted cases. Those checks can reject incorrect actions and a block-all control while allowing valid work. They do not measure live model quality, settlement or the time available to reviewers.

Deploy assigns the full release set to a named environment. Dev checks that the runtime can retain needed workflow state, request human review and handle failure before limiting exposure. A saved checkpoint can support later continuation. It cannot reverse an external refund. Isolation matters if a workflow executes code or handles files; this refund example does neither.

Monitor ties an input, model call, tool attempt and output to one case. That linked record is a trace. Join traces to policy checks, reviewer corrections, escalation and abandonment. Also watch overall latency, usage, spending and failures. A trace records an attempt, not proof that a refund completed.

These are delivery stages, not the refund service's lookup, review and payment steps. Access rules, spending limits and making approved components findable need care in every stage. Improvement also runs through all four. Domain engineering owns rollout, the platform supports it and refund specialists reveal bad decisions.

A passing test cannot book a reviewer's time

A passing test cannot establish qualified reviewer time, safe service permissions or an affordable operation. When routing changes, check the existing queue, reviewer skills, protected time, deadlines and bursts before admitting cases. Make human review capacity an AI rollout gate explains the need for an eligible reviewer and a real availability window for each case.

An average route rate helps plan work. It never guarantees a slot for the next case. Monitoring and investigation also need protected time. Otherwise the same people asked to review refunds must find spare time to notice and diagnose the release's failures.

Ready for covered cases, wait without coverage

Every number in this example is an illustrative planning assumption, not an observed result. Suppose 120 refund cases arrive per working day. Baseline routes 10% to specialists, or 12 cases. At full exposure, candidate routing of 25% means 30 specialist reviews.

At eight minutes per review, full candidate exposure needs 240 specialist minutes per day before other duties. Baseline needs 96 minutes. A better quality score cannot supply the extra 144 minutes or extend a customer's deadline.

Instead, 24 candidate cases replace 24 of the 120 daily cases. The other 96 stay on baseline. Candidate demand averages six reviews or 48 minutes. Baseline demand averages 9.6 reviews or 76.8 minutes. Together, that is 124.8 expected minutes per working day.

For the illustrative ready start, existing backlog is zero. Mara has 160 protected qualified specialist minutes daily across both release groups, excluding other duties. The planning difference is 35.2 minutes. That headroom is not guaranteed capacity. Real reviews are whole cases and arrivals can cluster.

The service is staffed from 09:00 to 17:00. It aims to decide each accepted case within the same working day. Mara assigns named eligible reviewer slots before each admission and confirms separate standard-review coverage wherever policy requires it. Late cases without a suitable slot wait under the service's allowed intake rule.

Mara checks the queue and reviewer slots before every admission and every 30 minutes while staffed. Four working hours for the oldest case is an early schedule warning, not permission to miss its original deadline. Two successive increases in unfinished specialist cases stop new candidate admissions until Mara checks the schedule.

The ready decision uses the assumed safety evidence and admits only covered cases within the 24-case daily limit. In the wait decision, the same evidence applies but protected qualified time is unconfirmed. Dev keeps candidate admissions closed. Neither novice nor unavailable staff count as coverage.

Illustrative expected specialist minutes per working day. Candidate cases replace baseline cases. Every admitted case still needs an eligible slot before its own deadline.
MeasureDecisionExpected review workQualified coverageNext action
Ready: at most 24 candidate casesReady: at most 24 candidate cases48 candidate plus 76.8 baseline = 124.8 minutes160 protected minutes daily, plus case slotsAdmit only cases with deadline coverage
Wait: no candidate admissionsWait: no candidate admissionsSame assumed routing and test evidenceProtected eligible time unconfirmedKeep candidate closed and check baseline capacity

Copy one record for the next change

This record makes a release decision readable after the prompt author leaves. Use immutable references for changed and unchanged components in both named sets. The filled example uses fictional labels and an illustrative decision, not a production release.

Reusable agent behavior release record

Release identity: Name baseline and candidate sets with immutable references. Intent and proposal author: State the change, expected work and accountable person. Component revisions: Enter baseline and candidate immutable references separately for application code, prompt, provider identity and deployment configuration, tool definitions, access policy, retrieval configuration, skills, rules and examples. Retrieval sources: Record policy snapshot, case snapshot rule and live business records checked at execution. Replay limits: Name provider or source changes that may prevent exact replay. Evaluation cases: Enter separate baseline and candidate case set references. Evaluation results: Enter separate baseline and candidate result references, affected cases, useful-work control, acceptance thresholds, failures and limits. Evidence acceptor: Name the person who checks relevance and accepts or rejects evidence. Unresolved risks: Give each risk an owner and a condition for action. Environment and exposure: Name target, daily admission limit and initial observation window. Review coverage: Record eligible reviewer slots, working hours, deadlines, backlog and standard-review coverage. Investigation coverage: Record protected calibration, monitoring and diagnosis time separately from case processing. Observation cadence: Name who checks routes, queue, outcomes and errors and when. Budget owner and scope: State currency, period, covered costs, ceiling and hold rule. Keep staffing time separate. Rollout owner and expansion: Name who authorizes rollout and the evidence and capacity needed for a new decision. Stop authority and conditions: Name who stops new admissions or unstarted actions and who handles already accepted work. Rollback target and in-flight rule: Check target safety and capacity. Decide how each started case keeps, pauses or migrates its identity and permissions. Reconciliation and verification: Name authoritative business records and owners for queued, paused, completed and uncertain actions.

Filled fictional refund release record

Release: Refund Baseline Birch to Refund Candidate Cedar. All component labels below are fictional immutable internal references. Intent and author: Nila, refund product engineer, proposes specialist routing for uncertain cases while safe automatic work still completes. Code baseline | candidate: app-birch-A | app-birch-A. Prompt baseline | candidate: prompt-birch-B | prompt-cedar-C. Provider and model deployment baseline | candidate: internal-provider-north with model-birch-D | internal-provider-north with model-birch-D. Tools baseline | candidate: tools-birch-E | tools-birch-E. Access policy baseline | candidate: policy-birch-F | policy-birch-F. Retrieval configuration baseline | candidate: retrieval-birch-G | retrieval-birch-G. Skills, rules and examples baseline | candidate: bundle-birch-H | bundle-birch-H. Retrieval source: Both sets use approved refund-policy snapshot policy-birch-snapshot. Each case records that snapshot ID. Live order and payment records are checked at execution. Changing business records or outside provider behavior limit exact replay. Evaluation cases baseline | candidate: cases-birch-J | cases-cedar-J. Illustrative evaluation result references, baseline | candidate: eval-birch-K | eval-cedar-K. Illustrative acceptance rule: Zero forbidden action attempts; correct stored outcome on every affected scripted case; one allowed-useful-work control completes rather than blocking everything. Sol checks proposed actions, service attempts and stored state. Assume these rules pass for the decision example. They do not establish live model quality or payment settlement. Evidence acceptor: Sol, evaluation lead, checks the comparison and records acceptance before any admission. Open risks: Mara owns clustered arrivals and late reviews. Ivo owns uncertain writes. Dev owns cost and outside provider drift. These risks stay open during observation. Environment and limit: Refund production, staffed 09:00 to 17:00. At most 24 candidate cases replace 24 baseline cases each working day for three staffed working days. No automatic expansion. Reviewer coverage: Mara has 160 protected qualified specialist minutes per working day for both groups, excluding other duties. Existing backlog starts at zero. She confirms named eligible slots, same-working-day deadlines and separate required standard-review coverage before admitting each case. A late arrival without a slot cannot enter. Investigation coverage: Sol has 30 protected minutes daily for sampling and calibration. Dev has 15 protected minutes daily for investigation. Operations confirms Mara's separate monitoring coverage. If any required coverage is missing, candidate admission waits. These are illustrative availability commitments, not staffing prices. Observation: Mara checks queue and slots before every admission and every 30 minutes during staffed hours. Daily she compares actual routing, decision deadlines and coverage. Sol samples decisions with qualified reviewers and checks forbidden attempts. Dev checks spending and authoritative completed and pending outcomes. Stop authority: Mara immediately stops affected new admissions when eligible coverage is lost, a deadline cannot be met, a permission breach occurs or a write remains unresolved. Two consecutive increases in unfinished specialist cases also trigger a hold and schedule check. Four working hours for the oldest case warns Mara to check its original deadline sooner. Baseline contingency: If baseline coverage also fails, Dev applies an allowed stop or deferral of new intake. The team still meets accepted deadlines and required approvals. Dev may cancel future unstarted candidate actions only after checking their safety; started business actions are not cancelled in bulk. Routing hold: More than six specialist-routed candidate cases among the day's 24 triggers a hold on expansion and a capacity reassessment. Small counts do not prove a stable rate. Budget: Dev owns an illustrative USD 20 per working day ceiling for model and runtime processing across the bounded workflow. Before further calls exceed it, Dev holds new candidate admissions and expansion. Reviewer and investigation time is a separate capacity need, not free within USD 20. Dev records any extra authorization needed to finish safely accepted cases. Rollout and expansion: Dev authorizes the initial bounded admission only after Sol accepts evidence and Mara confirms coverage. After three staffed working days, expansion requires updated evidence accepted by Sol, next-exposure capacity confirmed by Mara and Dev's recorded decision. Observation remains subject to the stop rules. Rollback and in-flight policy: Refund Baseline Birch is the new-case target only if its safety and capacity remain valid. Started cases retain their original identity. A reviewed compatibility and permission decision controls safe continuation, migration or suspension. Reconciliation: Ivo checks pending and uncertain writes against authoritative refund records before retry. Mara verifies queued and paused reviews. Dev checks routes and policy. Ivo verifies completed and uncertain business actions.

A stop does not undo a refund

Returning new cases to Baseline Birch does not erase a specialist decision or reverse an issued refund. Use that target only while its safety and reviewer capacity remain valid. Started cases retain their release identity but need not continue an unsafe release.

Review compatibility and permissions before resuming, migrating or safely suspending a started case. A saved workflow checkpoint can retain state and pending requests for continuation. It does not change the external business record by itself.

A timed-out write has an unknown result until an authoritative business record resolves it. Ivo checks that record before any retry. Treat timed-out agent writes as unknown, not failed explains why a timeout does not prove failure.

If baseline coverage fails as well, Dev cannot send unlimited cases back to it. Intake stops or waits under an allowed service rule. Mara checks queued and paused cases against accepted deadlines. Ivo checks completed actions and unresolved writes. Every accepted case still needs an owner.

Use actual cases to decide the next release

Classify each issue before changing behavior: poor decisions, unauthorized actions, too much review work, high processing cost or outside changes. Assign an owner to diagnose it. Add a reproducible case to the evaluation set, then compare a repair with baseline and the useful-work control.

Sol decides whether the new case and results address the issue. Dev authorizes the next exposure only after coverage and spending checks. Automated analysis may propose a diagnosis, patch or test. It cannot accept its own evidence or authorize its own deployment.

A changed provider, retrieved source or dependency can trigger targeted checks even when the team deploys nothing. Sample production decisions with qualified reviewers to compare judgments. Track review queues and time from first observed failure to verified fix. Reserve time for interpreting traces and investigating cases.

For the next change, put the behavior release set, evidence acceptance and operational stop authority in one record before expanding exposure.

Context Window dispatch

Practical AI engineering you can actually use

Working code, experiments, and production lessons. Published when there is something worth sending, never padded with AI news.

Confirmation is required. Read the privacy policy. Unsubscribe at any time.

The email edition is preparing to launch. No campaigns are being sent yet.