On this page
Learning objectives
- Classify actions by risk before the model can request them
- Persist a review packet that survives process restarts and delayed decisions
- Bind an approval to exact arguments, policy version, actor, and expiry
- Test injection, stale approval, duplicate resume, and reviewer-denial paths
Before you start
- • The typed tool executor from module 2
- • The authoritative run and effect records from module 3
- • A test identity for an operator and an authorized budget approver
Working definition
Human review
Human-in-the-loop control is a durable interruption in an application-owned state machine. The system evaluates a proposed action, records its exact arguments and policy context, exposes a decision to an authorized reviewer, and resumes only the same pending run. A guardrail is a narrow validation or policy check. It does not make the surrounding agent safe by itself.
A modal that asks whether an action looks acceptable is weak evidence. If the model can change arguments after approval, or if a retry can execute twice, the review is detached from the effect it was meant to govern.
Prompt injection is an untrusted-data problem as well as a prompt problem. Retrieved text, tool results, and user content must never acquire authority merely because a model repeats their instructions.
Review capacity is finite. A useful design routes low-risk reads automatically, blocks prohibited actions deterministically, and sends a concise evidence packet for the small set of actions where human judgment changes the risk decision.
Field situation
Beacon campaign budget review
Named synthetic scenario. Beacon Mobility, its campaign, identities, thresholds, and records are instructional fixtures, not a deployed system or client result.
- Owner
- You are the application engineer for a paid-media operations team.
- Decision
- Decide whether to block, request review, or execute each proposed action, then prove that approved execution is bound to the reviewed payload.
- Starting state
- The research agent may read campaign performance. A new tool can propose a daily budget change, but the advertising adapter remains a fake. The fixture includes a normal proposal, a 70 percent increase, a request containing injected instructions, and a proposal whose evidence changes while it waits for review.
- Expected outcome
- Normal reads continue, every budget proposal pauses with a durable review packet, prohibited or injected requests fail closed, and only an unexpired approval from a separate authorized reviewer can resume one matching fake effect.
Constraints
- • All budget writes stay inside a fake adapter; the lab is not authorized to call an advertising platform.
- • A change of up to 10 percent may be proposed but still requires a named reviewer; larger changes are blocked for this exercise.
- • Approval expires after 30 minutes and is valid only for the canonical hash of the reviewed arguments and evidence IDs.
- • A reviewer cannot approve a request they created, and a single approval token can be consumed once.
- • Retrieved content is data only and cannot alter tool permissions, approval policy, or system instructions.
Worked example
Beacon rejects a changed payload after approval
Evidence status: Named synthetic scenarioRun beacon-014 proposes changing campaign cmp_42 from USD 1,000 to USD 1,080 per day. The application validates the campaign ID, computes the 8 percent delta, records evidence ev_7 and ev_9, serializes the arguments in canonical key order, and stores the SHA-256 digest with policy budget-v3. The agent receives pending_review rather than a tool result that implies execution.
Reviewer rev_mina, who did not create the request, approves the displayed payload. Before resume, the fixture changes daily_budget to USD 1,120 while retaining the old approval ID. The executor recomputes the digest, detects the mismatch, marks the approval invalid, and writes no effect. A second unchanged fixture receives a fresh approval and produces one fake adapter receipt. A repeated resume returns the existing receipt rather than issuing a second write.
The trace contains proposed, pending_review, approved, effect_started, and effect_succeeded events for the valid run. The tampered run ends approval_invalid with expected and observed digests. The denial case ends denied and cannot be reopened. An automated assertion finds exactly one receipt for the valid idempotency key and none for the other fixtures.
Limits
All thresholds, identities, campaign data, and effects are synthetic. The exercise proves state-machine properties against a fake adapter, not the correctness of a real media-platform authorization model. OpenAI documents approval interruptions and resumable state, while exact SDK types can change. Treat the supplied record as an application-level pattern and verify current provider documentation before wiring an SDK.
Method
Build it, with checkpoints
Field situation
Decide whether to block, request review, or execute each proposed action, then prove that approved execution is bound to the reviewed payload.
- 01Classify every tool action
- 02Persist the immutable packet
- 03Authorize and resume
Acceptance checks
The fixture report shows one approved and consumed action, one denial, and four blocked adversarial paths. The only effect is a single fake receipt bound to the reviewed digest. Every terminal state names the policy check that decided it.
- 01
Classify every tool action
Create an allow, review, and prohibit table. Reads of approved campaign fields are allowed. Budget proposals enter review. Direct publish, credential access, and requests outside the campaign allowlist are prohibited before a model call can reinterpret them.
CHECKPOINT · A table-driven test maps every registered tool to one policy class and fails when a new tool has no classification.
- 02
Persist the immutable packet
Validate typed arguments, sort evidence IDs, hash the canonical payload, and store requester, policy version, expiry, and run ID. Return only a pending-review reference to the loop. Do not retain approval state solely in conversation text.
CHECKPOINT · Stop and restart the test process; the pending packet can still be loaded with the same digest and cannot be edited in place.
- 03
Authorize and resume
Check reviewer role and separation of duty, record approve or deny as an append-only event, then make resume recheck expiry, payload digest, policy version, and consumption state. Use the run's effect idempotency key at the adapter boundary.
CHECKPOINT · The valid fixture writes one fake receipt, denial writes none, and a second resume of the valid fixture returns the original receipt.
- 04
Red-team the boundary
Insert a retrieved document that says to bypass review, mutate an approved amount, expire a pending item, reuse a token for another run, and submit a self-approval. Assert explicit terminal reasons rather than a generic model apology.
CHECKPOINT · All five adversarial fixtures fail closed, preserve their audit events, and leave the fake campaign record unchanged.
- 05
Review the reviewer experience
Render current and proposed values, relative change, evidence links, policy version, requester, expiry, and a warning for untrusted source text. Time one reviewer using only the packet and record what information was missing.
CHECKPOINT · A reviewer can identify scope, magnitude, evidence, and expiry without opening the raw trace, and the packet never displays hidden chain-of-thought.
Hands-on lab
Implement a durable approval envelope
Extend the Beacon fake budget tool so a proposal becomes an immutable review packet, then attempt approval, denial, tampering, expiry, and duplicate resume.
Prepare
- • Keep the advertising adapter in memory and seed it with cmp_42 at USD 1,000 per day.
- • Provide fixed clocks and identities so expiry and separation-of-duty tests are deterministic.
- • Retain the raw retrieved text separately from the typed action arguments.
Deliverable
A TypeScript approval store and fake executor, six fixture tests, one rendered review-packet example, and an approval-state trace for every fixture.
Starter kit: Approval policy and record
TypeScripttype Decision = "pending" | "approved" | "denied" | "expired" | "consumed";
type BudgetArgs = {
campaignId: string;
currentDailyUsd: number;
proposedDailyUsd: number;
evidenceIds: string[];
};
type Approval = {
id: string;
runId: string;
tool: "propose_budget_change";
argsDigest: string;
policyVersion: "budget-v3";
requestedBy: string;
decidedBy?: string;
decision: Decision;
expiresAt: string;
consumedAt?: string;
};
const policy = {
allowedCampaign: /^cmp_[a-z0-9]+$/,
maxRelativeChange: 0.10,
approvalTtlMs: 30 * 60 * 1000,
rolesAllowedToApprove: new Set(["media-approver"]),
};
function canonicalBudgetArgs(args: BudgetArgs): string {
return JSON.stringify({
campaignId: args.campaignId,
currentDailyUsd: args.currentDailyUsd,
proposedDailyUsd: args.proposedDailyUsd,
evidenceIds: [...args.evidenceIds].sort(),
});
}
// Implement with node:crypto createHash("sha256"), then persist this digest.
function assertReviewable(args: BudgetArgs): void {
if (!policy.allowedCampaign.test(args.campaignId)) throw new Error("CAMPAIGN_DENIED");
const delta = Math.abs(args.proposedDailyUsd - args.currentDailyUsd) / args.currentDailyUsd;
if (!Number.isFinite(delta) || delta > policy.maxRelativeChange) throw new Error("POLICY_LIMIT");
if (args.evidenceIds.length < 2) throw new Error("EVIDENCE_REQUIRED");
}
// resume() must revalidate role, different reviewer, TTL, decision, digest,
// policy version, and idempotency key before calling the fake adapter once.Expected result
The fixture report shows one approved and consumed action, one denial, and four blocked adversarial paths. The only effect is a single fake receipt bound to the reviewed digest. Every terminal state names the policy check that decided it.
Carry forward
Keep the approval events, policy outcomes, and red-team fixtures. Module 9 converts them into regression cases, and module 10 uses their rates and handling time in its production gate.
Acceptance checks
- 01Every non-read action is classified, and unclassified tools fail closed.
- 02Approval persists exact arguments, evidence IDs, policy version, requester, reviewer, expiry, decision, and payload digest.
- 03Tampered, expired, self-approved, cross-run, and repeated-resume fixtures cannot create a new effect.
- 04The valid fixture resumes after a simulated restart and produces exactly one auditable fake receipt.
What breaks
Failure clinic
F1The approval screen shows a harmless action, but execution uses different arguments.
- Inspect
- Compare the canonical payload and digest at request, decision, and effect start.
- Likely cause
- Approval is bound only to a tool name or conversation turn.
- Repair
- Invalidate the decision and bind approval to immutable canonical arguments, evidence, policy version, run, and expiry.
- Prevent next time
- Recompute and compare the digest immediately beside the side-effect adapter.
F2A retry creates two budget-change receipts.
- Inspect
- Query effect records and adapter calls by the run's idempotency key.
- Likely cause
- Consumption state was saved after the external call or the adapter ignored the idempotency key.
- Repair
- Reserve the effect transactionally, pass the stable key, and reconcile unknown outcomes before retry.
- Prevent next time
- Test timeout-after-write and duplicate-resume cases on every release.
F3Retrieved copy convinces the agent that review is unnecessary.
- Inspect
- Trace provenance and check whether untrusted text entered policy or system-instruction fields.
- Likely cause
- The application allowed model-generated text to change authorization state.
- Repair
- Restore typed data boundaries and enforce tool policy in code outside the model context.
- Prevent next time
- Label untrusted fields, constrain tool inputs, and keep permission decisions deterministic.
F4Review queues grow even though nearly every request is approved unchanged.
- Inspect
- Segment review volume, decision time, edits, denials, and incident value by action class.
- Likely cause
- The team sends routine low-risk reads and high-risk writes through one undifferentiated gate.
- Repair
- Automate validated reads, prohibit disallowed actions, and reserve judgment for meaningful side effects.
- Prevent next time
- Review the policy matrix and reviewer burden against incident data each month.
Beyond the demo
Production boundary
- 01Classify every tool as allowed, reviewed, or prohibited and fail closed on missing policy.
- 02Enforce authorization and argument validation in application code beside the effect.
- 03Persist immutable review packets independently of model conversation state.
- 04Bind decisions to canonical payload, evidence, run, policy version, actor, and expiry.
- 05Require separation of duty for material actions and audit reviewer authorization.
- 06Use idempotency keys, effect ledgers, and reconciliation for unknown write outcomes.
- 07Keep untrusted retrieved content outside instruction and permission channels.
- 08Monitor approval volume, wait time, edit rate, denial rate, bypass attempts, and incidents.
Evidence status
Sources and claim limits
Sources support the named claims; they do not guarantee the same result in another system.
- [1]Guardrails and human reviewapproval interruptions · resumable state · tool-level guardrails
OpenAI · Official documentation · 2026-08-20
- [2]Safety in building agentsprompt injection · structured data boundaries · MCP approvals
OpenAI · Official documentation · 2026-08-20
- [3]MCP and Connectorsremote MCP configuration · approval modes · private server connectivity
OpenAI · Official documentation · 2026-08-20
- [4]AI Risk Management Frameworkrisk governance · measurement · operational accountability
NIST · Official documentation · 2026-08-20
Related Tenten resources