Skip to main content

bounded AI agent loop orchestration state machine

Lab

Agent loops and orchestration: make progress observable

Implement an external state machine with budgets, checkpoints, repetition detection, terminal states, and an escalation packet for a read-only research job.

DIFFICULTY
Advanced
ESTIMATED TIME
150 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Keep loop state and transition policy outside conversation text
  • Enforce step, time, cost, retry, and side-effect budgets deterministically
  • Detect repeated actions and missing progress from recorded observations
  • Persist checkpoints and produce a useful escalation packet on every non-success terminal state

Before you start

  • • Modules 02 through 05 artifacts
  • • Ability to test async code with fake clocks and injected tool failures

Working definition

Loops and orchestration

An agent loop repeatedly builds current context from state, asks a model for a typed next action, validates the action, executes an allowed tool, records the observation, and transitions through an application-owned state machine until completion or another terminal condition.

The loop concentrates the risks that a single model call can avoid: compounding assumptions, duplicate actions, cost growth, stalled progress, and partial effects after interruption.

A transcript can explain what the model said. Operators still need explicit state to know what completed, which limit stopped the run, and whether resuming is safe.

Field situation

Northstar evidence loop

Named synthetic scenario that reuses Northstar from Module 00. Tool observations, costs, limits, and traces are fixtures and not production measurements.

Owner
You are implementing the read-only loop for quarterly competitor governance research.
Decision
Define the state machine and deterministic stop rules around model-selected read actions.
Starting state
The model can search approved domains through a fake tool, read evidence, and mark a competitor covered or missing. The prototype asks the model whether it is done and retries vague calls indefinitely.
Expected outcome
Success produces a complete evidence ledger. Budget, timeout, non-progress, tool exhaustion, policy denial, and review-required each produce a resumable or terminal escalation packet.

Constraints

  • • Maximum eight model steps, 90 seconds, USD 1.00 estimated variable cost, and two transient retries per tool.
  • • No writes, publishing, outreach, login, or unrestricted browsing are available.
  • • Completion requires two current evidence records or one explicit evidence gap for each of four competitors.
  • • Two identical normalized actions without a changed observation trigger non_progress.

Worked example

The third identical search ends the run instead of spending another turn

Evidence status: Named synthetic scenario

The model requests search_domain for competitor Orion with query enterprise governance. The tool returns the same two irrelevant URLs twice. The model proposes the normalized same call for a third time, without changing query, source, or plan.

The orchestrator hashes the validated action and compares the last observation digest. After the second unchanged pair it transitions to non_progress before executing the third call. The escalation packet contains goal, coverage table, attempted query, two observations, remaining budget, and suggested operator choices.

The fixture proves the repository tool ran twice, the third request was not executed, cost accounting stopped, and the terminal reason is machine-readable. It does not claim that this heuristic catches every form of circular reasoning.

Limits

Action hashing catches exact or normalized repetition, not semantically equivalent loops with different wording. Cost estimates can lag provider billing. Long-running jobs need durable queues, leases, concurrency control, and recovery beyond the in-memory starter.

Method

Build it, with checkpoints

State machine for a bounded research agent with decide, validate, execute, observe, checkpoint, complete, non-progress, timeout, budget, policy, tool failure, and human review states.

Field situation

Define the state machine and deterministic stop rules around model-selected read actions.

  1. 01Define state and invariants
  2. 02Wrap the decision call
  3. 03Checkpoint around execution

Acceptance checks

An adaptive read loop whose decisions can vary while limits, evidence requirements, state, recovery, and terminal behavior remain deterministic and inspectable.

Why this visualRender the loop as a deterministic state machine and one annotated sequence trace. States, guard conditions, checkpoints, retries, and terminal reasons need source-controlled labels.
  1. 01

    Define state and invariants

    List mutable state, allowed transitions, completion evidence, and terminal reasons. Keep budgets and policy outside the next-action prompt.

    CHECKPOINT · Every transition has one code owner, and completed cannot be reached until all four coverage records pass.

  2. 02

    Wrap the decision call

    Ask the adapter for one typed action from the current state summary. Validate tool name, arguments, permissions, and remaining budget before dispatch.

    CHECKPOINT · An unknown or prohibited action transitions to policy_denied and never reaches a tool.

  3. 03

    Checkpoint around execution

    Persist planned action and idempotency key, execute with deadline and bounded retries, then persist observation, cost, and coverage changes.

    CHECKPOINT · An interruption fixture can resume without losing the observation or dispatching a completed action again.

  4. 04

    Detect non-progress

    Normalize action and observation digests. Increment unchangedPairs only when both repeat; reset it after genuinely new evidence or a changed plan.

    CHECKPOINT · The Orion fixture performs two searches, blocks the third identical dispatch, and ends non_progress.

  5. 05

    Serialize every stop

    Produce an escalation packet with goal, terminal reason, completed work, missing evidence, attempted actions, last errors, versions, spend estimate, and safe options.

    CHECKPOINT · Success and all six non-success fixtures contain enough state for an operator to close, revise, or resume the job without reading raw reasoning.

Hands-on lab

Build the Northstar loop controller

Wire a fake next-action adapter to search and read fixtures, then prove success, non-progress, timeout, budget exhaustion, policy denial, and retry exhaustion.

Prepare

  • • Reuse authoritative jobs and effects from Module 03.
  • • Use a fake clock and deterministic cost estimator for offline tests.
  • • Keep tools read-only and validate every action against an allowlist.
  • • Persist state before and after each tool dispatch in the completed version.

Deliverable

A typed state machine, loop controller, fake action adapter, two read tools, budget accounting, repetition detector, checkpoint store, seven terminal fixtures, and escalation serializer.

Starter kit: Bounded loop state machine

TypeScript
type Terminal = "completed" | "non_progress" | "budget_exhausted" | "timed_out" | "tool_exhausted" | "policy_denied" | "review_required";
type LoopState = {
  runId: string;
  status: "ready" | "deciding" | "executing" | Terminal;
  step: number;
  startedAtMs: number;
  estimatedCostUsd: number;
  coverage: Record<string, { evidenceIds: string[]; gap: boolean }>;
  lastActionHash: string | null;
  lastObservationHash: string | null;
  unchangedPairs: number;
};

const limits = { maxSteps: 8, maxWallMs: 90_000, maxCostUsd: 1.00, maxRetries: 2 } as const;

function shouldStop(s: LoopState, nowMs: number): Terminal | null {
  if (Object.values(s.coverage).every((x) => x.gap || x.evidenceIds.length >= 2)) return "completed";
  if (s.unchangedPairs >= 2) return "non_progress";
  if (s.step >= limits.maxSteps || s.estimatedCostUsd >= limits.maxCostUsd) return "budget_exhausted";
  if (nowMs - s.startedAtMs >= limits.maxWallMs) return "timed_out";
  return null;
}

Expected result

An adaptive read loop whose decisions can vary while limits, evidence requirements, state, recovery, and terminal behavior remain deterministic and inspectable.

Carry forward

Keep the loop state, terminal taxonomy, escalation packet, and trace fixtures. Module 07 will compare this single-agent baseline with a multi-agent breadth-first variant.

Acceptance checks

  1. 01All seven terminal fixtures finish within their configured step and time limits.
  2. 02No prohibited or over-budget action reaches a tool implementation.
  3. 03The repetition fixture executes no third identical action after two unchanged pairs.
  4. 04Every checkpoint can resume once without duplicating a completed tool effect or losing cost and evidence state.

What breaks

Failure clinic

F1The agent keeps searching without increasing evidence coverage.
Inspect
Compare normalized actions, observation digests, coverage deltas, query changes, and remaining budget across turns.
Likely cause
Completion and non-progress were left to model judgment, or the loop does not record a measurable delta.
Repair
Define evidence-level progress, block unchanged action-observation pairs, and escalate with the current gap.
Prevent next time
Keep non-progress fixtures and a maximum action count for every tool.
F2A resumed run repeats a tool call after the process crashed.
Inspect
Review planned and completed effect records, idempotency key, lease, and checkpoint order.
Likely cause
The system persisted the observation after acknowledging completion or lacked an execution ledger.
Repair
Use durable planned/completed effects and make tools idempotent or reconciliation-aware.
Prevent next time
Inject crashes at each persistence boundary and verify one logical effect.
F3The run exceeds its cost ceiling even though step count is below eight.
Inspect
Break down input, output, tool, retry, cache, and parallel call costs by step and compare estimate lag.
Likely cause
The budget tracks turns only or checks spend after dispatch.
Repair
Reserve estimated cost before execution and refuse a call that could breach the remaining ceiling.
Prevent next time
Set per-action upper bounds and reconcile estimated with billed cost after each run.
F4An operator receives a failed status with no safe way to continue.
Inspect
Check terminal reason, last valid checkpoint, completed effects, missing evidence, errors, and available authorized actions.
Likely cause
Failure handling focused on logging rather than operational continuation.
Repair
Serialize a structured escalation packet and map each terminal reason to close, revise, or resume options.
Prevent next time
Review failure packets with the actual on-call role during readiness testing.

Beyond the demo

Production boundary

  1. 01Keep goal, progress, approvals, effects, budgets, and terminal status in durable application state.
  2. 02Validate every proposed action against tool, argument, identity, policy, and remaining-budget rules.
  3. 03Set maximum steps, wall time, variable cost, retries, action counts, and payload size.
  4. 04Checkpoint planned and completed effects with stable idempotency keys.
  5. 05Detect repeated actions, unchanged observations, circular delegation, and validation churn.
  6. 06Trace state transition, model call, validated action, tool result, cost, and terminal reason.
  7. 07Provide staffed escalation paths with structured context and safe continuation choices.
  8. 08Rehearse crash, timeout, provider outage, policy denial, budget exhaustion, and manual failover.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Building effective agents

    Anthropic · Official documentation · 2026-08-20

    workflow and agent definitions · architecture patterns · cost and control tradeoffs
  2. [2]
    Orchestration and handoffs

    OpenAI · Official documentation · 2026-08-20

    handoffs · agents as tools · single-agent baseline
  3. [3]
    Guardrails and human review

    OpenAI · Official documentation · 2026-08-20

    approval interruptions · resumable state · tool-level guardrails
  4. [4]
    Evaluate agent workflows

    OpenAI · Official documentation · 2026-08-20

    trace grading · datasets · repeatable eval runs

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.