On this page
Learning objectives
- Describe workflow and agent boundaries in observable operational terms
- Separate interpretation uncertainty from stable business rules
- Compare quality, latency, cost, permission surface, and recovery effort
- Record why an agent loop is or is not justified for a named task
Before you start
- • Basic API and event-driven software concepts
- • A local TypeScript runtime or a spreadsheet for the non-code comparison
Working definition
Agents vs workflows
A workflow moves through control paths chosen in code. A model-assisted workflow gives a model one or more bounded decisions inside that path. An agent loop lets a model choose successive actions from an allowed tool set until application-owned completion, escalation, or budget conditions stop the run.
A stable process hidden inside an agent prompt is harder to review and recover than the same process expressed as code. The model can still help at the uncertain step without owning the rest of the control plane.
Autonomy changes the failure surface. Each additional turn can add cost, repeat a bad assumption, request a stronger tool, or make recovery depend on incomplete conversation history.
Field situation
Northstar research intake
Named synthetic scenario. Northstar Analytics and all measurements in this module are fixtures created for instruction; they are not a Tenten or client deployment.
- Owner
- You are the platform engineer supporting a six-person B2B marketing research team.
- Decision
- Choose a deterministic workflow, a model-assisted workflow, or a bounded agent for each request class.
- Starting state
- Researchers receive requests through a shared form. Every request needs scope validation, source-policy selection, evidence collection, and analyst review. The current queue contains routine competitor updates and open-ended market investigations.
- Expected outcome
- Routine updates follow deterministic code, ambiguous scoping uses one structured model decision, and only the breadth-first investigation qualifies for a read-only bounded loop after it beats the simpler baseline.
Constraints
- • The system is read-only and cannot publish, email, buy media, or alter CRM records.
- • A run has a USD 1.00 model and tool budget, a 90-second wall-clock limit, and at most eight model turns.
- • Only approved public domains and a synthetic internal corpus may be searched.
- • An analyst owns the final brief and can reject unsupported claims without explanation from the model.
Worked example
Three architectures, one request fixture
Evidence status: Named synthetic scenarioThe request is: Compare how four named competitors describe enterprise governance on their current product pages, preserve exact URLs, and identify missing evidence. Variant A follows a fixed list of URLs and extraction rules. Variant B uses the same path but asks a model to classify page passages into a fixed governance taxonomy. Variant C lets an agent search approved domains, open pages, deduplicate evidence, and stop when every competitor has either two current source records or a recorded evidence gap.
Variant A fails when a competitor moves its governance page. Variant B recovers only when the new URL is supplied, but its typed classifier handles wording differences well. Variant C finds moved pages and records gaps, yet it uses more calls and needs loop, tool, and citation controls. The team selects Variant B for weekly monitoring and reserves Variant C for quarterly investigations where path uncertainty is the actual work.
On the supplied 12-case fixture, Variant B is the intended answer for eight cases, Variant A for three stable exports, and Variant C for one open-ended investigation. The lab does not claim measured provider performance. Learners run their own comparison and keep the raw scorecard as the evidence.
Limits
The request set is small and synthetic. Cost and latency depend on provider, region, model, cache state, and tool implementation. Anthropic's workflow-versus-agent guidance supports the decision framework, while the numeric thresholds here are Northstar policy choices rather than provider recommendations.
Method
Build it, with checkpoints
Field situation
Choose a deterministic workflow, a model-assisted workflow, or a bounded agent for each request class.
- 01Mark stable and uncertain decisions
- 02Draw the three control paths
- 03Run the request matrix
Acceptance checks
A defensible architecture choice in which the agent is an exception for genuinely variable research, not the default wrapper around stable API calls.
- 01
Mark stable and uncertain decisions
List every branch in the request. Put stable policy, budgets, source allowlists, and completion checks under deterministic control. Put only interpretation or search choices in the uncertain column.
CHECKPOINT · The worksheet names at least one stable rule and one uncertain decision; no policy limit appears under model control.
- 02
Draw the three control paths
Sketch deterministic, model-assisted, and loop variants. Show who owns the next step, which tools each variant can request, where evidence is stored, and every terminal state.
CHECKPOINT · Each topology has completion, escalation, budget-exhausted, and failed states, even when a state is unreachable in the deterministic version.
- 03
Run the request matrix
Score 12 supplied or self-authored request fixtures for path uncertainty, action risk, value of adaptation, and cost of a wrong branch. Keep the scoring formula visible.
CHECKPOINT · A second reviewer can reproduce the selected architecture from the scores without reading an unstated rationale.
- 04
Compare operation, not prose
Record expected calls, maximum actions, recovery procedure, review load, and required telemetry. A fluent sample answer earns no credit unless the evidence and control path pass.
CHECKPOINT · The scorecard contains quality, latency, cost, permissions, observability, and rollback columns for every option.
- 05
Write the decision and a reversal trigger
Select the least autonomous passing option. State which measured failure would justify more autonomy and which failure would send the design back to a simpler workflow.
CHECKPOINT · The final record includes one promotion trigger, one simplification trigger, a named owner, and a review date.
Hands-on lab
Write an autonomy decision record
Classify 12 Northstar requests, sketch all three candidate architectures for one request, and defend the minimum autonomy that meets the stated outcome.
Prepare
- • Copy the YAML starter below into autonomy-review.yaml.
- • Choose one timer and one cost-estimation method and state their uncertainty.
- • Do not call an external API unless you already have an approved test account and spend cap.
- • Keep the scenario read-only throughout the exercise.
Deliverable
A completed YAML decision record, three small topology sketches, and a scorecard that compares the same request across the candidate designs.
Starter kit: Autonomy review worksheet
YAMLscenario: northstar-research-intake
request: "Compare governance claims across four approved competitor domains"
owner: marketing-research-lead
options:
deterministic:
path_known: true
uncertain_decisions: []
permissions: [read_fixture]
model_assisted:
path_known: true
uncertain_decisions: [classify_governance_passage]
permissions: [read_fixture]
bounded_agent:
path_known: false
uncertain_decisions: [choose_next_source, decide_evidence_gap]
permissions: [search_approved_domains, read_public_page]
limits:
max_turns: 8
max_wall_seconds: 90
max_cost_usd: 1.00
side_effects: 0
acceptance:
citation_coverage: 1.0
unsupported_claims: 0
forbidden_tool_calls: 0
decision: TODO
reason: TODOExpected result
A defensible architecture choice in which the agent is an exception for genuinely variable research, not the default wrapper around stable API calls.
Carry forward
Keep autonomy-review.yaml. Module 06 will turn its bounded-agent branch into an external loop state machine, and Module 10 will reuse its limits in the production readiness review.
Acceptance checks
- 01Every stable policy and limit is enforced outside model-generated text.
- 02The bounded-agent option has explicit turn, time, cost, permission, and terminal-state limits.
- 03All three options are scored against the same request fixtures and outcome rubric.
- 04The chosen design includes a manual continuation path when the model or a tool is unavailable.
What breaks
Failure clinic
F1The agent follows the same tool order on every run yet still costs more than the workflow.
- Inspect
- Compare five traces and diff their tool sequences, branch decisions, and termination reasons.
- Likely cause
- A stable process was placed inside a model loop even though no adaptive decision changes the path.
- Repair
- Move the sequence into code and retain a bounded model call only where classification or extraction is uncertain.
- Prevent next time
- Require trace diversity and measurable task improvement before approving an autonomous loop.
F2A prompt update silently changes a compliance branch that reviewers thought was fixed policy.
- Inspect
- Locate each business rule in code, prompt, tool description, and reviewer instructions; flag duplicates and prompt-only rules.
- Likely cause
- Deterministic policy was written as natural-language guidance rather than an application invariant.
- Repair
- Encode the rule as validation or routing code and add a regression fixture at the boundary.
- Prevent next time
- Maintain a policy inventory that names one executable owner for every hard rule.
F3The design review approves broad search and write permissions before the uncertain decision is understood.
- Inspect
- Map each permission to a specific objective and trace the proposed damage if that permission is misused.
- Likely cause
- The team selected a platform capability first and retrofitted a task boundary around it.
- Repair
- Return to the named request, remove unused tools, and begin with a read-only fixture.
- Prevent next time
- Make least privilege and a no-agent baseline mandatory fields in every architecture decision record.
F4The comparison crowns the agent because its prose sounds better, despite missing source records.
- Inspect
- Grade citation coverage, unsupported claims, completion state, calls, and review effort separately from writing quality.
- Likely cause
- The rubric measures presentation while the task outcome depends on evidence collection and operational control.
- Repair
- Use deterministic evidence checks first, then apply a calibrated writing rubric only to passing outputs.
- Prevent next time
- Lock the task-level acceptance checks before generating any candidate answers.
Beyond the demo
Production boundary
- 01Name the process owner and the person authorized to change the autonomy level.
- 02Keep policy, permissions, budgets, and terminal states in deterministic application code.
- 03Start with read-only tools and document why each additional capability is necessary.
- 04Version the decision record with prompt, model, tool, schema, and fixture identifiers.
- 05Trace each model decision and tool request with one run correlation ID.
- 06Define idempotency or replay behavior before any reversible write is introduced.
- 07Price a successful outcome using model, tool, infrastructure, review, and rework costs.
- 08Re-run the simpler baseline after material model or tool changes; complexity must keep earning its place.
Evidence status
Sources and claim limits
Sources support the named claims; they do not guarantee the same result in another system.
- [1]Building effective agentsworkflow and agent definitions · architecture patterns · cost and control tradeoffs
Anthropic · Official documentation · 2026-08-20
- [2]Orchestration and handoffshandoffs · agents as tools · single-agent baseline
OpenAI · Official documentation · 2026-08-20
- [3]New tools for building agentsResponses API architecture · Agents SDK capabilities · reported customer implementations
OpenAI · Public case · 2026-08-20
- [4]AI Risk Management Frameworkrisk governance · measurement · operational accountability
NIST · Official documentation · 2026-08-20
Related Tenten resources