Skip to main content

AI agent vs workflow architecture decision

Foundation

AI agents vs workflows: earn the loop

Compare three implementations of the same research intake task and keep the least autonomous design that clears the acceptance gate.

DIFFICULTY
Beginner
ESTIMATED TIME
75 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Describe workflow and agent boundaries in observable operational terms
  • Separate interpretation uncertainty from stable business rules
  • Compare quality, latency, cost, permission surface, and recovery effort
  • Record why an agent loop is or is not justified for a named task

Before you start

  • • Basic API and event-driven software concepts
  • • A local TypeScript runtime or a spreadsheet for the non-code comparison

Working definition

Agents vs workflows

A workflow moves through control paths chosen in code. A model-assisted workflow gives a model one or more bounded decisions inside that path. An agent loop lets a model choose successive actions from an allowed tool set until application-owned completion, escalation, or budget conditions stop the run.

A stable process hidden inside an agent prompt is harder to review and recover than the same process expressed as code. The model can still help at the uncertain step without owning the rest of the control plane.

Autonomy changes the failure surface. Each additional turn can add cost, repeat a bad assumption, request a stronger tool, or make recovery depend on incomplete conversation history.

Field situation

Northstar research intake

Named synthetic scenario. Northstar Analytics and all measurements in this module are fixtures created for instruction; they are not a Tenten or client deployment.

Owner
You are the platform engineer supporting a six-person B2B marketing research team.
Decision
Choose a deterministic workflow, a model-assisted workflow, or a bounded agent for each request class.
Starting state
Researchers receive requests through a shared form. Every request needs scope validation, source-policy selection, evidence collection, and analyst review. The current queue contains routine competitor updates and open-ended market investigations.
Expected outcome
Routine updates follow deterministic code, ambiguous scoping uses one structured model decision, and only the breadth-first investigation qualifies for a read-only bounded loop after it beats the simpler baseline.

Constraints

  • • The system is read-only and cannot publish, email, buy media, or alter CRM records.
  • • A run has a USD 1.00 model and tool budget, a 90-second wall-clock limit, and at most eight model turns.
  • • Only approved public domains and a synthetic internal corpus may be searched.
  • • An analyst owns the final brief and can reject unsupported claims without explanation from the model.

Worked example

Three architectures, one request fixture

Evidence status: Named synthetic scenario

The request is: Compare how four named competitors describe enterprise governance on their current product pages, preserve exact URLs, and identify missing evidence. Variant A follows a fixed list of URLs and extraction rules. Variant B uses the same path but asks a model to classify page passages into a fixed governance taxonomy. Variant C lets an agent search approved domains, open pages, deduplicate evidence, and stop when every competitor has either two current source records or a recorded evidence gap.

Variant A fails when a competitor moves its governance page. Variant B recovers only when the new URL is supplied, but its typed classifier handles wording differences well. Variant C finds moved pages and records gaps, yet it uses more calls and needs loop, tool, and citation controls. The team selects Variant B for weekly monitoring and reserves Variant C for quarterly investigations where path uncertainty is the actual work.

On the supplied 12-case fixture, Variant B is the intended answer for eight cases, Variant A for three stable exports, and Variant C for one open-ended investigation. The lab does not claim measured provider performance. Learners run their own comparison and keep the raw scorecard as the evidence.

Limits

The request set is small and synthetic. Cost and latency depend on provider, region, model, cache state, and tool implementation. Anthropic's workflow-versus-agent guidance supports the decision framework, while the numeric thresholds here are Northstar policy choices rather than provider recommendations.

Method

Build it, with checkpoints

Decision tree comparing a deterministic workflow, a model-assisted workflow, and a bounded agent by path uncertainty, action risk, and stopping controls.

Field situation

Choose a deterministic workflow, a model-assisted workflow, or a bounded agent for each request class.

  1. 01Mark stable and uncertain decisions
  2. 02Draw the three control paths
  3. 03Run the request matrix

Acceptance checks

A defensible architecture choice in which the agent is an exception for genuinely variable research, not the default wrapper around stable API calls.

Why this visualThe lesson compares control ownership. Use a deterministic Mermaid or SVG decision tree plus three side-by-side topologies; generated imagery cannot represent branches, tool authority, or terminal states accurately.
  1. 01

    Mark stable and uncertain decisions

    List every branch in the request. Put stable policy, budgets, source allowlists, and completion checks under deterministic control. Put only interpretation or search choices in the uncertain column.

    CHECKPOINT · The worksheet names at least one stable rule and one uncertain decision; no policy limit appears under model control.

  2. 02

    Draw the three control paths

    Sketch deterministic, model-assisted, and loop variants. Show who owns the next step, which tools each variant can request, where evidence is stored, and every terminal state.

    CHECKPOINT · Each topology has completion, escalation, budget-exhausted, and failed states, even when a state is unreachable in the deterministic version.

  3. 03

    Run the request matrix

    Score 12 supplied or self-authored request fixtures for path uncertainty, action risk, value of adaptation, and cost of a wrong branch. Keep the scoring formula visible.

    CHECKPOINT · A second reviewer can reproduce the selected architecture from the scores without reading an unstated rationale.

  4. 04

    Compare operation, not prose

    Record expected calls, maximum actions, recovery procedure, review load, and required telemetry. A fluent sample answer earns no credit unless the evidence and control path pass.

    CHECKPOINT · The scorecard contains quality, latency, cost, permissions, observability, and rollback columns for every option.

  5. 05

    Write the decision and a reversal trigger

    Select the least autonomous passing option. State which measured failure would justify more autonomy and which failure would send the design back to a simpler workflow.

    CHECKPOINT · The final record includes one promotion trigger, one simplification trigger, a named owner, and a review date.

Hands-on lab

Write an autonomy decision record

Classify 12 Northstar requests, sketch all three candidate architectures for one request, and defend the minimum autonomy that meets the stated outcome.

Prepare

  • • Copy the YAML starter below into autonomy-review.yaml.
  • • Choose one timer and one cost-estimation method and state their uncertainty.
  • • Do not call an external API unless you already have an approved test account and spend cap.
  • • Keep the scenario read-only throughout the exercise.

Deliverable

A completed YAML decision record, three small topology sketches, and a scorecard that compares the same request across the candidate designs.

Starter kit: Autonomy review worksheet

YAML
scenario: northstar-research-intake
request: "Compare governance claims across four approved competitor domains"
owner: marketing-research-lead
options:
  deterministic:
    path_known: true
    uncertain_decisions: []
    permissions: [read_fixture]
  model_assisted:
    path_known: true
    uncertain_decisions: [classify_governance_passage]
    permissions: [read_fixture]
  bounded_agent:
    path_known: false
    uncertain_decisions: [choose_next_source, decide_evidence_gap]
    permissions: [search_approved_domains, read_public_page]
limits:
  max_turns: 8
  max_wall_seconds: 90
  max_cost_usd: 1.00
  side_effects: 0
acceptance:
  citation_coverage: 1.0
  unsupported_claims: 0
  forbidden_tool_calls: 0
decision: TODO
reason: TODO

Expected result

A defensible architecture choice in which the agent is an exception for genuinely variable research, not the default wrapper around stable API calls.

Carry forward

Keep autonomy-review.yaml. Module 06 will turn its bounded-agent branch into an external loop state machine, and Module 10 will reuse its limits in the production readiness review.

Acceptance checks

  1. 01Every stable policy and limit is enforced outside model-generated text.
  2. 02The bounded-agent option has explicit turn, time, cost, permission, and terminal-state limits.
  3. 03All three options are scored against the same request fixtures and outcome rubric.
  4. 04The chosen design includes a manual continuation path when the model or a tool is unavailable.

What breaks

Failure clinic

F1The agent follows the same tool order on every run yet still costs more than the workflow.
Inspect
Compare five traces and diff their tool sequences, branch decisions, and termination reasons.
Likely cause
A stable process was placed inside a model loop even though no adaptive decision changes the path.
Repair
Move the sequence into code and retain a bounded model call only where classification or extraction is uncertain.
Prevent next time
Require trace diversity and measurable task improvement before approving an autonomous loop.
F2A prompt update silently changes a compliance branch that reviewers thought was fixed policy.
Inspect
Locate each business rule in code, prompt, tool description, and reviewer instructions; flag duplicates and prompt-only rules.
Likely cause
Deterministic policy was written as natural-language guidance rather than an application invariant.
Repair
Encode the rule as validation or routing code and add a regression fixture at the boundary.
Prevent next time
Maintain a policy inventory that names one executable owner for every hard rule.
F3The design review approves broad search and write permissions before the uncertain decision is understood.
Inspect
Map each permission to a specific objective and trace the proposed damage if that permission is misused.
Likely cause
The team selected a platform capability first and retrofitted a task boundary around it.
Repair
Return to the named request, remove unused tools, and begin with a read-only fixture.
Prevent next time
Make least privilege and a no-agent baseline mandatory fields in every architecture decision record.
F4The comparison crowns the agent because its prose sounds better, despite missing source records.
Inspect
Grade citation coverage, unsupported claims, completion state, calls, and review effort separately from writing quality.
Likely cause
The rubric measures presentation while the task outcome depends on evidence collection and operational control.
Repair
Use deterministic evidence checks first, then apply a calibrated writing rubric only to passing outputs.
Prevent next time
Lock the task-level acceptance checks before generating any candidate answers.

Beyond the demo

Production boundary

  1. 01Name the process owner and the person authorized to change the autonomy level.
  2. 02Keep policy, permissions, budgets, and terminal states in deterministic application code.
  3. 03Start with read-only tools and document why each additional capability is necessary.
  4. 04Version the decision record with prompt, model, tool, schema, and fixture identifiers.
  5. 05Trace each model decision and tool request with one run correlation ID.
  6. 06Define idempotency or replay behavior before any reversible write is introduced.
  7. 07Price a successful outcome using model, tool, infrastructure, review, and rework costs.
  8. 08Re-run the simpler baseline after material model or tool changes; complexity must keep earning its place.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Building effective agents

    Anthropic · Official documentation · 2026-08-20

    workflow and agent definitions · architecture patterns · cost and control tradeoffs
  2. [2]
    Orchestration and handoffs

    OpenAI · Official documentation · 2026-08-20

    handoffs · agents as tools · single-agent baseline
  3. [3]
    New tools for building agents

    OpenAI · Public case · 2026-08-20

    Responses API architecture · Agents SDK capabilities · reported customer implementations
  4. [4]
    AI Risk Management Framework

    NIST · Official documentation · 2026-08-20

    risk governance · measurement · operational accountability

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.