Skip to main content

multi agent research architecture evaluation tutorial

Lab

Multi-agent patterns: compare before you split

Build a breadth-first research variant beside the single-agent baseline and keep it only when independent work improves measured coverage enough to justify coordination cost.

DIFFICULTY
Advanced
ESTIMATED TIME
165 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Choose between handoff ownership and manager-controlled specialists
  • Define narrow worker contracts, isolated credentials, and a global budget
  • Measure duplication, contradiction, citation validity, latency, tokens, and synthesis quality
  • Reject multi-agent for simple or dependency-heavy requests

Before you start

  • • Module 06 single-agent loop, terminal taxonomy, and traces
  • • A fixture runner that can execute independent tasks concurrently

Working definition

Multi-agent patterns

A multi-agent system assigns separate model-driven workers distinct tasks, context, or tools under shared orchestration. The split is useful when work can proceed independently or policy and capability boundaries need isolation; shared personas alone do not create an architectural benefit.

Parallel workers can search independent branches faster and preserve more local context. They can also duplicate searches, disagree without evidence, exhaust a shared budget, and give each credential another place to fail.

Anthropic reports large gains for its breadth-first research setup alongside much higher token use and poor fit for tightly coupled work. Those findings justify a comparison, not a default architecture.

Field situation

Horizon competitor landscape

Named synthetic scenario. Horizon Research, company set, questions, budgets, and measured outputs are local fixtures. Anthropic's published multi-agent figures are cited separately and are not assigned to Horizon.

Owner
You are comparing a single research loop with a manager that calls bounded company-research workers as tools.
Decision
Route each question to deterministic, single-agent, or manager-plus-workers execution based on independence and expected value.
Starting state
The benchmark has 20 questions: five direct facts, ten independent multi-company comparisons, and five analyses where later work depends on shared earlier evidence.
Expected outcome
Direct facts avoid multi-agent overhead, independent comparisons can earn parallel workers, and dependency-heavy questions remain with one shared-context loop unless measured evidence says otherwise.

Constraints

  • • The manager may launch at most four workers and must keep final answer ownership.
  • • Each worker receives one company or source branch, read-only tools, a fixed output schema, and a local call budget.
  • • A global controller enforces wall time, spend, source allowlists, and total workers.
  • • The synthesis may use only evidence IDs present in accepted worker outputs.

Worked example

Four workers finish fast and still fail because two searched the same branch

Evidence status: Named synthetic scenario

Question HRZ-14 asks for governance claims from four companies. A vague manager instruction launches four workers, but two research the same company and none covers company Delta. The synthesis looks complete because it merges eight citations without checking company coverage.

The worker contract is changed to include assigned entity, excluded entities, required evidence count, source-policy version, and evidence IDs. The manager reserves one distinct entity per worker, validates coverage before synthesis, and returns evidence_gap if a branch fails. Duplicate URLs are deduplicated but not counted as extra coverage.

The repaired fixture covers all four assigned entities or records a gap, while the original fails deterministic coverage despite polished prose. The local benchmark decides whether latency savings outweigh token and coordination cost; no universal multi-agent gain is asserted.

Limits

The manager-as-tools pattern keeps one final owner, which suits this report. A customer-facing specialist may justify handoff instead. Model non-determinism requires repeated trials, and parallel calls can hit rate, connection, or provider concurrency limits that the in-memory runner does not reproduce.

Method

Build it, with checkpoints

Comparison of a single research agent and a manager with four scoped workers, showing global budgets, isolated tools, evidence contracts, synthesis gate, and evaluation outputs.

Field situation

Route each question to deterministic, single-agent, or manager-plus-workers execution based on independence and expected value.

  1. 01Classify task shape
  2. 02Define non-overlapping workers
  3. 03Enforce one global controller

Acceptance checks

A task router grounded in one shared benchmark. Multi-agent remains a measured option for independent breadth, while simple and tightly coupled work use smaller control surfaces.

Why this visualUse a deterministic supervisor-worker graph beside the single-agent topology and a benchmark chart generated from fixture results. The value claim depends on visible assignments, boundaries, and measured tradeoffs.
  1. 01

    Classify task shape

    Label each question direct, independently decomposable, or dependency-heavy. Write why shared context is or is not needed before running either system.

    CHECKPOINT · The routing label is frozen in the fixture and cannot be changed after seeing which architecture won.

  2. 02

    Define non-overlapping workers

    Give each worker an assigned entity, exclusions, source policy, tool cap, output schema, and no authority to create further workers.

    CHECKPOINT · The HRZ-14 plan assigns four distinct entities and the test rejects duplicate assignments.

  3. 03

    Enforce one global controller

    Reserve worker and cost budgets before launch, cancel outstanding work at the deadline, correlate traces, and validate every result before synthesis.

    CHECKPOINT · No combination of workers can exceed the global count or spend ceiling even when they finish concurrently.

  4. 04

    Grade baseline and workers

    Measure entity coverage, unique valid sources, contradictions, citation validity, wall time, total calls, tokens or estimated cost, and reviewer corrections.

    CHECKPOINT · A coverage failure cannot be hidden by a high writing-quality score, and duplicate URLs count once.

  5. 05

    Write the routing policy

    Select the architecture by task shape and observed value. Include a downgrade rule when coordination cost or failure rate exceeds the declared ceiling.

    CHECKPOINT · Direct questions do not launch workers, and every multi-agent route has a single-agent fallback plus a named review date.

Hands-on lab

Benchmark manager-plus-workers against one agent

Run all 20 Horizon questions through the Module 06 baseline and a four-worker-cap manager, then write a routing policy from the measured results.

Prepare

  • • Freeze the question set, source fixture, grader, and total budget for both variants.
  • • Use fake provider adapters first so orchestration tests are deterministic.
  • • Assign each worker a separate credential object even when all use the same fake backend.
  • • Run repeated trials if a live model is enabled and report the trial count.

Deliverable

A manager, worker contract, global budget, isolated role configs, single-agent baseline, 20-question eval, repeated-trial report, and measured routing policy.

Starter kit: Worker request and result contracts

TypeScript
type WorkerRequest = {
  taskId: string;
  assignedEntity: string;
  excludedEntities: string[];
  question: string;
  sourcePolicyVersion: string;
  maxToolCalls: number;
};

type WorkerResult = {
  taskId: string;
  assignedEntity: string;
  status: "covered" | "evidence_gap" | "failed";
  evidenceIds: string[];
  sourceUrls: string[];
  limitations: string[];
  toolCalls: number;
  estimatedCostUsd: number;
};

type GlobalBudget = {
  maxWorkers: 4;
  maxWallMs: number;
  maxCostUsd: number;
  usedCostUsd: number;
};

function canSynthesize(expected: string[], results: WorkerResult[]): boolean {
  return expected.every((entity) =>
    results.some((r) => r.assignedEntity === entity && (r.status === "covered" || r.status === "evidence_gap"))
  );
}

Expected result

A task router grounded in one shared benchmark. Multi-agent remains a measured option for independent breadth, while simple and tightly coupled work use smaller control surfaces.

Carry forward

Keep the task router, worker schema, global budget, evidence deduplicator, and comparison report. The capstone may enable workers only for questions that satisfy this policy.

Acceptance checks

  1. 01Both architectures use the same source corpus, question set, outcome graders, and total budget accounting.
  2. 02Each worker has a distinct assignment, isolated tool config, and no recursive-delegation capability.
  3. 03Synthesis starts only after every expected branch is covered or marked as an evidence gap.
  4. 04The report includes quality, duplication, contradictions, citations, wall time, calls, cost, and review effort by task shape.

What breaks

Failure clinic

F1Several workers return many sources, yet one required entity has no evidence.
Inspect
Compare assignment IDs, excluded entities, result coverage, duplicate URLs, and synthesis preconditions.
Likely cause
Manager instructions were vague and the synthesizer counted citations rather than required branches.
Repair
Reserve explicit non-overlapping assignments and validate entity coverage before synthesis.
Prevent next time
Reject duplicate assignments and make evidence gaps first-class worker outcomes.
F2Parallel execution is slower or more expensive than the single agent on simple questions.
Inspect
Break down scheduling, provider queue, calls, tokens, context duplication, synthesis, and retries by task class.
Likely cause
The manager launches a fixed team without considering task complexity or independent work.
Repair
Route direct facts to deterministic or single-agent paths and reserve workers for breadth-first tasks.
Prevent next time
Maintain a cost and latency downgrade threshold in the router.
F3A specialist performs an action that only the manager should authorize.
Inspect
Review worker credentials, exposed tools, handoff history, and shared execution identity.
Likely cause
Workers inherited the manager's capability set or a shared broad credential.
Repair
Issue role-specific tool allowlists and identities; keep approval and side effects at the manager boundary.
Prevent next time
Run permission-denial tests per role and fail on capability drift.
F4The evaluator prefers one architecture's writing style and masks factual contradictions.
Inspect
Separate deterministic citation and contradiction checks from model-judge style scores; audit disagreements manually.
Likely cause
One broad judge rubric conflates presentation with factual and operational outcomes.
Repair
Grade evidence and state deterministically, then calibrate a narrow judge on passing outputs.
Prevent next time
Keep human labels and architecture-blind evaluation samples in the regression suite.

Beyond the demo

Production boundary

  1. 01Require a single-agent or deterministic baseline before approving a multi-agent split.
  2. 02Use workers only for distinct capabilities, policies, context, ownership, or parallel branches.
  3. 03Define typed request, result, evidence, limitation, and failure contracts for every worker.
  4. 04Keep one global owner for budgets, permissions, termination, trace correlation, and final release.
  5. 05Give each role least-privilege credentials and prohibit recursive delegation unless separately bounded.
  6. 06Deduplicate evidence and validate branch coverage before synthesis.
  7. 07Evaluate repeated trials for quality, latency, cost, duplication, contradictions, and reviewer burden.
  8. 08Document downgrade, cancellation, partial-result, and single-agent fallback behavior.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Orchestration and handoffs

    OpenAI · Official documentation · 2026-08-20

    handoffs · agents as tools · single-agent baseline
  2. [2]
    How we built our multi-agent research system

    Anthropic · Public case · 2026-08-20

    breadth-first delegation · parallel research · multi-agent cost constraints
  3. [3]
    Building effective agents

    Anthropic · Official documentation · 2026-08-20

    workflow and agent definitions · architecture patterns · cost and control tradeoffs
  4. [4]
    Demystifying evals for AI agents

    Anthropic · Official documentation · 2026-08-20

    outcome and trajectory grading · repeated trials · isolated eval environments

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.