Skip to main content

build marketing research AI agent project

Project

Capstone: build an evidence-bound marketing research agent

Integrate typed planning, permission-filtered retrieval, bounded tools, resumable review, evals, and an operating decision into one inspectable system.

DIFFICULTY
Advanced
ESTIMATED TIME
5 to 8 hours
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Integrate course artifacts behind one typed run contract and evidence ledger
  • Produce claim-level citations and explicit unknowns from permission-filtered sources
  • Prove stopping, review, failure, and recovery behavior with a locked evaluation suite
  • Present architecture, quality, cost, risk, and scope as a defensible pilot decision

Before you start

  • • Completed artifacts from modules 0 through 10 or equivalent tested implementations
  • • A synthetic or approved research corpus with source IDs, permissions, freshness, and provenance
  • • A non-production environment, fake adapters, and reviewers for evidence and operating decisions

Working definition

Research agent capstone

The capstone is a bounded research service that accepts a structured question, creates a reviewable plan, gathers evidence through approved read-only tools, records claim-to-source links, stops on explicit conditions, and produces a brief with limitations. Its delivery includes tests, traces, evaluation results, cost model, runbook, and scope decision. The output document alone is not the product.

Marketing research combines open-ended discovery with strict evidence needs. It is a credible agent task only when search breadth, source policy, claim support, permission boundaries, and reviewer ownership are visible in the implementation.

Capstones often collapse into one impressive transcript. A reusable service needs a stable contract, versioned corpus, deterministic gates, failure cases, and replayable traces so another engineer can inspect a decision without trusting the demo narrator.

A recommendation may influence spend or positioning even when the system cannot write to external platforms. Evidence quality, uncertainty labels, and human sign-off are therefore operational controls, not cosmetic citations.

Field situation

Mariner category-entry research

Named synthetic scenario. Mariner Cloud, its market, documents, research questions, metrics, evaluations, and decisions are course fixtures. The worked result is not commercial advice or a Tenten client engagement.

Owner
You are the lead engineer and research operations owner for a B2B product marketing team.
Decision
Decide whether the evidence supports a limited discovery sprint, requires more research, or rejects the category thesis, and state which facts would change that decision.
Starting state
Mariner is considering a new workflow-automation category. The approved corpus has 48 synthetic interview excerpts, 12 product and pricing pages captured with dates, eight analyst-style notes, and six internal win-loss summaries. Ten documents are restricted to product leaders. A 30-case regression suite and fake read-only CRM are available.
Expected outcome
The system produces a concise evidence-bound brief, source ledger, unknowns, counterevidence, review record, trace, eval report, and scoped pilot decision. It abstains when minimum evidence coverage is not met.

Constraints

  • • The service can search, retrieve, calculate, and draft; it cannot publish, contact prospects, change CRM records, or purchase data.
  • • Every factual sentence in the final brief must cite one or more visible source IDs or be labeled inference, hypothesis, or unknown.
  • • Retrieval applies actor permissions before ranking, and restricted passages cannot appear in traces or outputs for unauthorized reviewers.
  • • Each run has at most ten model turns, sixteen tool calls, two parallel research branches, 120 seconds, and a fixture budget of USD 2.00.
  • • An analyst must approve the evidence packet and final recommendation; approval never authorizes an external side effect.

Worked example

Mariner narrows a claim after counterevidence

Evidence status: Named synthetic scenario

The planner decomposes the fixture question into buyer pain, switching trigger, alternative, and proof thresholds. One branch retrieves six approved interview excerpts suggesting reporting delays; another finds current pricing pages and two win-loss notes. A draft states that mid-market buyers consistently replace spreadsheets after weekly reporting failures. The claim ledger shows support from four interviews, while two excerpts describe spreadsheet use as adequate and the win-loss corpus is too small for a prevalence claim.

The verifier rejects the word consistently, marks prevalence unknown, and rewrites the finding as a bounded pattern observed in four cited interviews. It adds the two counterexamples and recommends a ten-interview discovery sprint focused on reporting cadence rather than a category launch. The analyst approves the revised evidence packet, records that interview selection may be biased, and declines to approve the original broad claim.

The final synthetic artifact has no uncited factual sentences. Each claim links to source ID, captured date, access class, supporting excerpt, counterevidence, and confidence rationale. The run stops after eight turns and twelve tool calls, remains under its fixture budget, and produces no external effect. The regression suite passes critical gates; the recommendation remains a research decision, not proof of market demand.

Limits

All Mariner evidence and results are fabricated. The Triple Whale source illustrates a real vendor-reported analytics-agent architecture and outcomes, but it does not validate this capstone or transfer its metrics. Anthropic's multi-agent figures are internal to its reported setup. The capstone must report its own corpus, permissions, model and tool versions, trial counts, errors, and reviewer decisions without implying general performance.

Method

Build it, with checkpoints

Mariner architecture from structured intake through permission-filtered retrieval, two bounded research branches, evidence ledger, claim verifier, analyst review, final brief, eval gate, telemetry, and kill control.

Field situation

Decide whether the evidence supports a limited discovery sprint, requires more research, or rejects the category thesis, and state which facts would change that decision.

  1. 01Lock the question and architecture
  2. 02Build permission-bearing retrieval
  3. 03Implement the bounded research loop

Acceptance checks

The capstone repository reproduces a reviewed evidence-bound brief and all critical gates from one documented command. The system abstains on insufficient evidence, denies unauthorized retrieval, stops on budgets, resumes safely, and produces no external side effect. The delivery ends with a scoped decision rather than a universal readiness claim.

Why this visualRender the final architecture from the versioned service manifest and produce the evidence-coverage chart from the claim ledger. A decorative generated image cannot communicate permission paths, gates, or unresolved evidence.
  1. 01

    Lock the question and architecture

    Write the structured intake, decision to support, allowed evidence, prohibited actions, budgets, reviewer, and abstention rule. Compare deterministic, single-agent, and optional two-branch designs on the same small baseline before selecting the control flow.

    CHECKPOINT · The architecture record names the simplest passing design, rejected alternatives, measurable reasons, permissions, stop conditions, and downgrade path.

  2. 02

    Build permission-bearing retrieval

    Ingest the 74 fixture documents with source ID, capture date, owner, access class, document context, and chunk provenance. Apply actor filters before semantic or lexical ranking, then benchmark required-source recall and citation correctness on the gold questions.

    CHECKPOINT · Unauthorized actors retrieve zero restricted chunks, every returned passage preserves provenance, and the retrieval report separates recall from answer quality.

  3. 03

    Implement the bounded research loop

    Use typed tool contracts, application-owned state, explicit stage and budget checks, at most two independent branches, evidence deduplication, and terminal reason codes. Persist compact notes and source IDs instead of replaying an unbounded transcript.

    CHECKPOINT · Fixtures prove turn, call, branch, time, and cost limits; a restart resumes from stored stage; and no path reaches an unregistered or write-capable tool.

  4. 04

    Verify claims before prose release

    Generate a claim ledger first. Check source existence, permissions, captured date, entailment, conflicts, inference labels, and missing evidence. Require an analyst to approve the ledger and limitations before rendering the final brief.

    CHECKPOINT · A script reports zero uncited factual claims, and seeded unsupported, stale, contradictory, and unauthorized claims are blocked with distinct codes.

  5. 05

    Evaluate and attack the system

    Run baseline and candidate three times over the 30-case suite. Include prompt injection, poisoned metadata, missing documents, malformed tool results, connector outage, stale approval, budget exhaustion, duplicate resume, and ambiguous evidence. Review critical traces manually.

    CHECKPOINT · Every critical gate passes or the decision is no-go; the report includes configuration, category slices, consistency, tails, judge calibration, incidents, and unresolved failures.

  6. 06

    Present the operating decision

    Demonstrate one normal run, one abstention, one permission denial, one restart, and one kill-switch event. Present unit economics, reviewer burden, residual risks, canary scope, owners, monitoring, runbook, rollback, and the facts that would justify expansion.

    CHECKPOINT · A reviewer who did not build the system can reproduce the demo from fixtures and sign a scoped pilot, conditional hold, or no-go memo from linked evidence.

Hands-on lab

Ship the evidence packet and pilot record

Build the Mariner read-only research service, evaluate it against locked fixtures, conduct adversarial review, and present a limited pilot or no-go decision with complete artifacts.

Prepare

  • • Fork the starter into a repository with lint, typecheck, unit tests, fixture tests, and one command that runs the locked eval suite.
  • • Document corpus ownership and remove any source that lacks permission, capture date, or stable identifier.
  • • Assign separate engineering, evidence-review, and operational-decision roles, even if one learner simulates them sequentially.

Deliverable

A runnable read-only repository, architecture record, corpus manifest, threat model, evidence ledger, reviewed brief, 30-case repeated-trial report, cost worksheet, five incident records, runbook, demo script, and signed pilot or no-go memo.

Starter kit: Capstone run contract

TypeScript
type Evidence = {
  sourceId: string;
  capturedAt: string;
  accessClass: "public" | "approved-internal" | "product-leads";
  excerpt: string;
  supports: string[];
  contradicts: string[];
};

type Claim = {
  id: string;
  text: string;
  status: "supported" | "inference" | "hypothesis" | "unknown";
  sourceIds: string[];
  limitation: string;
};

type ResearchRun = {
  runId: string;
  actorId: string;
  question: string;
  corpusVersion: "mariner-corpus-v1";
  policyVersion: "research-read-v1";
  budgets: { turns: 10; toolCalls: 16; parallelBranches: 2; wallMs: 120000; usd: 2 };
  evidence: Evidence[];
  claims: Claim[];
  openQuestions: string[];
  terminal: "complete" | "needs_evidence" | "review_denied" | "budget_stop" | "failed";
};

const releaseGates = {
  uncitedFactualClaims: 0,
  unauthorizedSourceExposures: 0,
  prohibitedEffects: 0,
  nonterminalRuns: 0,
  taskSuccessRateMin: 0.90,
};

// Required stages: validate intake, plan, retrieve with permissions,
// assemble evidence, draft claims, verify claims, request review, finalize.
// Persist an event after each stage and recheck budgets before the next one.

Expected result

The capstone repository reproduces a reviewed evidence-bound brief and all critical gates from one documented command. The system abstains on insufficient evidence, denies unauthorized retrieval, stops on budgets, resumes safely, and produces no external side effect. The delivery ends with a scoped decision rather than a universal readiness claim.

Carry forward

Keep the repository and evidence pack as the reference implementation. For a real engagement, replace fixtures only after data authorization, connector threat review, provider-semantics verification, target-environment evals, incident drills, and accountable business approval.

Acceptance checks

  1. 01Every factual sentence in the brief resolves to permitted source IDs, and inference, hypothesis, counterevidence, and unknowns are explicit.
  2. 02The runner enforces permissions, registered read-only tools, budgets, terminal states, approval binding, idempotency, and restart behavior with observable tests.
  3. 03Thirty cases run three times with zero critical failures, or the release memo records no-go without weakening the locked gates.
  4. 04The repository includes the architecture record, corpus manifest, threat model, trace samples, eval report, cost model, incident pack, runbook, review packet, and signed scope decision.

What breaks

Failure clinic

F1The brief contains citations that do not support the nearby claim.
Inspect
Open each cited excerpt, compare entailment and scope, and inspect where the claim changed after retrieval.
Likely cause
The writer attaches plausible sources after drafting rather than building prose from a verified claim ledger.
Repair
Reject the claim, reconstruct it from supporting passages, and expose contradiction or uncertainty.
Prevent next time
Make claim-source validation a blocking stage before document rendering and human approval.
F2An unauthorized reviewer can infer restricted material from the final summary.
Inspect
Trace source access, cached notes, branch synthesis, model context, and output under the reviewer's identity.
Likely cause
Permission filtering occurs after retrieval or restricted findings leak through shared state.
Repair
Purge the run, filter before ranking, isolate branch state, and regenerate under the actual actor policy.
Prevent next time
Test direct and inferential leakage with identity-specific fixtures on every corpus or prompt change.
F3The agent keeps searching because every new source suggests another question.
Inspect
Review coverage criteria, marginal evidence gain, repeated queries, budgets, and terminal-reason evaluation.
Likely cause
The planner has a broad goal but no evidence threshold, diminishing-return rule, or application stop.
Repair
Stop the run, surface open questions, and require a new approved scope for further research.
Prevent next time
Define claim-level evidence targets, branch and turn ceilings, deduplication, and explicit abstention.
F4The demonstration passes, but another engineer cannot reproduce it.
Inspect
Compare corpus, model, tool, policy, prompt, environment, seed, clock, and uncommitted local state.
Likely cause
The result depends on hidden setup, live mutable data, or manually selected favorable output.
Repair
Freeze fixtures and manifests, script the run, retain all trials, and document unavoidable external variance.
Prevent next time
Require clean-environment reproduction and artifact hashes before the final review.

Beyond the demo

Production boundary

  1. 01Name the research decision, evidence policy, abstention rule, reviewer, and prohibited actions.
  2. 02Filter source permissions before retrieval and preserve provenance through chunks, notes, claims, and output.
  3. 03Use typed inputs, tools, state, evidence, claims, approvals, terminal reasons, and error records.
  4. 04Enforce turns, calls, branches, time, bytes, and spend in application code.
  5. 05Grade retrieval, claims, outcomes, traces, policies, effects, consistency, latency, cost, and reviewer burden.
  6. 06Persist correlated redacted traces, versioned artifacts, review decisions, and reproducible release evidence.
  7. 07Rehearse injection, outage, stale data, malformed results, unknown effects, revocation, rollback, and manual fallback.
  8. 08Launch only within signed scope, with named ownership, kill authority, monitoring, review date, and expansion gates.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Structured model outputs

    OpenAI · Official documentation · 2026-08-20

    schema adherence · refusal handling · incomplete output handling
  2. [2]
    Function calling

    OpenAI · Official documentation · 2026-08-20

    strict function schemas · parallel calls · tool-call state machine
  3. [3]
    Evaluate agent workflows

    OpenAI · Official documentation · 2026-08-20

    trace grading · datasets · repeatable eval runs
  4. [4]
    How we built our multi-agent research system

    Anthropic · Public case · 2026-08-20

    breadth-first delegation · parallel research · multi-agent cost constraints
  5. [5]
    Demystifying evals for AI agents

    Anthropic · Official documentation · 2026-08-20

    outcome and trajectory grading · repeated trials · isolated eval environments
  6. [6]
    Triple Whale drives business growth with Claude

    Anthropic / Claude · Public case · 2026-08-20

    marketing analytics agents · specialized analytics tools · vendor-reported outcomes

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.