On this page
Learning objectives
- Integrate course artifacts behind one typed run contract and evidence ledger
- Produce claim-level citations and explicit unknowns from permission-filtered sources
- Prove stopping, review, failure, and recovery behavior with a locked evaluation suite
- Present architecture, quality, cost, risk, and scope as a defensible pilot decision
Before you start
- • Completed artifacts from modules 0 through 10 or equivalent tested implementations
- • A synthetic or approved research corpus with source IDs, permissions, freshness, and provenance
- • A non-production environment, fake adapters, and reviewers for evidence and operating decisions
Working definition
Research agent capstone
The capstone is a bounded research service that accepts a structured question, creates a reviewable plan, gathers evidence through approved read-only tools, records claim-to-source links, stops on explicit conditions, and produces a brief with limitations. Its delivery includes tests, traces, evaluation results, cost model, runbook, and scope decision. The output document alone is not the product.
Marketing research combines open-ended discovery with strict evidence needs. It is a credible agent task only when search breadth, source policy, claim support, permission boundaries, and reviewer ownership are visible in the implementation.
Capstones often collapse into one impressive transcript. A reusable service needs a stable contract, versioned corpus, deterministic gates, failure cases, and replayable traces so another engineer can inspect a decision without trusting the demo narrator.
A recommendation may influence spend or positioning even when the system cannot write to external platforms. Evidence quality, uncertainty labels, and human sign-off are therefore operational controls, not cosmetic citations.
Field situation
Mariner category-entry research
Named synthetic scenario. Mariner Cloud, its market, documents, research questions, metrics, evaluations, and decisions are course fixtures. The worked result is not commercial advice or a Tenten client engagement.
- Owner
- You are the lead engineer and research operations owner for a B2B product marketing team.
- Decision
- Decide whether the evidence supports a limited discovery sprint, requires more research, or rejects the category thesis, and state which facts would change that decision.
- Starting state
- Mariner is considering a new workflow-automation category. The approved corpus has 48 synthetic interview excerpts, 12 product and pricing pages captured with dates, eight analyst-style notes, and six internal win-loss summaries. Ten documents are restricted to product leaders. A 30-case regression suite and fake read-only CRM are available.
- Expected outcome
- The system produces a concise evidence-bound brief, source ledger, unknowns, counterevidence, review record, trace, eval report, and scoped pilot decision. It abstains when minimum evidence coverage is not met.
Constraints
- • The service can search, retrieve, calculate, and draft; it cannot publish, contact prospects, change CRM records, or purchase data.
- • Every factual sentence in the final brief must cite one or more visible source IDs or be labeled inference, hypothesis, or unknown.
- • Retrieval applies actor permissions before ranking, and restricted passages cannot appear in traces or outputs for unauthorized reviewers.
- • Each run has at most ten model turns, sixteen tool calls, two parallel research branches, 120 seconds, and a fixture budget of USD 2.00.
- • An analyst must approve the evidence packet and final recommendation; approval never authorizes an external side effect.
Worked example
Mariner narrows a claim after counterevidence
Evidence status: Named synthetic scenarioThe planner decomposes the fixture question into buyer pain, switching trigger, alternative, and proof thresholds. One branch retrieves six approved interview excerpts suggesting reporting delays; another finds current pricing pages and two win-loss notes. A draft states that mid-market buyers consistently replace spreadsheets after weekly reporting failures. The claim ledger shows support from four interviews, while two excerpts describe spreadsheet use as adequate and the win-loss corpus is too small for a prevalence claim.
The verifier rejects the word consistently, marks prevalence unknown, and rewrites the finding as a bounded pattern observed in four cited interviews. It adds the two counterexamples and recommends a ten-interview discovery sprint focused on reporting cadence rather than a category launch. The analyst approves the revised evidence packet, records that interview selection may be biased, and declines to approve the original broad claim.
The final synthetic artifact has no uncited factual sentences. Each claim links to source ID, captured date, access class, supporting excerpt, counterevidence, and confidence rationale. The run stops after eight turns and twelve tool calls, remains under its fixture budget, and produces no external effect. The regression suite passes critical gates; the recommendation remains a research decision, not proof of market demand.
Limits
All Mariner evidence and results are fabricated. The Triple Whale source illustrates a real vendor-reported analytics-agent architecture and outcomes, but it does not validate this capstone or transfer its metrics. Anthropic's multi-agent figures are internal to its reported setup. The capstone must report its own corpus, permissions, model and tool versions, trial counts, errors, and reviewer decisions without implying general performance.
Method
Build it, with checkpoints
Field situation
Decide whether the evidence supports a limited discovery sprint, requires more research, or rejects the category thesis, and state which facts would change that decision.
- 01Lock the question and architecture
- 02Build permission-bearing retrieval
- 03Implement the bounded research loop
Acceptance checks
The capstone repository reproduces a reviewed evidence-bound brief and all critical gates from one documented command. The system abstains on insufficient evidence, denies unauthorized retrieval, stops on budgets, resumes safely, and produces no external side effect. The delivery ends with a scoped decision rather than a universal readiness claim.
- 01
Lock the question and architecture
Write the structured intake, decision to support, allowed evidence, prohibited actions, budgets, reviewer, and abstention rule. Compare deterministic, single-agent, and optional two-branch designs on the same small baseline before selecting the control flow.
CHECKPOINT · The architecture record names the simplest passing design, rejected alternatives, measurable reasons, permissions, stop conditions, and downgrade path.
- 02
Build permission-bearing retrieval
Ingest the 74 fixture documents with source ID, capture date, owner, access class, document context, and chunk provenance. Apply actor filters before semantic or lexical ranking, then benchmark required-source recall and citation correctness on the gold questions.
CHECKPOINT · Unauthorized actors retrieve zero restricted chunks, every returned passage preserves provenance, and the retrieval report separates recall from answer quality.
- 03
Implement the bounded research loop
Use typed tool contracts, application-owned state, explicit stage and budget checks, at most two independent branches, evidence deduplication, and terminal reason codes. Persist compact notes and source IDs instead of replaying an unbounded transcript.
CHECKPOINT · Fixtures prove turn, call, branch, time, and cost limits; a restart resumes from stored stage; and no path reaches an unregistered or write-capable tool.
- 04
Verify claims before prose release
Generate a claim ledger first. Check source existence, permissions, captured date, entailment, conflicts, inference labels, and missing evidence. Require an analyst to approve the ledger and limitations before rendering the final brief.
CHECKPOINT · A script reports zero uncited factual claims, and seeded unsupported, stale, contradictory, and unauthorized claims are blocked with distinct codes.
- 05
Evaluate and attack the system
Run baseline and candidate three times over the 30-case suite. Include prompt injection, poisoned metadata, missing documents, malformed tool results, connector outage, stale approval, budget exhaustion, duplicate resume, and ambiguous evidence. Review critical traces manually.
CHECKPOINT · Every critical gate passes or the decision is no-go; the report includes configuration, category slices, consistency, tails, judge calibration, incidents, and unresolved failures.
- 06
Present the operating decision
Demonstrate one normal run, one abstention, one permission denial, one restart, and one kill-switch event. Present unit economics, reviewer burden, residual risks, canary scope, owners, monitoring, runbook, rollback, and the facts that would justify expansion.
CHECKPOINT · A reviewer who did not build the system can reproduce the demo from fixtures and sign a scoped pilot, conditional hold, or no-go memo from linked evidence.
Hands-on lab
Ship the evidence packet and pilot record
Build the Mariner read-only research service, evaluate it against locked fixtures, conduct adversarial review, and present a limited pilot or no-go decision with complete artifacts.
Prepare
- • Fork the starter into a repository with lint, typecheck, unit tests, fixture tests, and one command that runs the locked eval suite.
- • Document corpus ownership and remove any source that lacks permission, capture date, or stable identifier.
- • Assign separate engineering, evidence-review, and operational-decision roles, even if one learner simulates them sequentially.
Deliverable
A runnable read-only repository, architecture record, corpus manifest, threat model, evidence ledger, reviewed brief, 30-case repeated-trial report, cost worksheet, five incident records, runbook, demo script, and signed pilot or no-go memo.
Starter kit: Capstone run contract
TypeScripttype Evidence = {
sourceId: string;
capturedAt: string;
accessClass: "public" | "approved-internal" | "product-leads";
excerpt: string;
supports: string[];
contradicts: string[];
};
type Claim = {
id: string;
text: string;
status: "supported" | "inference" | "hypothesis" | "unknown";
sourceIds: string[];
limitation: string;
};
type ResearchRun = {
runId: string;
actorId: string;
question: string;
corpusVersion: "mariner-corpus-v1";
policyVersion: "research-read-v1";
budgets: { turns: 10; toolCalls: 16; parallelBranches: 2; wallMs: 120000; usd: 2 };
evidence: Evidence[];
claims: Claim[];
openQuestions: string[];
terminal: "complete" | "needs_evidence" | "review_denied" | "budget_stop" | "failed";
};
const releaseGates = {
uncitedFactualClaims: 0,
unauthorizedSourceExposures: 0,
prohibitedEffects: 0,
nonterminalRuns: 0,
taskSuccessRateMin: 0.90,
};
// Required stages: validate intake, plan, retrieve with permissions,
// assemble evidence, draft claims, verify claims, request review, finalize.
// Persist an event after each stage and recheck budgets before the next one.Expected result
The capstone repository reproduces a reviewed evidence-bound brief and all critical gates from one documented command. The system abstains on insufficient evidence, denies unauthorized retrieval, stops on budgets, resumes safely, and produces no external side effect. The delivery ends with a scoped decision rather than a universal readiness claim.
Carry forward
Keep the repository and evidence pack as the reference implementation. For a real engagement, replace fixtures only after data authorization, connector threat review, provider-semantics verification, target-environment evals, incident drills, and accountable business approval.
Acceptance checks
- 01Every factual sentence in the brief resolves to permitted source IDs, and inference, hypothesis, counterevidence, and unknowns are explicit.
- 02The runner enforces permissions, registered read-only tools, budgets, terminal states, approval binding, idempotency, and restart behavior with observable tests.
- 03Thirty cases run three times with zero critical failures, or the release memo records no-go without weakening the locked gates.
- 04The repository includes the architecture record, corpus manifest, threat model, trace samples, eval report, cost model, incident pack, runbook, review packet, and signed scope decision.
What breaks
Failure clinic
F1The brief contains citations that do not support the nearby claim.
- Inspect
- Open each cited excerpt, compare entailment and scope, and inspect where the claim changed after retrieval.
- Likely cause
- The writer attaches plausible sources after drafting rather than building prose from a verified claim ledger.
- Repair
- Reject the claim, reconstruct it from supporting passages, and expose contradiction or uncertainty.
- Prevent next time
- Make claim-source validation a blocking stage before document rendering and human approval.
F2An unauthorized reviewer can infer restricted material from the final summary.
- Inspect
- Trace source access, cached notes, branch synthesis, model context, and output under the reviewer's identity.
- Likely cause
- Permission filtering occurs after retrieval or restricted findings leak through shared state.
- Repair
- Purge the run, filter before ranking, isolate branch state, and regenerate under the actual actor policy.
- Prevent next time
- Test direct and inferential leakage with identity-specific fixtures on every corpus or prompt change.
F3The agent keeps searching because every new source suggests another question.
- Inspect
- Review coverage criteria, marginal evidence gain, repeated queries, budgets, and terminal-reason evaluation.
- Likely cause
- The planner has a broad goal but no evidence threshold, diminishing-return rule, or application stop.
- Repair
- Stop the run, surface open questions, and require a new approved scope for further research.
- Prevent next time
- Define claim-level evidence targets, branch and turn ceilings, deduplication, and explicit abstention.
F4The demonstration passes, but another engineer cannot reproduce it.
- Inspect
- Compare corpus, model, tool, policy, prompt, environment, seed, clock, and uncommitted local state.
- Likely cause
- The result depends on hidden setup, live mutable data, or manually selected favorable output.
- Repair
- Freeze fixtures and manifests, script the run, retain all trials, and document unavoidable external variance.
- Prevent next time
- Require clean-environment reproduction and artifact hashes before the final review.
Beyond the demo
Production boundary
- 01Name the research decision, evidence policy, abstention rule, reviewer, and prohibited actions.
- 02Filter source permissions before retrieval and preserve provenance through chunks, notes, claims, and output.
- 03Use typed inputs, tools, state, evidence, claims, approvals, terminal reasons, and error records.
- 04Enforce turns, calls, branches, time, bytes, and spend in application code.
- 05Grade retrieval, claims, outcomes, traces, policies, effects, consistency, latency, cost, and reviewer burden.
- 06Persist correlated redacted traces, versioned artifacts, review decisions, and reproducible release evidence.
- 07Rehearse injection, outage, stale data, malformed results, unknown effects, revocation, rollback, and manual fallback.
- 08Launch only within signed scope, with named ownership, kill authority, monitoring, review date, and expansion gates.
Evidence status
Sources and claim limits
Sources support the named claims; they do not guarantee the same result in another system.
- [1]Structured model outputsschema adherence · refusal handling · incomplete output handling
OpenAI · Official documentation · 2026-08-20
- [2]Function callingstrict function schemas · parallel calls · tool-call state machine
OpenAI · Official documentation · 2026-08-20
- [3]Evaluate agent workflowstrace grading · datasets · repeatable eval runs
OpenAI · Official documentation · 2026-08-20
- [4]How we built our multi-agent research systembreadth-first delegation · parallel research · multi-agent cost constraints
Anthropic · Public case · 2026-08-20
- [5]Demystifying evals for AI agentsoutcome and trajectory grading · repeated trials · isolated eval environments
Anthropic · Official documentation · 2026-08-20
- [6]Triple Whale drives business growth with Claudemarketing analytics agents · specialized analytics tools · vendor-reported outcomes
Anthropic / Claude · Public case · 2026-08-20
Related Tenten resources