Skip to main content

For engineers, technical operators, and AI product owners

Agent Building

Build an agent you can test, interrupt, inspect, and operate.

Twelve field modules take one marketing research problem through typed model calls, tools, state, retrieval, MCP, orchestration, review, evaluation, and a staged production decision. Every module produces an artifact that the next module can reuse.

Open module 00
Output
A bounded marketing research agent with claim-level evidence, explicit permissions, resumable human review, a regression suite, an operating cost model, and a documented go or no-go decision.
Effort
18 to 24 hours plus the capstone
Before you start
Comfort with HTTP, JSON, TypeScript, environment variables, and automated tests. Use only synthetic or approved non-sensitive data in the labs.

Foundation → lab → project

The ordered field path.

Each lesson produces an artifact, tests it against observable criteria, then carries it into the next decision. Enter where your workflow is stuck or follow the complete sequence.

  1. 00AI agents vs workflows: earn the loopCompare three implementations of the same research intake task and keep the least autonomous design that clears the acceptance gate.Foundation · 75 min
  2. 01LLM APIs and structured outputs: build a typed boundaryImplement a support classifier that treats valid JSON, schema adherence, refusal, truncation, timeout, and application validation as separate states.Lab · 105 min
  3. 02Tool use and function calling: keep execution outside the modelGive a model one narrow account lookup tool, then prove that identity, tenancy, argument validation, execution, and audit remain application responsibilities.Lab · 120 min
  4. 03Context, memory, and state: make the job resumableSeparate the model's current working context from authoritative job state and selectively retained memory, then test pause, resume, expiry, deletion, and contamination.Lab · 125 min
  5. 04RAG and knowledge retrieval: test the evidence layer firstBuild a two-tenant retrieval fixture, apply access controls before ranking, compare retrieval variants, and grade citations separately from answer prose.Lab · 150 min
  6. 05MCP and external systems: treat interoperability as a trust boundaryExpose one read-only knowledge lookup through MCP, bind it to an authenticated tenant, and test approval, server failure, malicious results, and removal.Lab · 140 min
  7. 06Agent loops and orchestration: make progress observableImplement an external state machine with budgets, checkpoints, repetition detection, terminal states, and an escalation packet for a read-only research job.Lab · 150 min
  8. 07Multi-agent patterns: compare before you splitBuild a breadth-first research variant beside the single-agent baseline and keep it only when independent work improves measured coverage enough to justify coordination cost.Lab · 165 min
  9. 08Human review and guardrails: make approval resumablePlace policy checks beside side effects, persist the approval packet, and resume the exact run after an authorized reviewer decides.Lab · 105 min
  10. 09Evals and observability: grade outcomes and tracesTurn real failure shapes into an isolated, repeatable suite that grades final state, evidence, policy behavior, and execution traces.Lab · 120 min
  11. 10Production security and cost: rehearse the operating boundaryBuild a staged readiness gate around identity, secrets, data handling, budgets, incidents, rollback, and accountable ownership.Lab · 120 min
  12. 11Capstone: build an evidence-bound marketing research agentIntegrate typed planning, permission-filtered retrieval, bounded tools, resumable review, evals, and an operating decision into one inspectable system.Project · 5 to 8 hours

Hands-on projects

Build evidence of operating ability.

PROJECT P1

Typed agent boundary

Combine a strict output contract, one read-only tool, authoritative state, and a fixture runner before adding an autonomous loop.

Output: TypeScript package with schemas, fake adapters, tool authorization, state fixtures, tests, and a short architecture decision record.

PROJECT P2

Production research agent

Apply retrieval, MCP, bounded orchestration, human approval, evals, and production controls to the capstone corpus.

Output: Staging-ready reference implementation, evidence ledger, review packet, evaluation report, cost worksheet, runbook, and go or no-go record.

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.