On this page
Learning objectives
- Decide when variable planning warrants an agent rather than fixed automation
- Write a mandate with tools, evidence rules, effort limits, and termination conditions
- Stage autonomy from recommendation to reversible execution
- Design approvals and traces that let reviewers make an informed decision
Before you start
- • A working bounded automation and a manual fallback
- • Versioned tools, policies, and representative evaluation cases
- • Named business, technical, risk, and incident owners
Working definition
Marketing agents
A marketing agent is a system that uses a model to choose among approved tools and intermediate steps toward a defined goal. It operates inside data, action, budget, time, and review limits. Fixed routes should remain conventional automation. Agent autonomy is earned per action through evidence, not granted to the whole workflow because a prototype appears capable.
Variable research paths can benefit from planning and tool choice, but the same flexibility makes behavior harder to predict. A narrow mandate gives evaluation and incident review a stable boundary.
An approval button is weak if the reviewer cannot see the proposed change, supporting evidence, uncertainty, cost, policy checks, and downstream effect. Approval quality depends on the packet.
Tool descriptions are control surfaces. They should state when a tool applies, required evidence, parameter bounds, side effects, and refusal conditions rather than merely naming an API.
Recommendation mode is useful production work. It creates traces and reviewer decisions that can justify, or reject, later permission for reversible actions.
Field situation
LumenFleet campaign research agent
Named synthetic scenario. LumenFleet is fictional and the agent remains in recommendation mode throughout the lesson.
- Owner
- You are the product owner for an agent that assembles evidence and proposes a campaign brief for fleet electrification software.
- Decision
- Specify whether the agent may recommend a brief for human approval and what evaluation evidence would be required before adding any reversible action.
- Starting state
- The prototype can search the web, read an approved document store, draft a brief, create a project task, and send a chat message. Its instruction simply says research the market and prepare a great campaign.
- Expected outcome
- An agent charter, tool policy, approval packet, trace schema, evaluation set, and incident stop rule.
Constraints
- • Search is limited to approved domains and logged queries
- • Customer files are read-only and retrieved passages carry access controls
- • The agent may propose a brief but cannot publish, message prospects, or change spend
- • Each run has tool-call, time, and cost ceilings plus an emergency stop
Worked example
Stopping when regional evidence is insufficient
Evidence status: Named synthetic scenarioLumenFleet asks the agent for a Singapore campaign brief. Search returns global reports and two local regulatory pages, but no approved local customer evidence. The prototype can still write a convincing brief by blending global themes.
The mandate requires two independent regional sources for pivotal market claims and forbids converting global evidence into local prevalence. The agent retrieves, records the gap, and returns an approval packet recommending targeted interviews rather than a launch brief. It stops after twelve tool calls or when the evidence rule fails. The reviewer can accept the research action, request a narrower brief, or reject the run.
The successful output is a justified refusal to overstate local knowledge. The trace shows queries, retrieved evidence IDs, the unmet rule, effort used, and proposed next step. No campaign or market result is claimed.
Limits
The scenario and evaluation are synthetic. A tool-call cap does not by itself control risk, retrieved sources can still be wrong, and human approval can become superficial under high volume. Anthropic's article offers engineering patterns, not a certification of this agent design.
Method
Build it, with checkpoints
Field situation
Specify whether the agent may recommend a brief for human approval and what evaluation evidence would be required before adding any reversible action.
- 01Prove that planning is needed
- 02Write the mandate and stop conditions
- 03Design narrow tools and trace fields
Acceptance checks
A recommendation agent whose useful behavior includes stopping, refusing, and escalating. Its tools and traces make the proposal reviewable, and any future autonomy increase is tied to task-specific evidence and a narrow action class.
- 01
Prove that planning is needed
Compare the task with the fixed workflow from the automation module. Identify where the next source or step depends on evidence found during the run. Keep stable validation and routing in code.
CHECKPOINT · The charter identifies at least one genuine variable decision and explains why a fixed sequence is insufficient.
- 02
Write the mandate and stop conditions
Define trigger, goal, allowed data, tools, evidence rules, prohibited actions, tool-call limit, time, cost, completion, refusal, escalation, and emergency stop. Make ambiguous success unacceptable.
CHECKPOINT · A reviewer can classify any proposed action as allowed, approval-required, or prohibited from the charter alone.
- 03
Design narrow tools and trace fields
Give each tool a clear purpose, typed parameters, authorization scope, result schema, side-effect label, and error behavior. Record model and policy versions, tool arguments, results, evidence IDs, cost, and state transitions.
CHECKPOINT · The agent cannot reach a public, customer, CRM, or spend side effect through any tool or shared credential.
- 04
Build a decision-ready approval packet
Show changed fields, supporting and conflicting evidence, uncertainty, policy checks, run cost, and what happens after approval. Let reviewers narrow scope or request a specific revision alongside approval and rejection.
CHECKPOINT · A reviewer can narrow scope or request a specific revision without opening raw execution logs.
- 05
Evaluate and stage autonomy
Run common, edge, adversarial, missing-evidence, tool-failure, conflicting-source, over-budget, and stop cases. Review traces and define the exact threshold for any future reversible action.
CHECKPOINT · All prohibited-action tests remain blocked, stop cases terminate within bounds, and the release decision stays recommendation-only unless every written threshold passes.
Hands-on lab
Specify and test a campaign research agent
Use LumenFleet with synthetic records. The agent should produce an approval packet or a controlled refusal and have no credential for public or budget actions.
Prepare
- • Complete the same task manually and keep one accepted and one rejected brief
- • Create read-only test tools with approved sample data
- • Name the person authorized to stop the system and resolve incidents
Deliverable
A versioned mandate, three tool contracts, approval packet schema, twelve-case evaluation set, trace review, and staged-autonomy decision.
Starter kit: Agent mandate and approval contract
Copyable YAMLagent: campaign-research-recommender
goal: propose a source-linked campaign brief for one market and audience
trigger: approved research request
allowed_tools: [approved_web_search, read_evidence_store, save_draft_packet]
prohibited_actions: [publish, send_message, create_ad, change_spend, write_crm]
evidence_rules:
pivotal_claims: two_independent_sources
regional_claims: regional_source_required
limits: {tool_calls: 12, elapsed_minutes: 8, run_cost_usd: 2}
stop_conditions: [evidence_gap, policy_conflict, tool_error_limit, budget_limit]
required_output: [proposal_diff, source_ids, uncertainty, policy_checks, cost, next_action]
approval_choices: [approve_recommendation, request_revision, narrow_scope, reject]
emergency_owner: ""
manual_fallback: ""
trace_retention: ""Expected result
A recommendation agent whose useful behavior includes stopping, refusing, and escalating. Its tools and traces make the proposal reviewable, and any future autonomy increase is tied to task-specific evidence and a narrow action class.
Carry forward
Use accepted and rejected runs, tool traces, reviewer reasons, and stop cases as the core evaluation set in the next module. Preserve recommendation mode as a supported deployment stage.
Acceptance checks
- 01The charter states why variable planning is needed and which steps remain deterministic
- 02Allowed tools, prohibited actions, evidence requirements, and hard limits are machine-testable
- 03The agent has no credential path to publication, outreach, CRM mutation, or spend
- 04Approval packets show proposed differences, evidence, uncertainty, checks, cost, and downstream effect
- 05At least twelve evaluation cases cover normal, edge, adversarial, and failure behavior
- 06A kill switch, incident owner, manual fallback, and trace review procedure are rehearsed
What breaks
Failure clinic
F1The agent is slower and more expensive than a fixed route while taking the same steps every time.
- Inspect
- Compare tool sequences and decisions across representative traces.
- Likely cause
- A deterministic workflow was wrapped in agent language without a variable planning need.
- Repair
- Replace the loop with explicit automation and keep the model only for bounded interpretation.
- Prevent next time
- Require evidence of path variability before approving an agent architecture.
F2Reviewers approve quickly but later cannot explain the evidence or changed fields.
- Inspect
- Review the packet, reviewer time, trace access, decision reason, and downstream effect shown.
- Likely cause
- Approval was designed as a button instead of an informed decision.
- Repair
- Pause execution authority and redesign packets around differences, evidence, uncertainty, and consequences.
- Prevent next time
- Test approval comprehension and record structured decision reasons.
F3The agent continues searching after evidence becomes inadequate or the run becomes uneconomic.
- Inspect
- Check tool-call count, elapsed time, cost, repeated queries, and stop-condition evaluation.
- Likely cause
- The goal rewards completion but gives no hard effort or evidence boundary.
- Repair
- Terminate the run, return the evidence gap, and enforce limits outside the model.
- Prevent next time
- Implement hard budgets and explicit refusal outputs in the orchestration layer.
F4A prompt injection in retrieved material changes the agent's plan or requests a forbidden tool.
- Inspect
- Review retrieved content, instruction hierarchy, tool request, policy decision, and untrusted-data treatment.
- Likely cause
- External content was allowed to act as instruction rather than evidence.
- Repair
- Block the action, quarantine the item, narrow tool access, and add the trace to adversarial tests.
- Prevent next time
- Separate instructions from untrusted content and authorize every tool call against policy.
Beyond the demo
Production boundary
- 01The task requires variable planning and has a bounded completion condition
- 02Mandate, evidence rules, prohibited actions, limits, and stop behavior are versioned
- 03Tools use least privilege, typed contracts, side-effect labels, and authorization checks
- 04Untrusted retrieved content cannot override system or workflow policy
- 05Approval packets expose differences, evidence, uncertainty, cost, checks, and consequences
- 06Recommendation mode, dry run, reversible action, and higher-risk action have separate release gates
- 07Evaluation covers typical, edge, adversarial, tool-failure, refusal, and budget cases
- 08Traces support replay and incident review while honoring data minimization and retention
- 09Kill switch, manual fallback, rollback, on-call owner, and escalation are rehearsed
Evidence status
Sources and claim limits
Sources support the named claims; they do not guarantee the same result in another system.
- [1]Building effective agentsworkflow versus agent choice · tool design · human review
Anthropic · Published research · 2026-08-20
- [2]How we built our multi-agent research systemresearch decomposition · citation quality · effort controls
Anthropic · Published research · 2026-08-20
- [3]Guardrails and human reviewapproval boundaries · guardrails · resumable review
OpenAI · Official documentation · 2026-08-20
- [4]Safety in building agentsprompt injection · data leakage · tool approval
OpenAI · Official documentation · 2026-08-20
- [5]Integrate AI into n8n workflowsworkflow structure · execution handling · testing
n8n · Official documentation · 2026-08-20
- [6]Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profilerisk ownership · measurement · human oversight
NIST · Official documentation · 2026-08-20
- [7]AI Agent Building learning trackagent architecture · tool use · evaluation progression
Tenten AI · Tenten field method · 2026-08-20
Related Tenten resources