Skip to main content

AI marketing agents with human approval

Lab

Marketing Agents and Human Approval

Give an agent a narrow mandate, inspectable tools, and approval packets that match the risk of each proposed marketing action.

DIFFICULTY
Advanced
ESTIMATED TIME
110 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Decide when variable planning warrants an agent rather than fixed automation
  • Write a mandate with tools, evidence rules, effort limits, and termination conditions
  • Stage autonomy from recommendation to reversible execution
  • Design approvals and traces that let reviewers make an informed decision

Before you start

  • • A working bounded automation and a manual fallback
  • • Versioned tools, policies, and representative evaluation cases
  • • Named business, technical, risk, and incident owners

Working definition

Marketing agents

A marketing agent is a system that uses a model to choose among approved tools and intermediate steps toward a defined goal. It operates inside data, action, budget, time, and review limits. Fixed routes should remain conventional automation. Agent autonomy is earned per action through evidence, not granted to the whole workflow because a prototype appears capable.

Variable research paths can benefit from planning and tool choice, but the same flexibility makes behavior harder to predict. A narrow mandate gives evaluation and incident review a stable boundary.

An approval button is weak if the reviewer cannot see the proposed change, supporting evidence, uncertainty, cost, policy checks, and downstream effect. Approval quality depends on the packet.

Tool descriptions are control surfaces. They should state when a tool applies, required evidence, parameter bounds, side effects, and refusal conditions rather than merely naming an API.

Recommendation mode is useful production work. It creates traces and reviewer decisions that can justify, or reject, later permission for reversible actions.

Field situation

LumenFleet campaign research agent

Named synthetic scenario. LumenFleet is fictional and the agent remains in recommendation mode throughout the lesson.

Owner
You are the product owner for an agent that assembles evidence and proposes a campaign brief for fleet electrification software.
Decision
Specify whether the agent may recommend a brief for human approval and what evaluation evidence would be required before adding any reversible action.
Starting state
The prototype can search the web, read an approved document store, draft a brief, create a project task, and send a chat message. Its instruction simply says research the market and prepare a great campaign.
Expected outcome
An agent charter, tool policy, approval packet, trace schema, evaluation set, and incident stop rule.

Constraints

  • • Search is limited to approved domains and logged queries
  • • Customer files are read-only and retrieved passages carry access controls
  • • The agent may propose a brief but cannot publish, message prospects, or change spend
  • • Each run has tool-call, time, and cost ceilings plus an emergency stop

Worked example

Stopping when regional evidence is insufficient

Evidence status: Named synthetic scenario

LumenFleet asks the agent for a Singapore campaign brief. Search returns global reports and two local regulatory pages, but no approved local customer evidence. The prototype can still write a convincing brief by blending global themes.

The mandate requires two independent regional sources for pivotal market claims and forbids converting global evidence into local prevalence. The agent retrieves, records the gap, and returns an approval packet recommending targeted interviews rather than a launch brief. It stops after twelve tool calls or when the evidence rule fails. The reviewer can accept the research action, request a narrower brief, or reject the run.

The successful output is a justified refusal to overstate local knowledge. The trace shows queries, retrieved evidence IDs, the unmet rule, effort used, and proposed next step. No campaign or market result is claimed.

Limits

The scenario and evaluation are synthetic. A tool-call cap does not by itself control risk, retrieved sources can still be wrong, and human approval can become superficial under high volume. Anthropic's article offers engineering patterns, not a certification of this agent design.

Method

Build it, with checkpoints

Autonomy ladder from recommendation to reversible execution, with an approval packet displaying proposed changes, evidence, uncertainty, checks, cost, and downstream effect

Field situation

Specify whether the agent may recommend a brief for human approval and what evaluation evidence would be required before adding any reversible action.

  1. 01Prove that planning is needed
  2. 02Write the mandate and stop conditions
  3. 03Design narrow tools and trace fields

Acceptance checks

A recommendation agent whose useful behavior includes stopping, refusing, and escalating. Its tools and traces make the proposal reviewable, and any future autonomy increase is tied to task-specific evidence and a narrow action class.

Why this visualAn autonomy ladder paired with an approval-packet anatomy makes permission boundaries concrete. It should show recommendation, reversible action, and consequential action as separate releases.
  1. 01

    Prove that planning is needed

    Compare the task with the fixed workflow from the automation module. Identify where the next source or step depends on evidence found during the run. Keep stable validation and routing in code.

    CHECKPOINT · The charter identifies at least one genuine variable decision and explains why a fixed sequence is insufficient.

  2. 02

    Write the mandate and stop conditions

    Define trigger, goal, allowed data, tools, evidence rules, prohibited actions, tool-call limit, time, cost, completion, refusal, escalation, and emergency stop. Make ambiguous success unacceptable.

    CHECKPOINT · A reviewer can classify any proposed action as allowed, approval-required, or prohibited from the charter alone.

  3. 03

    Design narrow tools and trace fields

    Give each tool a clear purpose, typed parameters, authorization scope, result schema, side-effect label, and error behavior. Record model and policy versions, tool arguments, results, evidence IDs, cost, and state transitions.

    CHECKPOINT · The agent cannot reach a public, customer, CRM, or spend side effect through any tool or shared credential.

  4. 04

    Build a decision-ready approval packet

    Show changed fields, supporting and conflicting evidence, uncertainty, policy checks, run cost, and what happens after approval. Let reviewers narrow scope or request a specific revision alongside approval and rejection.

    CHECKPOINT · A reviewer can narrow scope or request a specific revision without opening raw execution logs.

  5. 05

    Evaluate and stage autonomy

    Run common, edge, adversarial, missing-evidence, tool-failure, conflicting-source, over-budget, and stop cases. Review traces and define the exact threshold for any future reversible action.

    CHECKPOINT · All prohibited-action tests remain blocked, stop cases terminate within bounds, and the release decision stays recommendation-only unless every written threshold passes.

Hands-on lab

Specify and test a campaign research agent

Use LumenFleet with synthetic records. The agent should produce an approval packet or a controlled refusal and have no credential for public or budget actions.

Prepare

  • • Complete the same task manually and keep one accepted and one rejected brief
  • • Create read-only test tools with approved sample data
  • • Name the person authorized to stop the system and resolve incidents

Deliverable

A versioned mandate, three tool contracts, approval packet schema, twelve-case evaluation set, trace review, and staged-autonomy decision.

Starter kit: Agent mandate and approval contract

Copyable YAML
agent: campaign-research-recommender
goal: propose a source-linked campaign brief for one market and audience
trigger: approved research request
allowed_tools: [approved_web_search, read_evidence_store, save_draft_packet]
prohibited_actions: [publish, send_message, create_ad, change_spend, write_crm]
evidence_rules:
  pivotal_claims: two_independent_sources
  regional_claims: regional_source_required
limits: {tool_calls: 12, elapsed_minutes: 8, run_cost_usd: 2}
stop_conditions: [evidence_gap, policy_conflict, tool_error_limit, budget_limit]
required_output: [proposal_diff, source_ids, uncertainty, policy_checks, cost, next_action]
approval_choices: [approve_recommendation, request_revision, narrow_scope, reject]
emergency_owner: ""
manual_fallback: ""
trace_retention: ""

Expected result

A recommendation agent whose useful behavior includes stopping, refusing, and escalating. Its tools and traces make the proposal reviewable, and any future autonomy increase is tied to task-specific evidence and a narrow action class.

Carry forward

Use accepted and rejected runs, tool traces, reviewer reasons, and stop cases as the core evaluation set in the next module. Preserve recommendation mode as a supported deployment stage.

Acceptance checks

  1. 01The charter states why variable planning is needed and which steps remain deterministic
  2. 02Allowed tools, prohibited actions, evidence requirements, and hard limits are machine-testable
  3. 03The agent has no credential path to publication, outreach, CRM mutation, or spend
  4. 04Approval packets show proposed differences, evidence, uncertainty, checks, cost, and downstream effect
  5. 05At least twelve evaluation cases cover normal, edge, adversarial, and failure behavior
  6. 06A kill switch, incident owner, manual fallback, and trace review procedure are rehearsed

What breaks

Failure clinic

F1The agent is slower and more expensive than a fixed route while taking the same steps every time.
Inspect
Compare tool sequences and decisions across representative traces.
Likely cause
A deterministic workflow was wrapped in agent language without a variable planning need.
Repair
Replace the loop with explicit automation and keep the model only for bounded interpretation.
Prevent next time
Require evidence of path variability before approving an agent architecture.
F2Reviewers approve quickly but later cannot explain the evidence or changed fields.
Inspect
Review the packet, reviewer time, trace access, decision reason, and downstream effect shown.
Likely cause
Approval was designed as a button instead of an informed decision.
Repair
Pause execution authority and redesign packets around differences, evidence, uncertainty, and consequences.
Prevent next time
Test approval comprehension and record structured decision reasons.
F3The agent continues searching after evidence becomes inadequate or the run becomes uneconomic.
Inspect
Check tool-call count, elapsed time, cost, repeated queries, and stop-condition evaluation.
Likely cause
The goal rewards completion but gives no hard effort or evidence boundary.
Repair
Terminate the run, return the evidence gap, and enforce limits outside the model.
Prevent next time
Implement hard budgets and explicit refusal outputs in the orchestration layer.
F4A prompt injection in retrieved material changes the agent's plan or requests a forbidden tool.
Inspect
Review retrieved content, instruction hierarchy, tool request, policy decision, and untrusted-data treatment.
Likely cause
External content was allowed to act as instruction rather than evidence.
Repair
Block the action, quarantine the item, narrow tool access, and add the trace to adversarial tests.
Prevent next time
Separate instructions from untrusted content and authorize every tool call against policy.

Beyond the demo

Production boundary

  1. 01The task requires variable planning and has a bounded completion condition
  2. 02Mandate, evidence rules, prohibited actions, limits, and stop behavior are versioned
  3. 03Tools use least privilege, typed contracts, side-effect labels, and authorization checks
  4. 04Untrusted retrieved content cannot override system or workflow policy
  5. 05Approval packets expose differences, evidence, uncertainty, cost, checks, and consequences
  6. 06Recommendation mode, dry run, reversible action, and higher-risk action have separate release gates
  7. 07Evaluation covers typical, edge, adversarial, tool-failure, refusal, and budget cases
  8. 08Traces support replay and incident review while honoring data minimization and retention
  9. 09Kill switch, manual fallback, rollback, on-call owner, and escalation are rehearsed

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Building effective agents

    Anthropic · Published research · 2026-08-20

    workflow versus agent choice · tool design · human review
  2. [2]
    How we built our multi-agent research system

    Anthropic · Published research · 2026-08-20

    research decomposition · citation quality · effort controls
  3. [3]
    Guardrails and human review

    OpenAI · Official documentation · 2026-08-20

    approval boundaries · guardrails · resumable review
  4. [4]
    Safety in building agents

    OpenAI · Official documentation · 2026-08-20

    prompt injection · data leakage · tool approval
  5. [5]
    Integrate AI into n8n workflows

    n8n · Official documentation · 2026-08-20

    workflow structure · execution handling · testing
  6. [6]risk ownership · measurement · human oversight
  7. [7]
    AI Agent Building learning track

    Tenten AI · Tenten field method · 2026-08-20

    agent architecture · tool use · evaluation progression

Related Tenten resources

Apply the track

Start with one constrained workflow.

Tenten can work with your marketing, data, and technical owners to validate the workflow boundary, build the production controls, operate the first release, and transfer ownership against visible evidence.