Skip to main content

production AI agent security cost checklist

Lab

Production security and cost: rehearse the operating boundary

Build a staged readiness gate around identity, secrets, data handling, budgets, incidents, rollback, and accountable ownership.

DIFFICULTY
Advanced
ESTIMATED TIME
120 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Map identities, data classes, trust boundaries, permissions, and side effects
  • Estimate unit cost from measured fixture usage and scenario volumes
  • Rehearse revocation, outage, injection, duplicate effect, and manual fallback
  • Issue a staged go or no-go record with owners, gates, rollback, and kill authority

Before you start

  • • A bounded agent with persisted state, approvals, and effect records
  • • A passing evaluation baseline and representative traces from module 9
  • • Access to a non-production environment with fake or safely scoped integrations

Working definition

Production operations

Production readiness is evidence that a named service can operate within declared reliability, security, privacy, cost, and support boundaries. It combines technical controls with ownership: scoped identities, secret isolation, data retention, budgets, alerts, runbooks, rollback, change gates, and authority to stop the system. A successful demo is only one input to that decision.

An agent connects probabilistic decisions to data and tools. The model can suggest an action, but the application still owns authentication, authorization, validation, transaction handling, auditability, and recovery.

Average token cost says little about runaway loops, retrieval fan-out, failed retries, multi-agent amplification, reviewer time, or tool fees. Cost control belongs in the same run budget and release gate as quality and latency.

A control that has never been exercised is an assumption. Incident drills expose missing permissions, stale contact lists, ambiguous kill switches, untested restore paths, and dashboards that cannot answer what happened.

Field situation

Helix internal pilot readiness review

Named synthetic scenario. Helix Labs, volumes, costs, incidents, service levels, and drill results are instructional fixtures. No production deployment or customer outcome is claimed.

Owner
You are the technical owner presenting an internal pilot decision to security, finance, marketing operations, and support.
Decision
Approve, conditionally approve, or reject a 20-user read-only pilot after reviewing the control evidence, unit economics, drills, and residual risks.
Starting state
The read-only Helix research agent passes its locked eval suite and can pause for approval. The candidate pilot has 20 named users and an assumed 300 runs per month. A fake provider reports input and output units per trace; placeholder unit prices live in a versioned rate card rather than in code.
Expected outcome
The team approves only a staged read-only pilot when critical controls and drills pass. Any failed revocation, unknown effect, unowned alert, critical eval regression, or missing rollback produces a no-go result.

Constraints

  • • The pilot remains read-only and cannot publish content, send messages, change spend, or modify source systems.
  • • Only synthetic and separately approved internal documents may enter the staging index.
  • • Every integration uses a service identity with a documented scope, owner, rotation path, and revocation test.
  • • The run budget caps turns, tool calls, wall time, retrieved bytes, and calculated spend before execution begins.
  • • Pricing, quotas, retention, and API behavior are configuration inputs that must be reverified against current provider terms before a real launch.

Worked example

Helix stops a pilot after the credential drill

Evidence status: Named synthetic scenario

The fixture readiness report shows the quality gate passing and a modeled median run cost within its placeholder budget. During the credential-compromise drill, the operator revokes the connector secret. New requests fail, but a warm worker continues using a cached token for nine minutes. The dashboard groups these failures as generic tool errors and does not expose which identity was used.

The review board records no-go despite acceptable quality and modeled spend. The owner removes long-lived worker caching, adds identity ID and credential version to redacted tool events, tests revocation across every worker pool, and defines a maximum propagation objective. The whole critical drill set and locked eval suite must pass again before reconsideration. Thresholds stay fixed during remediation.

A second synthetic drill shows new and warm-worker calls failing closed after revocation, with no secret value in logs. The board conditionally approves a five-user read-only canary, not the full proposed pilot. The record sets a two-week review, named on-call owner, per-run and daily circuit breakers, rollback command, and explicit prohibition on writes.

Limits

The drill events, timing, costs, and approval are fabricated to teach a decision process. They do not establish a suitable revocation objective or budget for another system. Provider security guides reduce common risks but do not prove safety, and rate cards change. Replace every placeholder with verified contractual, architectural, and measured values before deployment.

Method

Build it, with checkpoints

Helix trust-boundary map connecting user, agent service, model API, read-only tools, corpus, state, approval, and telemetry, annotated with identities, data classes, budgets, and kill controls.

Field situation

Approve, conditionally approve, or reject a 20-user read-only pilot after reviewing the control evidence, unit economics, drills, and residual risks.

  1. 01Draw the service and trust inventory
  2. 02Enforce budgets and permissions
  3. 03Model unit economics from traces

Acceptance checks

The evidence pack supports a five-user read-only canary or a clear no-go. Costs are reproducible from measured fixtures and a replaceable rate card. All five drills have observable detection and recovery records, and one command or configuration flag stops new runs without a model decision.

Why this visualRender a deterministic trust-boundary diagram from the service inventory and place the readiness matrix beside it. Security claims require exact identities, directions, scopes, data classes, and owners.
  1. 01

    Draw the service and trust inventory

    List user, application, model provider, MCP or tool server, corpus, state store, approval service, and telemetry path. For each boundary, record identity, data classes, encryption responsibility, retention, region or contractual constraint, and permitted operations.

    CHECKPOINT · Every network call and stored field maps to an owner, purpose, data class, credential, permission, retention rule, and deletion or revocation path.

  2. 02

    Enforce budgets and permissions

    Load limits from the release manifest, give the service identity read-only scopes, and fail closed when any tool, class, or budget is missing. Add per-run, per-user, and daily circuit breakers plus a global disable flag independent of the model.

    CHECKPOINT · Seeded over-turn, over-call, over-time, over-byte, over-cost, and prohibited-tool fixtures stop before a disallowed effect and emit distinct reason codes.

  3. 03

    Model unit economics from traces

    Use measured input, output, tool, retry, and reviewer quantities from the eval suite. Apply a dated placeholder rate card, calculate median and p95 run cost, simulate monthly volume and failure spikes, then state which items remain unknown. Keep quality and reviewer burden beside the cost figure.

    CHECKPOINT · Another reviewer can reproduce every formula from trace exports and swap the rate card without editing application logic.

  4. 04

    Run five controlled incidents

    Revoke a credential, simulate provider outage, inject hostile retrieved text, force timeout after a fake effect, and exercise manual read-only fallback. For each, record detection, containment, user communication, recovery, evidence, and follow-up owner.

    CHECKPOINT · The drill pack proves fail-closed behavior, no completed prohibited effect, accountable decisions, and a known final state for every attempted run.

  5. 05

    Hold the readiness review

    Present eval gates, drill results, data inventory, permission tests, cost distribution, residual risks, on-call ownership, canary scope, rollback, and review date. Record dissent and conditions without editing gates to make the candidate pass.

    CHECKPOINT · The signed decision names release, hold, or rollback; exact scope; evidence links; owners; expiry; stop authority; and conditions for expansion.

Hands-on lab

Run a five-incident readiness review

Assemble the Helix service inventory and cost model, run five controlled incidents in staging, and present an evidence-linked decision for a read-only canary.

Prepare

  • • Use fake connectors or dedicated staging tenants with no production write permission.
  • • Name a facilitator, operator, observer, decision owner, and safety stop for each drill.
  • • Version the evaluation report, policy map, rate card, and runbook used by the review.

Deliverable

A versioned service and data-flow inventory, permission matrix, measured cost workbook, five drill reports, monitoring map, runbook, risk register, and signed canary decision.

Starter kit: Readiness manifest and cost equation

YAML
service: helix-research-agent
release: 0.9.0-rc1
scope: internal-read-only-canary
users: 5
prohibited_effects: [publish, send_message, change_budget, write_source_record]
budgets:
  max_turns_per_run: 8
  max_tool_calls_per_run: 12
  max_wall_seconds: 90
  max_retrieved_bytes: 250000
  max_calculated_usd_per_run: 1.00
identity:
  service_principal: helix-research-staging
  scopes: [corpus.read, crm_fixture.read]
  owner: platform-oncall
data:
  allowed_classes: [synthetic, approved-internal]
  retention_days: 30
critical_gates:
  prohibited_effect_count: 0
  critical_eval_failures: 0
  unresolved_unknown_effects: 0
  revocation_drill: pass
  rollback_drill: pass
drills:
  - credential_revocation
  - provider_outage
  - retrieved_prompt_injection
  - timeout_after_fake_effect
  - manual_read_only_fallback

# Use a dated, reviewed rate card. Do not paste current vendor prices into code.
# modeled_run_usd = input_units * input_rate + output_units * output_rate
#                 + tool_fees + retry_allowance + allocated_review_cost

Expected result

The evidence pack supports a five-user read-only canary or a clear no-go. Costs are reproducible from measured fixtures and a replaceable rate card. All five drills have observable detection and recovery records, and one command or configuration flag stops new runs without a model decision.

Carry forward

Use the readiness manifest, cost workbook, incident cases, runbook, and canary conditions as mandatory inputs to the capstone. The capstone cannot claim production readiness without replacing fixture assumptions and rerunning the gates in its target environment.

Acceptance checks

  1. 01The inventory covers every identity, integration, data class, storage location, permission, owner, retention rule, and revocation path.
  2. 02All six run-budget breaches and every prohibited operation fail closed with traceable reason codes.
  3. 03Median, p95, monthly, retry, tool, and human-review costs are reproducible and clearly labeled as fixture-based estimates.
  4. 04Five drills and the signed decision record prove detection, containment, recovery, rollback, stop authority, residual risk, and next review date.

What breaks

Failure clinic

F1Revoked credentials continue working on some workers.
Inspect
Correlate identity ID, credential version, worker pool, cache lifetime, and revocation timestamp.
Likely cause
Long-lived tokens or connection pools outlast the assumed revocation boundary.
Repair
Invalidate caches, shorten credential lifetime, recycle affected workers, and verify propagation before restoring service.
Prevent next time
Automate revocation drills across every execution pool and alert on obsolete credential versions.
F2A timeout leaves the team unsure whether an external write occurred.
Inspect
Check the local effect ledger, provider idempotency key, adapter receipt, and reconciliation endpoint.
Likely cause
The client retries an ambiguous result without reserving or reconciling the effect.
Repair
Quarantine the run, reconcile authoritative state, and resume only from the recorded effect outcome.
Prevent next time
Design idempotency and unknown-outcome recovery before granting write permission.
F3Monthly spend exceeds the model despite stable request count.
Inspect
Slice turns, tokens, tool fees, retries, retrieval bytes, worker count, and reviewer minutes by task class.
Likely cause
The model uses an average call estimate and omits tail behavior or retry amplification.
Repair
Recalculate from traces, cap expensive branches, and route predictable tasks to simpler paths.
Prevent next time
Alert on unit distributions and forecast volume, quality, and review load together.
F4The kill switch disables the UI but background runs continue.
Inspect
Check scheduler, queue consumers, resumptions, worker leases, and tool-server authorization after disablement.
Likely cause
Stop control exists only at request intake and is not enforced at every execution boundary.
Repair
Reject new work, cancel leases, prevent resume, revoke tool access, and reconcile in-flight effects.
Prevent next time
Exercise global stop and restore in scheduled drills with queue and worker assertions.

Beyond the demo

Production boundary

  1. 01Map every service, trust boundary, data class, identity, credential, and permitted operation.
  2. 02Use least-privilege service identities with tested rotation and revocation paths.
  3. 03Keep secrets out of prompts, tool output, traces, screenshots, and approval packets.
  4. 04Enforce per-run, per-user, daily, and global circuit breakers outside the model.
  5. 05Measure unit cost, quality, latency, retries, tool fees, and reviewer burden by task class.
  6. 06Define retention, deletion, access review, incident notification, and vendor-change review.
  7. 07Maintain runbooks for outage, injection, credential compromise, unknown effect, rollback, and manual fallback.
  8. 08Give a named human owner authority to halt runs and require evidence before scope expansion.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Safety in building agents

    OpenAI · Official documentation · 2026-08-20

    prompt injection · structured data boundaries · MCP approvals
  2. [2]
    Guardrails and human review

    OpenAI · Official documentation · 2026-08-20

    approval interruptions · resumable state · tool-level guardrails
  3. [3]
    Effective context engineering for AI agents

    Anthropic · Official documentation · 2026-08-20

    context budgets · compaction · structured note-taking
  4. [4]
    AI Risk Management Framework

    NIST · Official documentation · 2026-08-20

    risk governance · measurement · operational accountability

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.