Skip to main content

AI agent context memory state architecture tutorial

Lab

Context, memory, and state: make the job resumable

Separate the model's current working context from authoritative job state and selectively retained memory, then test pause, resume, expiry, deletion, and contamination.

DIFFICULTY
Intermediate
ESTIMATED TIME
125 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Define separate contracts for current context, authoritative state, and persistent memory
  • Build context from minimum high-signal records with provenance
  • Resume a paused job without replaying completed side effects
  • Test stale summaries, poisoned notes, cross-tenant retrieval, expiry, and deletion

Before you start

  • • Module 02 trusted Session and audit pattern
  • • Basic relational database concepts and JSON serialization

Working definition

Context, memory, and state

Context is the bounded information sent to the model for the current decision. State is the authoritative application record of goal, progress, approvals, and effects. Memory is a selected record retained for a later run under an owner, provenance, retention, and deletion policy.

A conversation transcript is useful evidence, yet it is a poor database. It can omit external effects, contain untrusted tool text, exceed the context window, and resist deterministic migration or repair.

Persistent memory increases the chance that stale, private, or malicious information influences a future task. Write policy and deletion tests belong in the first design, not in a cleanup phase.

Field situation

Atlas campaign reconciliation

Named synthetic scenario. Atlas Media, campaign records, approvals, memory entries, and timestamps are lab fixtures rather than a production history.

Owner
You maintain a job that compares approved campaign metadata with a synthetic analytics export and pauses when a discrepancy needs operator review.
Decision
Choose which fields live in the database, which enter the next model context, and which observations may persist as memory.
Starting state
A job can run longer than one model context. The prototype stores progress in chat messages, so a resumed run repeats a completed export read and treats an old model summary as the current approval status.
Expected outcome
The job pauses at review, resumes from database state exactly once, excludes expired or foreign-tenant memory, and proves deletion without trusting a generated summary.

Constraints

  • • The database is authoritative for job status, completed effects, and approvals.
  • • Context may include only the current goal, relevant records, latest approved notes, and provenance pointers.
  • • Memory entries require tenant, owner, source, created time, expiry, confidence, and deletion state.
  • • A resume token identifies serialized application state; it cannot authorize a new user or tenant.

Worked example

A polished summary loses to the approval table

Evidence status: Named synthetic scenario

Job job_atlas_17 has already fetched export version 44 and created discrepancy d_08. An earlier context summary says approval pending. The operator later rejects d_08 in the application. A resumed model context still contains the stale summary because the prototype replays chat history without rebuilding state.

The rebuilt version queries the job, effect ledger, discrepancy, and approval tables first. The context builder emits approval_status rejected with record version 3 and omits the stale summary. The loop moves to terminal status rejected and does not fetch export 44 again because the effect ledger contains the idempotency key.

Resume tests assert one export read, one discrepancy, current rejection status, and zero expired-memory records. These are deterministic fixture outcomes. They do not establish that provider-managed conversation features behave identically across APIs.

Limits

The starter models relational state in plain SQL and uses a simple context serializer. Real deployments need encryption, retention review, concurrency control, backups, data residency decisions, and provider-specific conversation semantics. Compaction remains lossy and must be evaluated on representative traces.

Method

Build it, with checkpoints

Architecture diagram separating current model context, authoritative job and approval state, persistent memory with expiry, and the authenticated pause and resume path.

Field situation

Choose which fields live in the database, which enter the next model context, and which observations may persist as memory.

  1. 01Place every fact in one category
  2. 02Seed the paused job
  3. 03Build minimum context

Acceptance checks

A job can survive context reset because progress and effects live outside the transcript. The model sees a compact current view built from authorized records rather than accumulated conversation text.

Why this visualUse a deterministic architecture and lifetime diagram showing context, database state, memory store, provenance, expiry, and resume flow. Generated art cannot communicate source-of-truth or retention boundaries.
  1. 01

    Place every fact in one category

    Map job goal, progress, records, approvals, tool observations, preferences, and summaries to context, state, memory, or discard. State may be rendered into context without ceasing to be authoritative.

    CHECKPOINT · Approval and effect status exist only as database facts; generated summaries cannot overwrite them.

  2. 02

    Seed the paused job

    Insert job_atlas_17, export effect version 44, discrepancy d_08, pending approval, one valid memory, one expired memory, and one foreign-tenant memory.

    CHECKPOINT · The raw fixture contains all three memory cases and one completed idempotency key before resume begins.

  3. 03

    Build minimum context

    Query by authenticated tenant and job, include current record versions and provenance refs, filter memory by owner, expiry, and deletion, then serialize to a bounded object.

    CHECKPOINT · The context contains one valid memory and the current approval; expired and foreign records are absent.

  4. 04

    Reject and resume

    Update d_08 approval to rejected with an optimistic version check. Resume from serialized job state after re-authenticating the caller and rebuilding context.

    CHECKPOINT · The job terminates rejected, export version 44 is not fetched again, and no new discrepancy is created.

  5. 05

    Delete and contaminate

    Soft-delete the valid memory and verify it leaves future context. Add a poisoned memory that claims policy changed, then prove that policy still comes from versioned application configuration.

    CHECKPOINT · Deleted memory is absent, poisoned text cannot alter policy, and the audit log identifies every memory record considered and filtered.

Hands-on lab

Build a resumable Atlas job record

Create the schema below, populate job_atlas_17, pause it for review, change approval state, and resume through a context builder that reads current records.

Prepare

  • • Use SQLite or translate the starter to your local relational database.
  • • Keep the resume token separate from authentication and authorization.
  • • Set fixture time to 2026-08-20T09:00:00Z so expiry tests are repeatable.
  • • Do not persist raw provider reasoning or secrets as memory.

Deliverable

A migration, fixture seed, context-builder output, pause/resume test, effect ledger test, contamination tests, and memory deletion verification.

Starter kit: Authoritative state and memory schema

SQL
CREATE TABLE jobs (
  id TEXT PRIMARY KEY,
  tenant_id TEXT NOT NULL,
  status TEXT NOT NULL CHECK (status IN ('ready','running','awaiting_review','completed','rejected','failed')),
  goal TEXT NOT NULL,
  version INTEGER NOT NULL,
  updated_at TEXT NOT NULL
);

CREATE TABLE effects (
  job_id TEXT NOT NULL,
  idempotency_key TEXT NOT NULL,
  kind TEXT NOT NULL,
  result_ref TEXT NOT NULL,
  completed_at TEXT NOT NULL,
  PRIMARY KEY (job_id, idempotency_key)
);

CREATE TABLE approvals (
  job_id TEXT NOT NULL,
  subject_id TEXT NOT NULL,
  status TEXT NOT NULL CHECK (status IN ('pending','approved','rejected')),
  reviewer_id TEXT,
  version INTEGER NOT NULL,
  decided_at TEXT,
  PRIMARY KEY (job_id, subject_id)
);

CREATE TABLE memories (
  id TEXT PRIMARY KEY,
  tenant_id TEXT NOT NULL,
  owner_id TEXT NOT NULL,
  source_ref TEXT NOT NULL,
  content TEXT NOT NULL,
  confidence REAL NOT NULL,
  created_at TEXT NOT NULL,
  expires_at TEXT NOT NULL,
  deleted_at TEXT
);

Expected result

A job can survive context reset because progress and effects live outside the transcript. The model sees a compact current view built from authorized records rather than accumulated conversation text.

Carry forward

Keep the jobs, effects, approvals, memories, and context-builder contracts. Retrieval metadata in Module 04 and loop checkpoints in Module 06 will attach to these records.

Acceptance checks

  1. 01Resume re-authenticates the caller and never treats a resume token as tenant authority.
  2. 02Completed effects do not repeat, and concurrent approval updates use a version check.
  3. 03Expired, deleted, and foreign-tenant memories never enter model context.
  4. 04A generated or retrieved memory cannot replace configured policy or current approval state.

What breaks

Failure clinic

F1A resumed job repeats a completed read or write.
Inspect
Compare the effect ledger, idempotency key, checkpoint commit, and tool trace around the interruption.
Likely cause
Progress was inferred from conversation history or persisted after the external effect instead of atomically around it.
Repair
Record planned and completed effects with stable idempotency keys, then make resume consult the ledger.
Prevent next time
Test interruption before dispatch, after dispatch, and before result persistence.
F2The agent follows an old approval or preference after an operator changed it.
Inspect
Compare context record versions with the authoritative table and locate cached summaries or memories.
Likely cause
A stale summary was treated as state, or cache invalidation ignored the updated record version.
Repair
Rebuild consequential fields from current state and include versions plus provenance in context.
Prevent next time
Set freshness rules per field and reject context assembled from mismatched versions.
F3A note from another tenant appears in the prompt.
Inspect
Trace retrieval filters, cache keys, memory ownership, session binding, and context-builder logs.
Likely cause
Memory lookup used semantic similarity without applying tenant and owner filters first.
Repair
Scope the query and caches by authenticated tenant before ranking or serialization.
Prevent next time
Keep cross-tenant canary records and assert zero unauthorized context hits.
F4Compaction saves tokens but drops an unresolved constraint needed later.
Inspect
Compare compacted context with a golden retention list across long representative traces.
Likely cause
The summarizer optimized brevity without a recall-first contract for decisions, blockers, and provenance.
Repair
Preserve structured state separately and tune compaction against required retention checks.
Prevent next time
Measure recall before token savings and keep recent critical records outside free-form summaries.

Beyond the demo

Production boundary

  1. 01Keep permissions, effects, approvals, and task status in authoritative application state.
  2. 02Re-authenticate and re-authorize every resumed run.
  3. 03Version state transitions and protect concurrent updates with optimistic or transactional control.
  4. 04Give every memory an owner, tenant, source, confidence, creation time, expiry, and deletion path.
  5. 05Encrypt sensitive state and review provider retention for any context sent off-system.
  6. 06Build context with freshness, provenance, size, and access checks before each consequential decision.
  7. 07Test context overflow, compaction loss, poisoned memory, deletion, and cross-tenant contamination.
  8. 08Log memory selection and state versions without copying unnecessary private content into telemetry.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Effective context engineering for AI agents

    Anthropic · Official documentation · 2026-08-20

    context budgets · compaction · structured note-taking
  2. [2]
    Conversation state

    OpenAI · Official documentation · 2026-08-20

    response continuation · conversation objects · application-managed history
  3. [3]
    Compaction

    OpenAI · Official documentation · 2026-08-20

    long-running response context · compacted state · context-window management
  4. [4]
    AI Risk Management Framework

    NIST · Official documentation · 2026-08-20

    risk governance · measurement · operational accountability

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.