Skip to main content

permission aware RAG retrieval evaluation tutorial

Lab

RAG and knowledge retrieval: test the evidence layer first

Build a two-tenant retrieval fixture, apply access controls before ranking, compare retrieval variants, and grade citations separately from answer prose.

DIFFICULTY
Intermediate
ESTIMATED TIME
150 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Preserve ownership, version, freshness, and permission metadata through chunking
  • Apply tenant and document access filters before semantic or lexical ranking
  • Compare a baseline with contextual chunks, lexical search, and reranking
  • Measure retrieval misses and citation fidelity independently from answer quality

Before you start

  • • Module 03 tenant-scoped state and provenance model
  • • Basic search, embeddings, and test-fixture concepts

Working definition

RAG and retrieval

Retrieval-augmented generation selects authorized external evidence at run time and supplies it to a model with provenance. The retrieval system and answer generator have separate failure modes and need separate evaluation sets.

A polished answer cannot repair missing or unauthorized evidence. If the required passage never reaches context, the model either admits the gap or fills it from prior knowledge and inference.

Document permissions must survive ingestion, chunk creation, indexing, filtering, ranking, caching, citation, and deletion. Applying access checks after retrieval already exposes forbidden content to the system path.

Field situation

Cedar policy answer service

Named synthetic scenario. Cedar Systems, employees, policy documents, expected passages, and retrieval scores are fabricated for instruction.

Owner
You are implementing retrieval for an internal policy assistant used by two synthetic regional teams.
Decision
Choose a retrieval configuration that meets Cedar's hit-rate and zero-unauthorized-hit checks on the supplied corpus.
Starting state
Cedar has six public policies, four EU-only addenda, and four US-only addenda. A prototype embeds plain chunks without permission metadata and evaluates only final answer fluency.
Expected outcome
The selected configuration retrieves the expected authorized chunk for the gold questions, records honest empty results, and never returns a canary from the wrong region.

Constraints

  • • Each chunk inherits tenant, audience, document version, effective date, and deletion status.
  • • Authorization filters run before semantic or lexical scoring and before cache lookup returns content.
  • • Answers cite chunk ID, document ID, title, version, and source URL or internal record path.
  • • The gold set includes authorized hits, empty results, stale versions, and forbidden evidence canaries.

Worked example

The correct passage is forbidden for the active user

Evidence status: Named synthetic scenario

Question CDR-09 asks whether unused training budget rolls into the next quarter. The only passage containing rollover belongs to the EU addendum. The active user is in Cedar US, whose current policy says managers must decide case by case and contains no rollover rule. An unfiltered semantic search ranks the EU passage first.

The retrieval layer filters to Cedar US and public documents before scoring. It returns the US passage plus an evidence_gap status. The answer states that the available US policy does not define rollover and routes the question to the policy owner. It does not cite, summarize, or reveal the EU rule.

The acceptance report records an authorized hit for the US passage, zero forbidden canary hits, and one expected evidence gap. The lab's metrics apply only to the seeded 30-question corpus; no claim is made about general RAG accuracy.

Limits

Anthropic's published Contextual Retrieval results were measured on its experimental setup and are not expected gains for Cedar. Hosted file-search products abstract parts of ingestion and ranking, so learners must verify which permission, deletion, metadata, and result-inspection controls their chosen provider exposes.

Method

Build it, with checkpoints

Permission-aware RAG pipeline showing document metadata inheritance, pre-retrieval authorization, hybrid ranking, evidence assembly, citation validation, and deletion propagation.

Field situation

Choose a retrieval configuration that meets Cedar's hit-rate and zero-unauthorized-hit checks on the supplied corpus.

  1. 01Carry metadata through ingestion
  2. 02Filter before scoring
  3. 03Run retrieval variants

Acceptance checks

A retrieval report that can choose a configuration based on authorized evidence recall, empty-result honesty, latency, and cost before anyone debates prose quality.

Why this visualUse a deterministic ingestion and retrieval diagram with permission filters before ranking, plus a small citation UI sourced from real fixture output. Generated images cannot assert which user can see which chunk.
  1. 01

    Carry metadata through ingestion

    Chunk each document while copying tenant, audience, region, version, effective date, source, and deletion fields. Reject orphan chunks with no parent record.

    CHECKPOINT · Every indexed chunk resolves to one current document and has a complete permission envelope.

  2. 02

    Filter before scoring

    Derive filters from the authenticated session, apply them before vector or lexical ranking, and include the same fields in cache keys.

    CHECKPOINT · CDR-09 never places eu_004_02 in the candidate set for a US user, even when its semantic score would be highest.

  3. 03

    Run retrieval variants

    Compare baseline chunks, contextualized chunks, hybrid semantic plus lexical retrieval, and hybrid retrieval with reranking. Keep top-k and corpus fixed.

    CHECKPOINT · The report shows top-5 and top-20 miss rates, empty results, forbidden hits, latency, and cost for each configuration.

  4. 04

    Generate only from passing evidence

    Build answer context from authorized results with stable citation fields. Require evidence_gap when the available passage cannot support the requested conclusion.

    CHECKPOINT · Every answer sentence marked factual maps to an included chunk, and the CDR-09 answer reveals no EU rollover rule.

  5. 05

    Delete and re-run

    Mark one document deleted, propagate removal through chunks, index, cache, and citation lookup, then repeat the associated gold questions.

    CHECKPOINT · No deleted chunk appears in candidates or citations, and the result records an expected gap instead of stale evidence.

Hands-on lab

Build and evaluate Cedar retrieval

Index the synthetic policies with inherited metadata, run four retrieval variants on 30 gold questions, and reject any configuration with an unauthorized hit.

Prepare

  • • Create a local corpus with public, EU, and US documents plus two canary strings.
  • • Use fake embeddings or a local deterministic similarity stub if no approved embedding service is available.
  • • Freeze document versions and effective dates in the fixture manifest.
  • • Keep answer grading disabled until retrieval results pass access and hit checks.

Deliverable

A corpus manifest, permission-aware chunker, four retrieval configurations, 30-question gold set, retrieval report, citation validator, and deletion test.

Starter kit: Permission-bearing chunk fixture

JSONL
{"chunkId":"pub_001_00","documentId":"pub_001","tenant":"cedar","audience":["public"],"region":null,"version":3,"effectiveAt":"2026-07-01","deletedAt":null,"title":"Learning budget policy","text":"Managers approve learning spend against the current quarter budget."}
{"chunkId":"eu_004_02","documentId":"eu_004","tenant":"cedar","audience":["employee"],"region":"EU","version":2,"effectiveAt":"2026-06-15","deletedAt":null,"title":"EU learning addendum","text":"Unused approved learning budget may roll into the next quarter. CANARY_EU_771."}
{"chunkId":"us_003_01","documentId":"us_003","tenant":"cedar","audience":["employee"],"region":"US","version":4,"effectiveAt":"2026-08-01","deletedAt":null,"title":"US learning addendum","text":"Managers decide exceptions case by case. This addendum does not define automatic rollover."}
{"queryId":"CDR-09","user":{"tenant":"cedar","audience":["employee"],"region":"US"},"question":"Does unused training budget roll over?","expectedChunkIds":["us_003_01"],"forbiddenChunkIds":["eu_004_02"],"expectedStatus":"evidence_gap"}

Expected result

A retrieval report that can choose a configuration based on authorized evidence recall, empty-result honesty, latency, and cost before anyone debates prose quality.

Carry forward

Keep the corpus, permission envelope, gold questions, evidence_gap status, and citation validator. The capstone will reuse them as the governed knowledge layer.

Acceptance checks

  1. 01Unauthorized-hit count is zero for every configuration and user fixture.
  2. 02The chosen variant meets the declared authorized top-k target on the 30-question gold set.
  3. 03Every factual answer sentence cites an included authorized chunk with document version.
  4. 04Deletion removes content from index, cache, candidates, and citations within the tested propagation window.

What breaks

Failure clinic

F1A wrong-region passage appears in candidates even though the final answer omits it.
Inspect
Inspect pre-rank candidates, session-derived filters, cache keys, reranker inputs, and trace redaction.
Likely cause
Permissions were applied after retrieval or omitted from an intermediate cache or reranking service.
Repair
Enforce authorization before candidate generation and preserve it at every intermediate store.
Prevent next time
Use forbidden canary chunks and fail the release on any unauthorized candidate hit.
F2Answers are fluent while the gold passage rarely appears in top-k results.
Inspect
Grade retrieval independently and compare misses by query type, chunk boundary, metadata, and ranking stage.
Likely cause
The team evaluated answer style and allowed the model to rely on prior knowledge or inference.
Repair
Repair corpus, chunking, query, hybrid retrieval, or reranking before tuning the generator.
Prevent next time
Make retrieval thresholds a prerequisite for answer-level evaluation.
F3A superseded policy keeps appearing after the source is updated.
Inspect
Trace parent version, chunk version, index document, cache entry, and citation resolver.
Likely cause
Version and deletion events did not propagate to every derived artifact.
Repair
Use stable parent IDs, tombstones, reindex jobs, cache invalidation, and a freshness check at read time.
Prevent next time
Test update and deletion propagation with a maximum allowed stale window.
F4Reranking improves recall but doubles latency beyond the service target.
Inspect
Break down query, embedding, lexical, reranker, cache, and generation timing by percentile.
Likely cause
Every query uses the expensive path regardless of ambiguity or baseline confidence.
Repair
Route only uncertain queries to reranking or reduce candidate count after measuring recall impact.
Prevent next time
Publish quality, latency, and cost together for every retrieval configuration.

Beyond the demo

Production boundary

  1. 01Maintain a document manifest with owner, source, version, effective date, permissions, and deletion status.
  2. 02Propagate permission metadata to chunks, indexes, caches, rerankers, and citation records.
  3. 03Derive access filters from trusted session state and apply them before ranking.
  4. 04Evaluate authorized retrieval recall separately from answer correctness and writing quality.
  5. 05Preserve exact evidence provenance and expose honest empty or insufficient-evidence states.
  6. 06Reconcile updates and deletions across every derived store within a tested service window.
  7. 07Monitor misses, forbidden canary hits, freshness, latency percentiles, tokens, and cost.
  8. 08Red-team retrieved text for prompt injection before allowing it to influence tool calls or privileged context.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Introducing Contextual Retrieval

    Anthropic · Published research · 2026-08-20

    retrieval baselines · contextual chunks · BM25 and reranking
  2. [2]
    File search

    OpenAI · Official documentation · 2026-08-20

    hosted retrieval · vector stores · retrieval result handling
  3. [3]
    Safety in building agents

    OpenAI · Official documentation · 2026-08-20

    prompt injection · structured data boundaries · MCP approvals
  4. [4]
    AI Risk Management Framework

    NIST · Official documentation · 2026-08-20

    risk governance · measurement · operational accountability

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.