Skip to main content

AI marketing evaluation and attribution framework

Lab

Measurement, Evals, and Attribution

Measure output quality, workflow performance, full operating cost, and business movement without pretending one model or touchpoint caused every outcome.

DIFFICULTY
Advanced
ESTIMATED TIME
105 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Separate task quality, workflow health, and business indicators
  • Build a representative evaluation set with written grading rules
  • Calculate operating cost beyond model tokens
  • Use attribution and experiment evidence with explicit uncertainty

Before you start

  • • A defined workflow with versioned run and review records
  • • Access to accepted, rejected, edge, and incident cases
  • • A named owner for the business decision the scorecard will support

Working definition

Measurement and evals

AI marketing measurement combines task evaluations, end-to-end operating metrics, and downstream business evidence. Each layer answers a different question: whether an output meets its rubric, whether the workflow is reliable and economical, and whether the business signal moved. Attribution is a model for decisions under uncertainty, not a ledger of certain causality.

A model can produce high-scoring text while the workflow creates long queues, excessive review, poor adoption, or costly incident recovery. Output evaluation alone misses the operating system.

Evaluation cases drawn only from successful demos hide the work that matters in production: ambiguity, missing fields, conflicting evidence, language variation, edge policy, and dependency failure.

Token charges can be a small part of total cost. Integration, review, retries, monitoring, rework, vendor fees, maintenance, and incident response belong in the same decision record.

Business movement has many causes. Controlled experiments, matched comparisons, and incrementality methods can strengthen inference, but the report must retain assumptions, interference, and uncertainty.

Field situation

Keelhouse Analytics weekly brief scorecard

Named synthetic scenario. Keelhouse Analytics is fictional and all measurements are illustrative field definitions rather than observed results.

Owner
You are the growth operations analyst deciding whether a research-to-brief workflow should remain in recommendation mode, be revised, or expand to one reversible action.
Decision
Create a weekly scorecard and release rule that can support continue, revise, pause, or narrowly expand without inflating attribution.
Starting state
The pilot logs prompt cost and counts generated briefs. Reviewers use free-text comments, rejected work is under-sampled, and the dashboard labels any opportunity touched by a brief as AI-influenced revenue.
Expected outcome
A versioned evaluation set, grading guide, scorecard, full-cost model, and staged release decision.

Constraints

  • • The evaluation set must include hard, rejected, and policy-sensitive cases
  • • Reviewer identities may be pseudonymous but decisions need stable IDs
  • • Business reporting cannot assign deterministic revenue credit to the workflow
  • • Any autonomy expansion requires written quality, cost, reliability, and risk thresholds

Worked example

Rejecting an expansion despite an attractive draft-quality average

Evidence status: Named synthetic scenario

Keelhouse's synthetic weekly dashboard shows a strong average editorial score, but the sample excludes rejected briefs. Trace review reveals that regional claims fail more often, reviewers spend substantial time fixing citations, and two runs created duplicate tasks. The revenue field counts opportunities that would have received marketing support anyway.

The analyst stratifies the evaluation set by common, regional, conflicting-source, missing-evidence, and policy cases. It uses claim accuracy, evidence completeness, decision usefulness, tone, and policy compliance as separate dimensions. The operating layer adds cycle time, active review, acceptance without material correction, duplicate side effects, error recovery, adoption, and full cost. The business layer reports influenced opportunities only as a descriptive proxy and proposes a controlled rollout comparison.

The synthetic release decision is revise: keep recommendation mode, correct source validation and idempotency, expand the holdout set, and do not add an execution tool. The scorecard shows why a high average did not satisfy the written risk and reliability gates.

Limits

No metric is a real result. Rubric scores require calibration, reviewers can disagree, and a holdout may not remove all selection effects. Business proxies should not be relabeled causal. NIST and OpenAI documentation support evaluation practice but do not prescribe the business thresholds used here.

Method

Build it, with checkpoints

Three-layer measurement model for task quality, workflow operation, and business evidence leading to continue, revise, pause, or narrow expansion

Field situation

Create a weekly scorecard and release rule that can support continue, revise, pause, or narrowly expand without inflating attribution.

  1. 01Define the decision and measurement layers
  2. 02Build a representative evaluation set
  3. 03Write and calibrate graders

Acceptance checks

A scorecard that reveals tradeoffs instead of compressing them into one vanity number. The team can find which cases fail, where time and money go, whether people use the output, and what evidence controls the next release.

Why this visualA three-layer scorecard diagram prevents task accuracy, workflow efficiency, and business impact from collapsing into one number. A small release-gate flow shows how thresholds control autonomy.
  1. 01

    Define the decision and measurement layers

    Write what the scorecard can change. Assign task quality, workflow operation, and business evidence to separate sections. Give every metric an owner, formula, source, cadence, and limitation.

    CHECKPOINT · No output-quality measure is presented as revenue evidence, and every field supports a named operating or release decision.

  2. 02

    Build a representative evaluation set

    Stratify historical or approved synthetic cases by frequency, difficulty, segment, language, missing data, conflict, policy risk, and failure cost. Hold back a protected portion from tuning.

    CHECKPOINT · The set contains accepted, rejected, edge, adversarial, tool-failure, and stop cases with documented selection logic.

  3. 03

    Write and calibrate graders

    Define scoring anchors for claims, evidence, usefulness, tone, and policy. Mark critical failures separately from averages. Have two reviewers grade an overlap and discuss disagreements before freezing the rubric version.

    CHECKPOINT · Reviewers can explain each score using visible evidence, and unresolved disagreement is reported rather than averaged away.

  4. 04

    Calculate workflow health and full cost

    Measure latency distribution, active labor, acceptance without material correction, retries, duplicate actions, incidents, recovery, adoption, and cost across model, tools, review, maintenance, and failure.

    CHECKPOINT · The scorecard reconciles to run and labor records and does not treat unknown cost as zero.

  5. 05

    Assess business evidence and make the release call

    Describe the proxy, attribution window, comparison, selection, interference, and concurrent changes. Apply the prewritten gate and record continue, revise, pause, or expansion of one action.

    CHECKPOINT · The release decision follows the threshold even when a favorable average or anecdote points elsewhere, and limitations remain beside the business measure.

Hands-on lab

Build a three-layer workflow scorecard

Use Keelhouse or a real workflow with versioned records. If business data is unavailable, label the field unavailable and complete task and operating layers honestly.

Prepare

  • • Export a sample of accepted, revised, rejected, edge, and incident runs
  • • Name the decision and action that each scorecard threshold controls
  • • Separate any protected holdout before changing prompts or policies

Deliverable

A stratified evaluation set, grader instructions, inter-reviewer calibration note, weekly scorecard, cost worksheet, and signed release decision.

Starter kit: Evaluation and weekly scorecard schema

Copyable CSV headers
# eval_cases.csv
case_id,stratum,input_pointer,expected_behavior,claim_accuracy_rule,evidence_rule,usefulness_rule,policy_rule,critical_failure,holdout,reviewer_id,review_status

# weekly_scorecard.csv
week,workflow_version,eval_pass_rate,critical_failures,claim_accuracy,evidence_completeness,accept_without_material_change,cycle_time_p50,cycle_time_p90,active_review_minutes,duplicate_side_effects,recovery_minutes,adoption_rate,model_cost,tool_cost,review_cost,maintenance_cost,incident_cost,business_proxy,proxy_definition,confidence,decision,decision_reason

# release_rule.md
Current stage:
Required quality thresholds:
Required reliability thresholds:
Maximum full cost:
Zero-tolerance failures:
Minimum adoption evidence:
Business evidence and limitations:
Decision: continue / revise / pause / expand one action

Expected result

A scorecard that reveals tradeoffs instead of compressing them into one vanity number. The team can find which cases fail, where time and money go, whether people use the output, and what evidence controls the next release.

Carry forward

Use the scorecard as the acceptance gate for the forward-deployed operating system. Preserve baseline, holdout, and incident cases across model or vendor changes.

Acceptance checks

  1. 01Task quality, workflow operation, and business evidence appear as separate labeled layers
  2. 02The evaluation set documents strata and includes protected holdout, rejected, edge, and failure cases
  3. 03Rubrics have observable anchors and critical failures cannot be hidden by averages
  4. 04Full cost includes model, tools, review, maintenance, rework, and incidents, with unknowns visible
  5. 05Attribution language distinguishes descriptive association, experiment, and causal estimate
  6. 06A written threshold produces a signed continue, revise, pause, or narrow expansion decision

What breaks

Failure clinic

F1Evaluation improves every week while reviewer complaints and corrections persist.
Inspect
Compare tuned cases, protected holdout, production strata, grader versions, and rejected runs.
Likely cause
The team optimized against a narrow or contaminated evaluation set.
Repair
Restore the holdout, add missed failure strata, and report scores by segment rather than only an average.
Prevent next time
Version sets and prevent prompt authors from editing protected expected answers.
F2ROI appears positive only when the dashboard ignores human review and maintenance.
Inspect
Reconcile labor logs, vendor bills, incident tickets, retries, and opportunity cost with the model bill.
Likely cause
Token cost was treated as total operating cost.
Repair
Recalculate using loaded labor and all recurring and failure costs; mark unavailable inputs.
Prevent next time
Assign finance and workflow owners to approve the cost definition before launch.
F3The dashboard credits the workflow with every touched opportunity.
Inspect
Review touch definition, counterfactual, selection rule, concurrent activity, and attribution window.
Likely cause
A descriptive association was labeled as causal revenue impact.
Repair
Rename the proxy, remove deterministic credit, and design a credible comparison where feasible.
Prevent next time
Require evidence labels and limitations beside every business measure.
F4A high average pass rate masks rare but serious policy or duplicate-action failures.
Inspect
Break out critical failures, severity, tail behavior, and high-cost strata.
Likely cause
The release gate relies on a mean score with no zero-tolerance conditions.
Repair
Pause expansion, repair the failure class, and add severity-weighted and hard-stop gates.
Prevent next time
Report critical events separately and test them on every release.

Beyond the demo

Production boundary

  1. 01Every metric has a definition, source, owner, cadence, and known limitation
  2. 02Evaluation cases represent frequency, difficulty, segment, language, policy, and failure cost
  3. 03A protected holdout is separated from prompt and policy tuning
  4. 04Human graders are calibrated and disagreement is measured
  5. 05Critical failures and tail latency remain visible outside aggregate scores
  6. 06Full cost includes model, tools, labor, review, maintenance, retries, incidents, and recovery
  7. 07Adoption and acceptance without material correction are measured from actual use
  8. 08Business proxies, controlled comparisons, and causal estimates use distinct labels
  9. 09Release thresholds and the authority to continue, pause, revise, or expand are documented

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Evaluate agent workflows

    OpenAI · Official documentation · 2026-08-20

    trace datasets · graders · continuous evaluation
  2. [2]
    Demystifying evals for AI agents

    Anthropic · Published research · 2026-08-20

    representative task sets · repeated trials · grader calibration
  3. [3]
    AI RMF Core: Measure

    NIST AI Resource Center · Official documentation · 2026-08-20

    measurement design · monitoring · risk metrics
  4. [4]risk ownership · measurement · human oversight
  5. [5]
    Building effective agents

    Anthropic · Published research · 2026-08-20

    workflow versus agent choice · tool design · human review
  6. [6]
    AI ROI Calculator

    Tenten AI · Tenten field method · 2026-08-20

    cost framing · readiness inputs · commercial evaluation

Related Tenten resources

Apply the track

Start with one constrained workflow.

Tenten can work with your marketing, data, and technical owners to validate the workflow boundary, build the production controls, operate the first release, and transfer ownership against visible evidence.