On this page
Learning objectives
- Separate task quality, workflow health, and business indicators
- Build a representative evaluation set with written grading rules
- Calculate operating cost beyond model tokens
- Use attribution and experiment evidence with explicit uncertainty
Before you start
- • A defined workflow with versioned run and review records
- • Access to accepted, rejected, edge, and incident cases
- • A named owner for the business decision the scorecard will support
Working definition
Measurement and evals
AI marketing measurement combines task evaluations, end-to-end operating metrics, and downstream business evidence. Each layer answers a different question: whether an output meets its rubric, whether the workflow is reliable and economical, and whether the business signal moved. Attribution is a model for decisions under uncertainty, not a ledger of certain causality.
A model can produce high-scoring text while the workflow creates long queues, excessive review, poor adoption, or costly incident recovery. Output evaluation alone misses the operating system.
Evaluation cases drawn only from successful demos hide the work that matters in production: ambiguity, missing fields, conflicting evidence, language variation, edge policy, and dependency failure.
Token charges can be a small part of total cost. Integration, review, retries, monitoring, rework, vendor fees, maintenance, and incident response belong in the same decision record.
Business movement has many causes. Controlled experiments, matched comparisons, and incrementality methods can strengthen inference, but the report must retain assumptions, interference, and uncertainty.
Field situation
Keelhouse Analytics weekly brief scorecard
Named synthetic scenario. Keelhouse Analytics is fictional and all measurements are illustrative field definitions rather than observed results.
- Owner
- You are the growth operations analyst deciding whether a research-to-brief workflow should remain in recommendation mode, be revised, or expand to one reversible action.
- Decision
- Create a weekly scorecard and release rule that can support continue, revise, pause, or narrowly expand without inflating attribution.
- Starting state
- The pilot logs prompt cost and counts generated briefs. Reviewers use free-text comments, rejected work is under-sampled, and the dashboard labels any opportunity touched by a brief as AI-influenced revenue.
- Expected outcome
- A versioned evaluation set, grading guide, scorecard, full-cost model, and staged release decision.
Constraints
- • The evaluation set must include hard, rejected, and policy-sensitive cases
- • Reviewer identities may be pseudonymous but decisions need stable IDs
- • Business reporting cannot assign deterministic revenue credit to the workflow
- • Any autonomy expansion requires written quality, cost, reliability, and risk thresholds
Worked example
Rejecting an expansion despite an attractive draft-quality average
Evidence status: Named synthetic scenarioKeelhouse's synthetic weekly dashboard shows a strong average editorial score, but the sample excludes rejected briefs. Trace review reveals that regional claims fail more often, reviewers spend substantial time fixing citations, and two runs created duplicate tasks. The revenue field counts opportunities that would have received marketing support anyway.
The analyst stratifies the evaluation set by common, regional, conflicting-source, missing-evidence, and policy cases. It uses claim accuracy, evidence completeness, decision usefulness, tone, and policy compliance as separate dimensions. The operating layer adds cycle time, active review, acceptance without material correction, duplicate side effects, error recovery, adoption, and full cost. The business layer reports influenced opportunities only as a descriptive proxy and proposes a controlled rollout comparison.
The synthetic release decision is revise: keep recommendation mode, correct source validation and idempotency, expand the holdout set, and do not add an execution tool. The scorecard shows why a high average did not satisfy the written risk and reliability gates.
Limits
No metric is a real result. Rubric scores require calibration, reviewers can disagree, and a holdout may not remove all selection effects. Business proxies should not be relabeled causal. NIST and OpenAI documentation support evaluation practice but do not prescribe the business thresholds used here.
Method
Build it, with checkpoints
Field situation
Create a weekly scorecard and release rule that can support continue, revise, pause, or narrowly expand without inflating attribution.
- 01Define the decision and measurement layers
- 02Build a representative evaluation set
- 03Write and calibrate graders
Acceptance checks
A scorecard that reveals tradeoffs instead of compressing them into one vanity number. The team can find which cases fail, where time and money go, whether people use the output, and what evidence controls the next release.
- 01
Define the decision and measurement layers
Write what the scorecard can change. Assign task quality, workflow operation, and business evidence to separate sections. Give every metric an owner, formula, source, cadence, and limitation.
CHECKPOINT · No output-quality measure is presented as revenue evidence, and every field supports a named operating or release decision.
- 02
Build a representative evaluation set
Stratify historical or approved synthetic cases by frequency, difficulty, segment, language, missing data, conflict, policy risk, and failure cost. Hold back a protected portion from tuning.
CHECKPOINT · The set contains accepted, rejected, edge, adversarial, tool-failure, and stop cases with documented selection logic.
- 03
Write and calibrate graders
Define scoring anchors for claims, evidence, usefulness, tone, and policy. Mark critical failures separately from averages. Have two reviewers grade an overlap and discuss disagreements before freezing the rubric version.
CHECKPOINT · Reviewers can explain each score using visible evidence, and unresolved disagreement is reported rather than averaged away.
- 04
Calculate workflow health and full cost
Measure latency distribution, active labor, acceptance without material correction, retries, duplicate actions, incidents, recovery, adoption, and cost across model, tools, review, maintenance, and failure.
CHECKPOINT · The scorecard reconciles to run and labor records and does not treat unknown cost as zero.
- 05
Assess business evidence and make the release call
Describe the proxy, attribution window, comparison, selection, interference, and concurrent changes. Apply the prewritten gate and record continue, revise, pause, or expansion of one action.
CHECKPOINT · The release decision follows the threshold even when a favorable average or anecdote points elsewhere, and limitations remain beside the business measure.
Hands-on lab
Build a three-layer workflow scorecard
Use Keelhouse or a real workflow with versioned records. If business data is unavailable, label the field unavailable and complete task and operating layers honestly.
Prepare
- • Export a sample of accepted, revised, rejected, edge, and incident runs
- • Name the decision and action that each scorecard threshold controls
- • Separate any protected holdout before changing prompts or policies
Deliverable
A stratified evaluation set, grader instructions, inter-reviewer calibration note, weekly scorecard, cost worksheet, and signed release decision.
Starter kit: Evaluation and weekly scorecard schema
Copyable CSV headers# eval_cases.csv
case_id,stratum,input_pointer,expected_behavior,claim_accuracy_rule,evidence_rule,usefulness_rule,policy_rule,critical_failure,holdout,reviewer_id,review_status
# weekly_scorecard.csv
week,workflow_version,eval_pass_rate,critical_failures,claim_accuracy,evidence_completeness,accept_without_material_change,cycle_time_p50,cycle_time_p90,active_review_minutes,duplicate_side_effects,recovery_minutes,adoption_rate,model_cost,tool_cost,review_cost,maintenance_cost,incident_cost,business_proxy,proxy_definition,confidence,decision,decision_reason
# release_rule.md
Current stage:
Required quality thresholds:
Required reliability thresholds:
Maximum full cost:
Zero-tolerance failures:
Minimum adoption evidence:
Business evidence and limitations:
Decision: continue / revise / pause / expand one actionExpected result
A scorecard that reveals tradeoffs instead of compressing them into one vanity number. The team can find which cases fail, where time and money go, whether people use the output, and what evidence controls the next release.
Carry forward
Use the scorecard as the acceptance gate for the forward-deployed operating system. Preserve baseline, holdout, and incident cases across model or vendor changes.
Acceptance checks
- 01Task quality, workflow operation, and business evidence appear as separate labeled layers
- 02The evaluation set documents strata and includes protected holdout, rejected, edge, and failure cases
- 03Rubrics have observable anchors and critical failures cannot be hidden by averages
- 04Full cost includes model, tools, review, maintenance, rework, and incidents, with unknowns visible
- 05Attribution language distinguishes descriptive association, experiment, and causal estimate
- 06A written threshold produces a signed continue, revise, pause, or narrow expansion decision
What breaks
Failure clinic
F1Evaluation improves every week while reviewer complaints and corrections persist.
- Inspect
- Compare tuned cases, protected holdout, production strata, grader versions, and rejected runs.
- Likely cause
- The team optimized against a narrow or contaminated evaluation set.
- Repair
- Restore the holdout, add missed failure strata, and report scores by segment rather than only an average.
- Prevent next time
- Version sets and prevent prompt authors from editing protected expected answers.
F2ROI appears positive only when the dashboard ignores human review and maintenance.
- Inspect
- Reconcile labor logs, vendor bills, incident tickets, retries, and opportunity cost with the model bill.
- Likely cause
- Token cost was treated as total operating cost.
- Repair
- Recalculate using loaded labor and all recurring and failure costs; mark unavailable inputs.
- Prevent next time
- Assign finance and workflow owners to approve the cost definition before launch.
F3The dashboard credits the workflow with every touched opportunity.
- Inspect
- Review touch definition, counterfactual, selection rule, concurrent activity, and attribution window.
- Likely cause
- A descriptive association was labeled as causal revenue impact.
- Repair
- Rename the proxy, remove deterministic credit, and design a credible comparison where feasible.
- Prevent next time
- Require evidence labels and limitations beside every business measure.
F4A high average pass rate masks rare but serious policy or duplicate-action failures.
- Inspect
- Break out critical failures, severity, tail behavior, and high-cost strata.
- Likely cause
- The release gate relies on a mean score with no zero-tolerance conditions.
- Repair
- Pause expansion, repair the failure class, and add severity-weighted and hard-stop gates.
- Prevent next time
- Report critical events separately and test them on every release.
Beyond the demo
Production boundary
- 01Every metric has a definition, source, owner, cadence, and known limitation
- 02Evaluation cases represent frequency, difficulty, segment, language, policy, and failure cost
- 03A protected holdout is separated from prompt and policy tuning
- 04Human graders are calibrated and disagreement is measured
- 05Critical failures and tail latency remain visible outside aggregate scores
- 06Full cost includes model, tools, labor, review, maintenance, retries, incidents, and recovery
- 07Adoption and acceptance without material correction are measured from actual use
- 08Business proxies, controlled comparisons, and causal estimates use distinct labels
- 09Release thresholds and the authority to continue, pause, revise, or expand are documented
Evidence status
Sources and claim limits
Sources support the named claims; they do not guarantee the same result in another system.
- [1]Evaluate agent workflowstrace datasets · graders · continuous evaluation
OpenAI · Official documentation · 2026-08-20
- [2]Demystifying evals for AI agentsrepresentative task sets · repeated trials · grader calibration
Anthropic · Published research · 2026-08-20
- [3]AI RMF Core: Measuremeasurement design · monitoring · risk metrics
NIST AI Resource Center · Official documentation · 2026-08-20
- [4]Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profilerisk ownership · measurement · human oversight
NIST · Official documentation · 2026-08-20
- [5]Building effective agentsworkflow versus agent choice · tool design · human review
Anthropic · Published research · 2026-08-20
- [6]AI ROI Calculatorcost framing · readiness inputs · commercial evaluation
Tenten AI · Tenten field method · 2026-08-20
Related Tenten resources