Skip to main content

LLM structured outputs TypeScript validation tutorial

Lab

LLM APIs and structured outputs: build a typed boundary

Implement a support classifier that treats valid JSON, schema adherence, refusal, truncation, timeout, and application validation as separate states.

DIFFICULTY
Beginner
ESTIMATED TIME
105 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Choose structured output for final data and function calling for application actions
  • Enforce a narrow schema at the provider and application boundaries
  • Handle refusal, incomplete output, timeout, rate limit, and semantic rejection explicitly
  • Build a fixture suite that survives model or prompt changes

Before you start

  • • Node.js 20 or later and TypeScript
  • • Familiarity with async functions, JSON Schema, and environment variables

Working definition

APIs and structured outputs

A typed model boundary specifies the request contract, constrains the returned shape, validates it again in application code, records the provider outcome, and routes every non-success state without asking downstream code to guess what happened.

Structured Outputs can enforce a supported response schema. They do not verify whether a field is true, authorized, current, or useful for the business decision.

A production caller needs a complete parse path. Refusal, output limits, transient transport errors, and valid but unsupported classifications lead to different recovery actions.

Field situation

Harbor support intake

Named synthetic scenario. Harbor Cloud, ticket text, labels, counts, and thresholds are instructional fixtures with no connection to a real support operation.

Owner
You are implementing the classification boundary used before tickets enter Harbor's routing queue.
Decision
Design the provider schema and application validator that produce a queue-safe classification record.
Starting state
The queue receives billing, access, incident, and unknown requests. A legacy regex sends ambiguous tickets to billing and cannot explain which phrase supported the decision.
Expected outcome
Ten fixtures produce typed records or named failure states. No ticket reaches the routing queue through an unvalidated parse or invented evidence string.

Constraints

  • • The classifier may label and escalate but cannot reply to users or change account data.
  • • Unknown is a valid result and must not be retried as if it were a transport error.
  • • Ticket text may contain malicious instructions and is always untrusted user data.
  • • The application requires a source quote copied from the ticket and rejects evidence not present verbatim.

Worked example

Valid JSON is still rejected when its evidence is invented

Evidence status: Named synthetic scenario

Ticket HBR-07 reads: I was charged twice after changing plans. The expected category is billing, urgency is normal, and the exact evidence span is charged twice. The provider returns a schema-valid object with category billing and evidence duplicate subscription charge. The phrase is plausible but does not occur in the ticket.

Provider schema validation passes. Harbor's application validator rejects the object with EVIDENCE_NOT_IN_INPUT, records the attempt, and places the original ticket in the review queue. It does not retry because a second model call could produce a different unsupported quote while hiding the semantic defect.

The completed fixture distinguishes six terminal states: accepted, human_review, provider_refusal, incomplete, transient_error, and invalid_output. The lesson's expected result is a local test report, not a claim that any model reaches a particular accuracy.

Limits

OpenAI documents schema adherence and a separate refusal field for Structured Outputs. Application evidence checks, Harbor's categories, and its retry policy are local design choices. Anthropic's API uses different request and response details, so the starter keeps provider transport behind one adapter.

Method

Build it, with checkpoints

State diagram for a typed LLM request showing schema validation, refusal and incomplete branches, application evidence checks, bounded retry, acceptance, and human review.

Field situation

Design the provider schema and application validator that produce a queue-safe classification record.

  1. 01Freeze the contract
  2. 02Write the deterministic validator
  3. 03Model every provider outcome

Acceptance checks

A caller that returns a closed Result union. Downstream routing code can handle each state without parsing prose or assuming schema-valid data is semantically safe.

Why this visualUse a deterministic request-state diagram generated from Mermaid or SVG source. It must show provider completion, refusal, incomplete output, transport error, application rejection, acceptance, and human review; an editorial image would obscure the contract.
  1. 01

    Freeze the contract

    Copy the union types and schema. Add a schema version and request ID to your envelope without widening the classification fields.

    CHECKPOINT · TypeScript rejects an extra category at compile time, and the JSON Schema rejects additional properties at runtime.

  2. 02

    Write the deterministic validator

    Check enum membership, required values, evidence presence in the original input, and business escalation rules after the provider parse.

    CHECKPOINT · HBR-07 returns invalid_output with EVIDENCE_NOT_IN_INPUT even though its object matches the provider schema.

  3. 03

    Model every provider outcome

    Make the adapter map refusal, output limit, timeout, rate limit, malformed payload, and success to the Result union. Preserve the provider request ID separately from ticket content.

    CHECKPOINT · No catch block converts every error into the same retry; each fixture reaches its named state.

  4. 04

    Add bounded retry policy

    Retry only selected transient transport outcomes. Use jitter, honor provider guidance when available, cap attempts, and route exhaustion to review.

    CHECKPOINT · The test runner proves semantic rejection and refusal are never retried, while one simulated rate limit retries at most twice.

  5. 05

    Run the version comparison

    Execute all fixtures against the fake baseline. If you enable a real provider, pin the tested model identifier and save aggregate results without storing unnecessary ticket content.

    CHECKPOINT · The report names schema, prompt, model or fake adapter, test-set version, accepted count, review count, error count, latency, and estimated cost.

Hands-on lab

Implement and test the Harbor classifier

Use the copyable TypeScript starter with a fake provider adapter first. Add a real provider only after all deterministic tests pass.

Prepare

  • • Create a clean directory and install TypeScript plus your preferred test runner.
  • • Keep OPENAI_API_KEY or ANTHROPIC_API_KEY out of source control; the default lab path needs neither.
  • • Save the ten tickets as fixtures so future prompt and model changes run against the same inputs.
  • • Set an explicit test timeout and do not add automatic retries around semantic rejection.

Deliverable

A provider-neutral classifier module, ten fixtures, deterministic unit tests, one adapter contract, and a run report that counts every terminal state.

Starter kit: Classifier contract and validator

TypeScript
type Category = "billing" | "access" | "incident" | "unknown";

type Classification = {
  category: Category;
  urgency: "normal" | "urgent";
  evidence: string;
  escalationReason: string | null;
};

type Result =
  | { status: "accepted"; value: Classification }
  | { status: "human_review"; code: string }
  | { status: "provider_refusal"; message: string }
  | { status: "incomplete"; reason: string }
  | { status: "transient_error"; retryAfterMs: number }
  | { status: "invalid_output"; code: string };

export const schema = {
  type: "object",
  additionalProperties: false,
  required: ["category", "urgency", "evidence", "escalationReason"],
  properties: {
    category: { type: "string", enum: ["billing", "access", "incident", "unknown"] },
    urgency: { type: "string", enum: ["normal", "urgent"] },
    evidence: { type: "string", minLength: 1 },
    escalationReason: { type: ["string", "null"] }
  }
} as const;

export function validateEvidence(ticket: string, value: Classification): Result {
  if (!ticket.includes(value.evidence)) {
    return { status: "invalid_output", code: "EVIDENCE_NOT_IN_INPUT" };
  }
  if (value.category === "unknown" || value.escalationReason) {
    return { status: "human_review", code: "CLASSIFICATION_UNCERTAIN" };
  }
  return { status: "accepted", value };
}

Expected result

A caller that returns a closed Result union. Downstream routing code can handle each state without parsing prose or assuming schema-valid data is semantically safe.

Carry forward

Keep the Result union, adapter boundary, schema version, and fixture runner. Module 02 will use the same pattern for tool requests and tool results.

Acceptance checks

  1. 01All ten fixtures terminate and no test uses an unbounded retry loop.
  2. 02The valid-but-invented evidence fixture fails application validation.
  3. 03Refusal, incomplete output, transient failure, unknown classification, and accepted output remain distinct states.
  4. 04Logs contain request and version metadata but exclude full ticket text by default.

What breaks

Failure clinic

F1The parser succeeds, yet downstream code crashes on a missing or unexpected field.
Inspect
Compare the provider schema, application type, raw output status, and runtime validator; check for permissive additional properties.
Likely cause
The application trusted JSON syntax or a TypeScript assertion instead of validating the actual runtime value.
Repair
Use a strict supported schema and validate the parsed object again before routing.
Prevent next time
Run invalid, extra-field, null, refusal, and truncated fixtures in CI for every schema change.
F2A harmful or out-of-scope ticket causes a parse error rather than a visible refusal state.
Inspect
Check the response envelope before attempting to parse the structured content.
Likely cause
The caller assumes every provider response must conform to the business schema.
Repair
Branch on the provider's documented refusal or incomplete surface, then parse only completed content.
Prevent next time
Keep refusal and output-limit fixtures beside every structured-output test set.
F3Rate limiting creates a retry storm and duplicate queue records.
Inspect
Review attempt count, jitter, idempotency key, queue insert trace, and concurrent worker behavior.
Likely cause
Retries are unbounded or repeat the whole workflow, including a downstream write.
Repair
Retry only the provider call with a cap, then perform one idempotent queue transition.
Prevent next time
Test exhaustion, concurrent retries, and duplicate delivery before adding live traffic.
F4Schema-valid evidence does not appear in the source ticket.
Inspect
Perform an exact or normalized span check against the untrusted input and record the mismatch code.
Likely cause
Structured output constrained form, while the model still generated unsupported content.
Repair
Reject the record or require source offsets that the application can verify.
Prevent next time
Include semantic validators for identifiers, quoted evidence, permissions, dates, and business invariants.

Beyond the demo

Production boundary

  1. 01Pin and log schema, prompt, provider adapter, and tested model versions.
  2. 02Keep untrusted user text in user input, outside higher-priority instructions.
  3. 03Use strict provider schemas where supported and retain application validation.
  4. 04Represent unknown, refusal, incomplete output, timeout, rate limit, and invalid content separately.
  5. 05Cap retries, add jitter, and make downstream transitions idempotent.
  6. 06Redact sensitive inputs while retaining request IDs and enough metadata for authorized diagnosis.
  7. 07Re-run a fixed regression set before changing model, prompt, schema, or validation rules.
  8. 08Route unresolved cases to a staffed queue with an owned response-time target.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Structured model outputs

    OpenAI · Official documentation · 2026-08-20

    schema adherence · refusal handling · incomplete output handling
  2. [2]
    Function calling

    OpenAI · Official documentation · 2026-08-20

    strict function schemas · parallel calls · tool-call state machine
  3. [3]
    Tool use with Claude

    Anthropic · Official documentation · 2026-08-20

    tool definitions · tool result loop · client-side execution responsibility
  4. [4]
    Evaluate agent workflows

    OpenAI · Official documentation · 2026-08-20

    trace grading · datasets · repeatable eval runs

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.