Skip to main content

AI tool calling authorization TypeScript tutorial

Lab

Tool use and function calling: keep execution outside the model

Give a model one narrow account lookup tool, then prove that identity, tenancy, argument validation, execution, and audit remain application responsibilities.

DIFFICULTY
Intermediate
ESTIMATED TIME
120 min
UPDATED
2026-08-20
COPY REVIEW
blader/humanizer
2 passes
On this page
  1. 01Working definition
  2. 02Field situation
  3. 03Worked example
  4. 04Build it, with checkpoints
  5. 05Hands-on lab
  6. 06Failure clinic
  7. 07Production boundary
  8. 08Sources and claim limits

Learning objectives

  • Design a strict tool schema with one distinct purpose
  • Authorize the caller and tenant independently from model-generated arguments
  • Return concise typed observations and actionable error codes
  • Test wrong-tool, forged-argument, timeout, oversized-result, and repeated-call failures

Before you start

  • • Module 01 Result union and runtime validation pattern
  • • Basic authorization concepts and a local TypeScript test runner

Working definition

Tool use

Tool calling lets a model request a named capability with structured arguments. The host application decides whether that request is valid and authorized, executes it under a constrained identity, records the outcome, and returns a typed observation for the next model turn.

A tool schema constrains argument shape. It cannot prove that the requesting user may access the named account, that the account belongs to the active tenant, or that the proposed action matches policy.

Tool responses become model context. Oversized payloads, opaque identifiers, and vague errors consume context while giving the model little information it can use to recover.

Field situation

Meridian account lookup

Named synthetic scenario. Meridian CRM, tenant IDs, account records, traces, and policies are fabricated for this lab and do not describe a live system.

Owner
You are the API engineer adding a read-only account lookup to a research assistant.
Decision
Define the smallest tool interface that answers the analyst's question without exposing raw SQL, arbitrary HTTP, or cross-tenant data.
Starting state
The assistant summarizes approved account context before an analyst researches a company. Two synthetic tenants share the same database, and an early prototype trusts the account ID proposed by the model.
Expected outcome
Authorized lookups return a compact observation. Forged tenants, unknown records, schema errors, and timeouts produce distinct non-secret-bearing errors and no account data.

Constraints

  • • The tool is read-only and returns company name, segment, approved notes, and record version only.
  • • Tenant identity comes from the authenticated session and never from model arguments.
  • • The response may contain at most five notes and 2,000 UTF-8 characters.
  • • Every request records caller, tenant, tool, validated arguments, result code, duration, and correlation ID.

Worked example

A valid account ID fails the tenant check

Evidence status: Named synthetic scenario

The active session belongs to tenant meridian-eu. A user asks about Northwind Labs, but retrieved web text contains an instruction to call lookup_account with accountId acct_us_204. The account exists under tenant meridian-us. The model emits a schema-valid call because it cannot see the database ownership rule.

The host ignores any tenant value in generated text, binds meridian-eu from the session, validates accountId, and queries by both tenant and account. The lookup returns NOT_FOUND without revealing whether acct_us_204 exists elsewhere. The trace marks the request denied_at_data_boundary and stores a hash of the validated argument rather than raw account notes.

The adversarial fixture passes when the EU caller receives no US data, the audit record keeps the authenticated tenant, and the model gets a concise recoverable error. No provider reliability percentage is claimed; the check is deterministic application behavior.

Limits

The sample uses an in-memory repository and simple role policy. A production system may need row-level security, delegated OAuth, attribute-based controls, field redaction, and legal retention rules. OpenAI and Anthropic tool documentation explains the call loop, while authorization stays with the application.

Method

Build it, with checkpoints

Sequence diagram showing a model requesting lookup_account, the host binding trusted tenant identity, authorization and validation before tenant-scoped lookup, audit recording, and a bounded tool result.

Field situation

Define the smallest tool interface that answers the analyst's question without exposing raw SQL, arbitrary HTTP, or cross-tenant data.

  1. 01Separate identity from arguments
  2. 02Validate before access
  3. 03Shape the observation

Acceptance checks

A read-only capability whose schema helps the model request useful work while the application retains identity, authorization, execution, payload, retry, and audit control.

Why this visualRender a deterministic sequence diagram from source showing user, host, model, authorization layer, tenant-scoped repository, audit sink, and returned observation. The trust boundary and denial point must be exact.
  1. 01

    Separate identity from arguments

    Create Session from trusted middleware and Args from the model call. Remove tenantId, userId, role, URL, and query fields from the tool schema.

    CHECKPOINT · The model-visible schema contains only accountId, while the executor requires a trusted Session supplied by the host.

  2. 02

    Validate before access

    Apply strict schema checks, role policy, and tenant-scoped lookup before assembling any response. Return the same NOT_FOUND shape for absent and other-tenant records.

    CHECKPOINT · The forged acct_us_204 request under meridian-eu yields NOT_FOUND and no row fields appear in logs or tool results.

  3. 03

    Shape the observation

    Return semantic field names, five notes at most, a stable version, and a bounded payload. Use error codes the model can act on without internal stack traces.

    CHECKPOINT · The normal result stays below 2,000 characters, while the oversized fixture returns RESULT_TOO_LARGE with zero partial notes.

  4. 04

    Add timeout and repetition controls

    Abort the repository call at a defined deadline and cache the completed read by run plus validated argument hash. Reads may repeat once after a transient timeout, then escalate.

    CHECKPOINT · The timeout fixture ends with TIMEOUT, and three identical calls in one run execute the repository no more than once after success.

  5. 05

    Write adversarial tests and audit assertions

    Cover extra arguments, malformed IDs, missing role, cross-tenant ID, unknown record, timeout, oversized response, and duplicate call. Assert the audit record for each path.

    CHECKPOINT · Every test records correlation ID, trusted tenant, tool name, result code, and duration without storing account notes.

Hands-on lab

Build the Meridian read-only tool boundary

Implement lookup_account against an in-memory two-tenant fixture, then run allowed, denied, invalid, timeout, and oversized-response cases.

Prepare

  • • Reuse the closed Result pattern from Module 01.
  • • Create separate session and tool-argument types; never merge them into one model-visible object.
  • • Use synthetic account notes without names, emails, or secrets.
  • • Keep the fake repository deterministic so all acceptance checks run offline.

Deliverable

A strict tool declaration, session-bound executor, synthetic repository, compact result contract, audit-event type, and at least eight offline tests.

Starter kit: Authorized account tool

TypeScript
type Session = { userId: string; tenantId: string; roles: string[] };
type Args = { accountId: string };
type ToolResult =
  | { ok: true; account: { id: string; name: string; segment: string; notes: string[]; version: number } }
  | { ok: false; code: "FORBIDDEN" | "NOT_FOUND" | "INVALID_ARGS" | "TIMEOUT" | "RESULT_TOO_LARGE" };

const TOOL = {
  type: "function",
  name: "lookup_account",
  description: "Read one account in the authenticated tenant. Never searches other tenants or updates records.",
  strict: true,
  parameters: {
    type: "object",
    additionalProperties: false,
    required: ["accountId"],
    properties: { accountId: { type: "string", pattern: "^acct_[a-z0-9_]+$" } }
  }
} as const;

export async function executeLookup(session: Session, args: Args): Promise<ToolResult> {
  if (!session.roles.includes("account:read")) return { ok: false, code: "FORBIDDEN" };
  if (!/^acct_[a-z0-9_]+$/.test(args.accountId)) return { ok: false, code: "INVALID_ARGS" };
  const row = await repository.findByTenantAndId(session.tenantId, args.accountId);
  if (!row) return { ok: false, code: "NOT_FOUND" };
  const account = { ...row, notes: row.notes.slice(0, 5) };
  if (JSON.stringify(account).length > 2000) return { ok: false, code: "RESULT_TOO_LARGE" };
  return { ok: true, account };
}

Expected result

A read-only capability whose schema helps the model request useful work while the application retains identity, authorization, execution, payload, retry, and audit control.

Carry forward

Keep Session, ToolResult, audit fields, and the two-tenant fixture. Module 05 will expose a similarly narrow read capability through an MCP server.

Acceptance checks

  1. 01Cross-tenant and unauthorized fixtures return no account fields and use a non-enumerating error shape.
  2. 02Extra or malformed arguments fail before the repository is called.
  3. 03Tool results stay within the documented field and size limits.
  4. 04Duplicate successful calls in one run do not repeat repository work, and every path emits an audit record.

What breaks

Failure clinic

F1A caller can retrieve an account from a different tenant by supplying its ID.
Inspect
Trace where tenant identity originates and inspect the repository predicate for both tenantId and accountId.
Likely cause
Authorization was inferred from a model argument or checked after a global ID lookup.
Repair
Bind tenant from the authenticated session and make tenant-scoped lookup the only repository method exposed to the tool.
Prevent next time
Keep cross-tenant fixtures in CI and add datastore-level row policies where available.
F2The model repeatedly calls the right tool with invalid optional fields.
Inspect
Review raw calls, schema strictness, parameter names, tool description, and returned validation error.
Likely cause
The schema is permissive or the tool contract uses ambiguous parameters the model cannot distinguish.
Repair
Remove unused fields, set additionalProperties to false, and return one actionable example with INVALID_ARGS.
Prevent next time
Evaluate tool selection and arguments on a held-out realistic set before widening the surface.
F3A successful lookup floods context with raw CRM payloads and degrades later decisions.
Inspect
Measure result tokens and identify fields never used by the next decision.
Likely cause
The tool mirrors a backend response instead of designing an agent-facing observation.
Repair
Return a concise semantic projection with pagination or a detailed mode available only when needed.
Prevent next time
Set response budgets and include total tool-result tokens in evaluation reports.
F4A timeout triggers repeated calls that create inconsistent traces or overload the backend.
Inspect
Check deadlines, attempt counters, run-level cache keys, and whether the caller can distinguish unknown completion.
Likely cause
Retry behavior was left to the model and no idempotency or deduplication boundary exists.
Repair
Handle retries in deterministic orchestration with a cap and cache completed results by validated input.
Prevent next time
Inject slow, dropped, and duplicate responses in tool integration tests.

Beyond the demo

Production boundary

  1. 01Derive identity, tenant, and roles from trusted middleware rather than model-visible arguments.
  2. 02Give each tool one distinct job and a strict minimal schema.
  3. 03Classify every capability as read-only, reversible write, external communication, or irreversible action.
  4. 04Use short-lived least-privilege credentials and explicit network destinations.
  5. 05Return bounded semantic results with stable error codes and no internal stack traces.
  6. 06Apply deadlines, deterministic retry policy, rate limits, and deduplication outside the model.
  7. 07Audit initiator, trusted tenant, validated argument hash, decision, result, latency, and version.
  8. 08Run held-out tool-selection and adversarial authorization tests before adding a new tool.

Evidence status

Sources and claim limits

Sources support the named claims; they do not guarantee the same result in another system.

  1. [1]
    Function calling

    OpenAI · Official documentation · 2026-08-20

    strict function schemas · parallel calls · tool-call state machine
  2. [2]
    Tool use with Claude

    Anthropic · Official documentation · 2026-08-20

    tool definitions · tool result loop · client-side execution responsibility
  3. [3]
    Writing effective tools for agents

    Anthropic · Official documentation · 2026-08-20

    tool interface design · held-out tool evaluations · token-efficient results
  4. [4]
    Safety in building agents

    OpenAI · Official documentation · 2026-08-20

    prompt injection · structured data boundaries · MCP approvals

Related Tenten resources

When the lab reaches production

Bring the artifacts, not a blank brief.

A useful implementation review starts with your task fixtures, permission map, traces, eval report, failure cases, and cost ceiling. Tenten can review that evidence and help close the integration or operating gaps without reopening decisions the course already proved.