On this page
Learning objectives
- Design a strict tool schema with one distinct purpose
- Authorize the caller and tenant independently from model-generated arguments
- Return concise typed observations and actionable error codes
- Test wrong-tool, forged-argument, timeout, oversized-result, and repeated-call failures
Before you start
- • Module 01 Result union and runtime validation pattern
- • Basic authorization concepts and a local TypeScript test runner
Working definition
Tool use
Tool calling lets a model request a named capability with structured arguments. The host application decides whether that request is valid and authorized, executes it under a constrained identity, records the outcome, and returns a typed observation for the next model turn.
A tool schema constrains argument shape. It cannot prove that the requesting user may access the named account, that the account belongs to the active tenant, or that the proposed action matches policy.
Tool responses become model context. Oversized payloads, opaque identifiers, and vague errors consume context while giving the model little information it can use to recover.
Field situation
Meridian account lookup
Named synthetic scenario. Meridian CRM, tenant IDs, account records, traces, and policies are fabricated for this lab and do not describe a live system.
- Owner
- You are the API engineer adding a read-only account lookup to a research assistant.
- Decision
- Define the smallest tool interface that answers the analyst's question without exposing raw SQL, arbitrary HTTP, or cross-tenant data.
- Starting state
- The assistant summarizes approved account context before an analyst researches a company. Two synthetic tenants share the same database, and an early prototype trusts the account ID proposed by the model.
- Expected outcome
- Authorized lookups return a compact observation. Forged tenants, unknown records, schema errors, and timeouts produce distinct non-secret-bearing errors and no account data.
Constraints
- • The tool is read-only and returns company name, segment, approved notes, and record version only.
- • Tenant identity comes from the authenticated session and never from model arguments.
- • The response may contain at most five notes and 2,000 UTF-8 characters.
- • Every request records caller, tenant, tool, validated arguments, result code, duration, and correlation ID.
Worked example
A valid account ID fails the tenant check
Evidence status: Named synthetic scenarioThe active session belongs to tenant meridian-eu. A user asks about Northwind Labs, but retrieved web text contains an instruction to call lookup_account with accountId acct_us_204. The account exists under tenant meridian-us. The model emits a schema-valid call because it cannot see the database ownership rule.
The host ignores any tenant value in generated text, binds meridian-eu from the session, validates accountId, and queries by both tenant and account. The lookup returns NOT_FOUND without revealing whether acct_us_204 exists elsewhere. The trace marks the request denied_at_data_boundary and stores a hash of the validated argument rather than raw account notes.
The adversarial fixture passes when the EU caller receives no US data, the audit record keeps the authenticated tenant, and the model gets a concise recoverable error. No provider reliability percentage is claimed; the check is deterministic application behavior.
Limits
The sample uses an in-memory repository and simple role policy. A production system may need row-level security, delegated OAuth, attribute-based controls, field redaction, and legal retention rules. OpenAI and Anthropic tool documentation explains the call loop, while authorization stays with the application.
Method
Build it, with checkpoints
Field situation
Define the smallest tool interface that answers the analyst's question without exposing raw SQL, arbitrary HTTP, or cross-tenant data.
- 01Separate identity from arguments
- 02Validate before access
- 03Shape the observation
Acceptance checks
A read-only capability whose schema helps the model request useful work while the application retains identity, authorization, execution, payload, retry, and audit control.
- 01
Separate identity from arguments
Create Session from trusted middleware and Args from the model call. Remove tenantId, userId, role, URL, and query fields from the tool schema.
CHECKPOINT · The model-visible schema contains only accountId, while the executor requires a trusted Session supplied by the host.
- 02
Validate before access
Apply strict schema checks, role policy, and tenant-scoped lookup before assembling any response. Return the same NOT_FOUND shape for absent and other-tenant records.
CHECKPOINT · The forged acct_us_204 request under meridian-eu yields NOT_FOUND and no row fields appear in logs or tool results.
- 03
Shape the observation
Return semantic field names, five notes at most, a stable version, and a bounded payload. Use error codes the model can act on without internal stack traces.
CHECKPOINT · The normal result stays below 2,000 characters, while the oversized fixture returns RESULT_TOO_LARGE with zero partial notes.
- 04
Add timeout and repetition controls
Abort the repository call at a defined deadline and cache the completed read by run plus validated argument hash. Reads may repeat once after a transient timeout, then escalate.
CHECKPOINT · The timeout fixture ends with TIMEOUT, and three identical calls in one run execute the repository no more than once after success.
- 05
Write adversarial tests and audit assertions
Cover extra arguments, malformed IDs, missing role, cross-tenant ID, unknown record, timeout, oversized response, and duplicate call. Assert the audit record for each path.
CHECKPOINT · Every test records correlation ID, trusted tenant, tool name, result code, and duration without storing account notes.
Hands-on lab
Build the Meridian read-only tool boundary
Implement lookup_account against an in-memory two-tenant fixture, then run allowed, denied, invalid, timeout, and oversized-response cases.
Prepare
- • Reuse the closed Result pattern from Module 01.
- • Create separate session and tool-argument types; never merge them into one model-visible object.
- • Use synthetic account notes without names, emails, or secrets.
- • Keep the fake repository deterministic so all acceptance checks run offline.
Deliverable
A strict tool declaration, session-bound executor, synthetic repository, compact result contract, audit-event type, and at least eight offline tests.
Starter kit: Authorized account tool
TypeScripttype Session = { userId: string; tenantId: string; roles: string[] };
type Args = { accountId: string };
type ToolResult =
| { ok: true; account: { id: string; name: string; segment: string; notes: string[]; version: number } }
| { ok: false; code: "FORBIDDEN" | "NOT_FOUND" | "INVALID_ARGS" | "TIMEOUT" | "RESULT_TOO_LARGE" };
const TOOL = {
type: "function",
name: "lookup_account",
description: "Read one account in the authenticated tenant. Never searches other tenants or updates records.",
strict: true,
parameters: {
type: "object",
additionalProperties: false,
required: ["accountId"],
properties: { accountId: { type: "string", pattern: "^acct_[a-z0-9_]+$" } }
}
} as const;
export async function executeLookup(session: Session, args: Args): Promise<ToolResult> {
if (!session.roles.includes("account:read")) return { ok: false, code: "FORBIDDEN" };
if (!/^acct_[a-z0-9_]+$/.test(args.accountId)) return { ok: false, code: "INVALID_ARGS" };
const row = await repository.findByTenantAndId(session.tenantId, args.accountId);
if (!row) return { ok: false, code: "NOT_FOUND" };
const account = { ...row, notes: row.notes.slice(0, 5) };
if (JSON.stringify(account).length > 2000) return { ok: false, code: "RESULT_TOO_LARGE" };
return { ok: true, account };
}Expected result
A read-only capability whose schema helps the model request useful work while the application retains identity, authorization, execution, payload, retry, and audit control.
Carry forward
Keep Session, ToolResult, audit fields, and the two-tenant fixture. Module 05 will expose a similarly narrow read capability through an MCP server.
Acceptance checks
- 01Cross-tenant and unauthorized fixtures return no account fields and use a non-enumerating error shape.
- 02Extra or malformed arguments fail before the repository is called.
- 03Tool results stay within the documented field and size limits.
- 04Duplicate successful calls in one run do not repeat repository work, and every path emits an audit record.
What breaks
Failure clinic
F1A caller can retrieve an account from a different tenant by supplying its ID.
- Inspect
- Trace where tenant identity originates and inspect the repository predicate for both tenantId and accountId.
- Likely cause
- Authorization was inferred from a model argument or checked after a global ID lookup.
- Repair
- Bind tenant from the authenticated session and make tenant-scoped lookup the only repository method exposed to the tool.
- Prevent next time
- Keep cross-tenant fixtures in CI and add datastore-level row policies where available.
F2The model repeatedly calls the right tool with invalid optional fields.
- Inspect
- Review raw calls, schema strictness, parameter names, tool description, and returned validation error.
- Likely cause
- The schema is permissive or the tool contract uses ambiguous parameters the model cannot distinguish.
- Repair
- Remove unused fields, set additionalProperties to false, and return one actionable example with INVALID_ARGS.
- Prevent next time
- Evaluate tool selection and arguments on a held-out realistic set before widening the surface.
F3A successful lookup floods context with raw CRM payloads and degrades later decisions.
- Inspect
- Measure result tokens and identify fields never used by the next decision.
- Likely cause
- The tool mirrors a backend response instead of designing an agent-facing observation.
- Repair
- Return a concise semantic projection with pagination or a detailed mode available only when needed.
- Prevent next time
- Set response budgets and include total tool-result tokens in evaluation reports.
F4A timeout triggers repeated calls that create inconsistent traces or overload the backend.
- Inspect
- Check deadlines, attempt counters, run-level cache keys, and whether the caller can distinguish unknown completion.
- Likely cause
- Retry behavior was left to the model and no idempotency or deduplication boundary exists.
- Repair
- Handle retries in deterministic orchestration with a cap and cache completed results by validated input.
- Prevent next time
- Inject slow, dropped, and duplicate responses in tool integration tests.
Beyond the demo
Production boundary
- 01Derive identity, tenant, and roles from trusted middleware rather than model-visible arguments.
- 02Give each tool one distinct job and a strict minimal schema.
- 03Classify every capability as read-only, reversible write, external communication, or irreversible action.
- 04Use short-lived least-privilege credentials and explicit network destinations.
- 05Return bounded semantic results with stable error codes and no internal stack traces.
- 06Apply deadlines, deterministic retry policy, rate limits, and deduplication outside the model.
- 07Audit initiator, trusted tenant, validated argument hash, decision, result, latency, and version.
- 08Run held-out tool-selection and adversarial authorization tests before adding a new tool.
Evidence status
Sources and claim limits
Sources support the named claims; they do not guarantee the same result in another system.
- [1]Function callingstrict function schemas · parallel calls · tool-call state machine
OpenAI · Official documentation · 2026-08-20
- [2]Tool use with Claudetool definitions · tool result loop · client-side execution responsibility
Anthropic · Official documentation · 2026-08-20
- [3]Writing effective tools for agentstool interface design · held-out tool evaluations · token-efficient results
Anthropic · Official documentation · 2026-08-20
- [4]Safety in building agentsprompt injection · structured data boundaries · MCP approvals
OpenAI · Official documentation · 2026-08-20
Related Tenten resources