API
At most ten steps, in order of how much they help. Each step is from Anthropic's own pages, linked. Re-checked against those pages every week; anything that can't be confirmed is marked unverified rather than removed.
- 01
Pick a model with your own tests, pin its ID, and plan for retirement
Anthropic suggests starting with Claude Opus 5.5 (
claude-opus-5-5) for most workloads, or starting efficiency-first with Claude Haiku 4.5 and moving up only if your tests show a gap. Every model ID is a pinned snapshot, including dateless ones such asclaude-sonnet-5-5; Anthropic ships updates under a new ID. Older aliases such asclaude-sonnet-4-5do move, so use the full ID. Keep the ID in config, not scattered through code. Models are deprecated with at least 60 days' notice, and requests to a retired model fail. For example,claude-sonnet-4-5-20250929retires on 30 November 2026, withclaude-sonnet-5-5as the replacement. To find old models in use, export the CSV on the Console Usage page, which breaks usage down by API key and model. Re-run your tests before you switch. (models overview, choosing a model, model IDs and versioning, model deprecations, migration guides) - 02
Build an evaluation set before you tune anything
Anthropic calls a good evaluation set "the most important step" in deciding whether to change models. Write specific, measurable success criteria, often across several dimensions such as accuracy, tone, latency and cost. Collect test cases that look like real traffic, plus edge cases: irrelevant, overly long or ambiguous input. Grade automatically where you can, with code checks or a model as grader. Use a different model to grade than the one that produced the output. Many cases with rough automatic grading beat a few graded by hand. Run the set on every change of model, effort level or prompt. (define success and build evals, choosing a model)
- 03
Cache the parts of the request that do not change
Put stable content first, in the order tools, then system, then messages, and anything that changes per request (timestamps, the user's message) after it. Add a top-level
cache_control: {"type": "ephemeral"}for automatic caching in conversations, or placecache_controlon the last block that is identical across requests. The cache lasts 5 minutes by default;"ttl": "1h"gives an hour. Prompts below a minimum length are not cached and no error is returned, so checkcache_read_input_tokensandcache_creation_input_tokensinusage. Changing tools, thinking settings or top-leveloutput_config.effortbreaks the cache from that point. For most models, cache reads also do not count towards your input-tokens-per-minute rate limit. Caches are kept separate per workspace. (prompt caching, rate limits, workspaces) - 04
Get machine-readable output with structured outputs and strict tools
When your code parses Claude's reply, set
output_config.formatwith a JSON schema rather than parsing free text. When Claude calls your functions, addstrict: trueto each tool so its inputs always match the schema. Both are generally available with no beta header. The SDKs can build the schema for you:client.messages.parse()with Pydantic in Python, or Zod in TypeScript. Know the limits: no recursive schemas,additionalPropertiesmust befalse, no numeric or string length constraints, and at most 20 strict tools per request. A refusal or amax_tokensstop can still return output that does not match, so checkstop_reasonbefore you parse. For tools, either run the loop yourself (tool_usein,tool_resultback) or let the SDK's Tool Runner do it. (structured outputs, tool use overview, stop reasons) - 05
Set effort on purpose and give
max_tokensroomEffort (
output_config.effort:low,medium,high,xhigh,max) is the main control for cost, speed and depth on current models. It affects all output, including thinking and tool calls. Defaults differ:mediumon Claude Opus 5.5,highon most others, so set it explicitly and choose it with an effort sweep on your evaluation set rather than copying an old setting. Uselowfor simple, high-volume or latency-sensitive calls.max_tokensis a hard limit on thinking plus reply, so set it large at higher effort. Thinking is adaptive and always on for some models; sendingthinking: {"type": "disabled"}to them returns a 400 error. From Claude Opus 4.7 onwards, settingtemperature,top_portop_kto a non-default value also returns a 400 error. (effort, models overview, errors, model deprecations) - 06
Stream anything long
Set
"stream": true, or use the SDK'smessages.stream()helper, so users see text as it arrives. For long or large-max_tokensrequests, streaming is not optional: the SDKs refuse non-streaming requests expected to run past 10 minutes, and idle connections can be dropped. If you do not need the text as it arrives, callget_final_message()(TypeScript:finalMessage()) to get the complete message over a stream. Errors can arrive mid-stream after a 200 response, so handleerrorevents too. To resume an interrupted stream on Claude 4.6 and later, send the partial text back in a user message and ask Claude to continue. (streaming, errors: long requests) - 07
Handle errors, rate limits and stop reasons properly
The SDKs retry connection errors, 429s and 5xx errors twice by default with exponential backoff and honour
retry-after; setmax_retriesto suit your app. Catch the SDK's typed exceptions rather than matching message text. A 529 means the API is overloaded. A 429 witherror_codeenforced_spend_limit_reachedand noretry-afteris the monthly spend cap, and retrying will not help. Limits use a token bucket, so short bursts can trip them: ramp traffic up gradually and watch theanthropic-ratelimit-*response headers. Log therequest-idheader for support. Checkstop_reasonon every response:max_tokensmeans the reply was cut off,refusalmeans Claude declined, andtool_usemeans your code must run a tool. (errors, rate limits, stop reasons) - 08
Send work that can wait to the Message Batches API
Batches suit bulk jobs such as evaluations, classification and back-fills. They cost 50% less than standard calls and have their own rate limits. Most batches finish within an hour; any request not done in 24 hours expires and is not billed. A batch holds up to 100,000 requests or 256 MB. Give each request a unique
custom_id, because results may not come back in order. Download results within 29 days. With shared context across a batch, use the 1-hour cache lifetime, since batches often take longer than 5 minutes. (batch processing, rate limits) - 09
Measure tokens and cap spend
The token counting endpoint (
messages.count_tokens) takes the same input as a message and is free, with its own rate limit; the count is an estimate. Claude Opus 4.7 and later use a newer tokenizer that gives roughly 30% more tokens for the same text, so recount when you migrate rather than reusing old figures. Log theusageobject from each response. Total input iscache_read_input_tokens+cache_creation_input_tokens+input_tokens. Set your own spend limit under Settings > Billing, and lower spend and rate limits per workspace so one project cannot use up another's share. The Console Usage page shows your cache hit rate. (token counting, rate limits, workspaces) - 10
Keep API keys on the server, scoped and short-lived
Never ship a key in browser or mobile code. Read it from the
ANTHROPIC_API_KEYenvironment variable or a secrets manager, rotate it, and disable or delete any key you think has leaked. Set an expiry when you create a key. Use a service account key for shared or production workloads, not a personal key, which stops working when its owner leaves; older workspace keys are now legacy. For production on a cloud platform or in CI, Workload Identity Federation swaps static keys for short-lived tokens. For iOS and macOS apps that call Claude directly, App Attest issues short-lived tokens to genuine installs. Use separate workspaces for development, staging and production, each with its own keys and limits. (authentication, workspaces)
Unverified
What Anthropic's own pages disagree on, or could not be confirmed.
- Claude Haiku 4.5's retirement is listed as "not sooner than October 15, 2026", eleven days after this check. No deprecation notice was on the deprecations page on 2026-10-04. If you start efficiency-first on Haiku 4.5, watch that page.
- The tool use overview says you can require a tool call by setting
tool_choice. The errors page says Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1 rejecttool_choiceof typeanyortoolwith a 400 error, and suggestsautowith strict tools or structured outputs instead. Step 4 follows the errors page. - The structured outputs page names
output_config.formatas the parameter and callsoutput_formatdeprecated, but its Pythonmessages.parse()example passesoutput_format=. Check your SDK version's docs for the exact helper argument. - Changing effort partway through a conversation without breaking the cache (a per-message
output_config) is in beta and only on some models, so it is left out of Step 5. - Fast mode (
speed: "fast") is a research preview with premium pricing and separate rate limits; it was not read in full, so it is not a step. - The multi-model "advisor" and "orchestrator" patterns on the cost and intelligence page were not read, so they are not a step. (choosing a model links to them.)