Claude Haiku 5.5 and Nano Banana 2.1 are now live on CometAPI โ†’
guide/CometAPI research

How to Use Claude Haiku 5.5 API: Complete Developer Guide

Learn how to use Claude Haiku 5.5 API with CometAPI. Explore setup, code examples, streaming, effort settings, pricing, and migration.

CometAPI
Deon GoodwinAI model and API research team
Updated Oct 10, 2026 16 min read
How to Use Claude Haiku 5.5 API: Complete Developer Guide
Use this pattern

Make the first API call.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_COMETAPI_KEY",
    base_url="https://api.cometapi.com/v1",
)

response = client.chat.completions.create(
    model="gpt-5-mini",
    messages=[{"role": "user", "content": "Build this workflow."}],
)

print(response.choices[0].message.content)

TLDR Claude Haiku 5.5 calling it the cheapest, fastest, and most capable small model it has ever shipped. Designed for high-volume, latency-sensitive workloads such as classification, extraction, routing, summarization, and subagent tasks, it delivers a 1-million-token context window, up to 128K output tokens, adaptive thinking with adjustable effort levels.

On average it costs around 75% less to run than Haiku 4.5 while posting large gains on agentic and computer-use benchmarks. Developers can access it via the official Claude API (claude-haiku-5-5), Amazon Bedrock, Google Cloud, Microsoft Foundry, or cost-optimized gateways such as CometAPI (starting at $0.08 / $0.40).

Key Takeaways

  • Release & Positioning: Launched October 7, 2026 as the small, high-throughput model in the Claude 5.5 familyโ€”ideal for real-time support, document processing, and delegated agent steps.
  • Core Specs: 1M-token context, 128K max output (300K via Batch API beta), adaptive thinking (default medium effort), text + image input โ†’ text output, knowledge cutoff June 2026.
  • Pricing Advantage: $0.10 input / $0.50 output per million tokens for prompts up to 100K (โ‰ˆ90% cheaper than Haiku 4.5 on most requests); higher tier above 100K. Cache reads as low as $0.01. CometAPI offers โ‰ˆ20% lower Standard rates.
  • Performance Leap: Major gains vs Haiku 4.5 (e.g., OSWorld 72.4% vs 15.7%, Terminal-Bench 39.2% vs 0%, GDPval-AA 1620 vs 735 Elo) while remaining competitive with or ahead of GPT-6 Luna on many tests.
  • Developer Features: Effort parameter (lowโ€“max), prompt caching, Batch API (50% off), computer/browser use (beta SDK support). Temperature/top_p/top_k must be omitted or a 400 error occurs.

What Is Claude Haiku 5.5 and Why It Matters in 2026

Claude Haiku 5.5 is Anthropicโ€™s latest entry in the lightweight tier of the Claude family. Unlike the more expensive Opus 5.5 and Sonnet 5.5 models optimized for complex reasoning and long-horizon agents, Haiku 5.5 is purpose-built for high-volume, cost-sensitive, and latency-critical workloads. Anthropic describes it as โ€œthe cheapest, fastest, and most capable small model weโ€™ve ever released.โ€

Typical production uses include:

  • Ticket classification and routing in customer support
  • Field extraction and summarization from long documents
  • Compaction of conversation history for larger models
  • Database-style queries and structured data tasks
  • Fast subagents that handle narrow coding, tool use, or browser steps while a stronger model orchestrates

Because it is dramatically cheaper and faster than its predecessor while closing much of the quality gap on practical agentic tasks, many teams are re-architecting pipelines so that Haiku 5.5 handles the majority of requests and only escalates difficult cases to Sonnet or Opus. Early customer feedback from Asana, HubSpot, Box, and others highlights 30%+ latency reductions and measurable accuracy lifts on high-volume internal evals.

Technical Specifications of Claude Haiku 5.5

SpecificationValueNotes
Model ID (Claude API)claude-haiku-5-5No date suffix
Amazon Bedrockanthropic.claude-haiku-5-5 (or global/us/eu/au/jp profiles)
Context window1,000,000 tokensโ‰ˆ555k words with new tokenizer
Max output (Messages API)128,000 tokens300K with Batch API + beta header
Input modalitiesText + images
Output modalityText
ThinkingAdaptive (on by default)Controlled via effort
Default effortmediumOptions: low, medium, high, xhigh, max
Knowledge cutoffJune 2026
Retirement commitmentNot before October 7, 2027
TokenizerNewer (same family as Claude 4.7+)Same text โ‰ˆ30% more tokens than Haiku 4.5

Source: Anthropic model overview and platform documentation.

The larger context and output limits, combined with adaptive thinking, make Haiku 5.5 far more usable for real production systems than earlier small models that required aggressive truncation or fixed thinking budgets.

Responses and completion checks

Read content blocks by type, rather than assuming content[0] is the visible answer. Inspect stop_reason before accepting output: max_tokens indicates a truncated generation and refusal requires an explicit application response. For tool-enabled requests, process tool-use blocks and return tool results in native format instead of treating them as a completed text answer.

What Changed Compared With Claude Haiku 4.5 API?

Capacity, behavior and pricing changes

DimensionHaiku 4.5Haiku 5.5Sonnet 5.5
Context window200K1M1M
Maximum output64K128K128K
Adaptive effortNoYesYes
Anthropic input / MTok$1.00$0.10$2.00
Anthropic output / MTok$5.00$0.50$10.00
PositioningLegacy fast modelRoutine work at scaleHarder complex tasks

Haiku 5.5 prices in this comparison apply to prompts up to 100,000 tokens; larger prompts cost more. These are Anthropic Standard rates, not CometAPI rates. The comparison is a routing starting point, not a claim that the models perform identically.

For Haiku 4.5 integrations, the key implementation changes are adaptive thinking instead of a manual budget, effort controls, selective parsing of content blocks and updated token counts. The same text can use approximately 30% more input tokens. Recount representative prompts rather than reusing Haiku 4.5 token estimates. The migration section turns these changes into a deployment checklist.

Reported benchmark improvements

BenchmarkHaiku 4.5Haiku 5.5Sonnet 5.5
OSWorld 2.1 (offline subset; partial credit)15.7%72.4%83.9%
Terminal-Bench 4.00.0%39.2%70.6%
Humanityโ€™s Last Exam (no tools)10.2%45.9%56.9%
Chartography (no tools)6.4%46.4%61.6%
GDPval-AA v2.1 (Elo)7351,6201,840

Anthropic reports the scores above in its launch benchmark summary. They are reported evaluation results, not live tests of the CometAPI examples in this guide. OSWorld's 72.4% is partial credit, not a strict task-completion rate.

The Haiku 5.5 system card describes the evaluation conditions. Standard Haiku 5.5 evaluations use adaptive thinking at max effort and default sampling unless noted; the product default is medium. OSWorld uses 82 offline tasks, no internet access, 1080p resolution and a 500-action limit. Terminal-Bench 4.0 uses Claude Code in bare mode, no internet egress and 10 trials per task for Haiku 5.5; Sonnet's setup differs. HLE and Chartography rows above are tool-free. GDPval-AA reports an Elo rating from Artificial Analysis evaluations. Match the relevant harness, effort and tool settings when comparing results.

How to Use Claude Haiku 5.5 API: Complete Developer Guide

Anthropic's original OSWorld graphic shows the relationship between effort, per-task cost and partial-credit score. It does not measure API latency or guarantee the same performance on your workload.

The scores indicate much stronger computer-use and general-task performance than Haiku 4.5, but they do not guarantee the same results in your application. Run a representative evaluation set before committing production traffic.

The benchmark table and supplied graphic are retained as model comparisons. They are reported evaluations rather than live tests of the CometAPI examples. Use the same task distribution, effort and tool settings when evaluating an application; a benchmark score does not establish gateway latency or your own success rate.

How to Use Claude Haiku 5.5 API Through CometAPI

Create a key and choose the native route

Create a CometAPI account, confirm access to claude-haiku-5-5 and generate a server-side API key. Set COMETAPI_KEY in your environment. Keep the key outside source control and browser code. Anthropic and CometAPI keys belong to different providers; use the CometAPI key for every example in this revised guide.

export COMETAPI_KEY="your_cometapi_api_key"
pip install --upgrade anthropic
$env:COMETAPI_KEY="your_cometapi_api_key"
IntegrationSDK base URLRequest pathCredential
Anthropic-compatible Messageshttps://api.cometapi.com/v1/messagesCOMETAPI_KEY
OpenAI-compatible Chat Completionshttps://api.cometapi.com/v1/chat/completionsCOMETAPI_KEY

CometAPI supports Claudeโ€™s native Messages and content-block format. Use the Anthropic SDK with the CometAPI root base URL. Native Messages is the preferred route in this guide for Claude-specific thinking and content blocks. The compatible Chat Completions route is an option for an existing OpenAI-style application. Confirm advanced-field support on the exact gateway route.

Send a minimal request and verify the result

curl https://api.cometapi.com/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: $COMETAPI_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-haiku-5-5",
    "max_tokens": 2048,
    "output_config": {"effort": "low"},
    "messages": [{"role": "user", "content": "Summarize three AI uses in customer support."}]
  }'
import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["COMETAPI_KEY"],
    base_url="https://api.cometapi.com",
    timeout=60.0,
    max_retries=2,
)

response = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2048,
    output_config={"effort": "low"},
    messages=[{"role": "user", "content": "Classify this ticket: I cannot log in."}],
)
if response.stop_reason != "end_turn":
    raise RuntimeError(f"Unaccepted completion: {response.stop_reason}")
print("".join(block.text for block in response.content if block.type == "text"))
print(response.usage)

Save the Python example as example.py and run python example.py. A successful checkpoint is a completed response with visible text and usage fields. An authentication error means the key needs checking; a model/route error means the model ID, endpoint or account access needs checking.

Use the same native format from JavaScript

npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.COMETAPI_KEY,
  baseURL: "https://api.cometapi.com",
  timeout: 60_000,
  maxRetries: 2,
});
const response = await client.messages.create({
  model: "claude-haiku-5-5",
  max_tokens: 2048,
  output_config: {effort: "low"},
  messages: [{role: "user", content: "Extract product, quantity and price: 3 laptops at $900 each."}],
});
if (response.stop_reason !== "end_turn") {
  throw new Error("Unaccepted completion: " + response.stop_reason);
}
for (const block of response.content) {
  if (block.type === "text") console.log(block.text);
}

Use a supported Node.js release. Save the file as example.mjs, set COMETAPI_KEY and run node example.mjs. Record the SDK version you validate for deployment.

Adapt an existing OpenAI-compatible application

If your application already uses Chat Completions, install the openai package and configure the separate client below. This route returns choices instead of native content blocks. Keep route-specific parsing separate and test thinking, images and tools before relying on gateway translation.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["COMETAPI_KEY"],
    base_url="https://api.cometapi.com/v1",
)
response = client.chat.completions.create(
    model="claude-haiku-5-5",
    max_tokens=2048,
    messages=[
        {"role": "system", "content": "Be concise and helpful."},
        {"role": "user", "content": "Explain API rate limits."},
    ],
)
print(response.choices[0].message.content)

Sending images to Claude Haiku 5.5 API

Use an image together with an instruction that specifies the evidence and output you need. The official vision interface accepts image content blocks with base64 data, an image URL, or a supported uploaded file reference. This example reads a local PNG, places the image before the question, and asks for values that can be checked against the source.

Analyze a real local PNG through the native Anthropic Messages API.

import os
import base64
from pathlib import Path
import anthropic

# Replace dashboard.png with a real PNG file you are authorized to analyze.
image_data = base64.b64encode(Path("dashboard.png").read_bytes()).decode("ascii")
anthropic_client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
response = anthropic_client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2048,
    output_config={"effort": "low"},
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {
                "type": "base64", "media_type": "image/png", "data": image_data
            }},
            {"type": "text", "text": (
                "Read the visible dashboard metrics. List anomalies with their "
                "labels and values, and state when text is unreadable."
            )},
        ],
    }],
)
for block in response.content:
    if block.type == "text":
        print(block.text)

Use the MIME type that matches the actual file. Check readability, image dimensions and request-size limits before submission; avoid sending high-resolution media that does not help the task. For CometAPI, change the key and root base URL as shown in its Messages example, and verify image support on the exact route.

Controlling Claude Haiku 5.5 reasoning

Haiku 5.5 uses adaptive thinking by default. The official thinking configuration explains when disabled is accepted and how thinking blocks are returned. Set effort through output_config.effort; do not reuse legacy budget_tokens.

EffortGood starting use caseTrade-off
lowShort classification, simple extractionTypically lower token use and latency
mediumSupport responses, routine agent workBalanced default
highMulti-step document analysisPotentially more reasoning
xhighDifficult evaluationsMore processing and cost
maxReasoning-intensive casesHighest potential reasoning expenditure

Configure adaptive thinking with low effort.

import osimport anthropicโ€‹anthropic_client = anthropic.Anthropic( ย   api_key=os.environ["ANTHROPIC_API_KEY"],)โ€‹response = anthropic_client.messages.create( ย   model="claude-haiku-5-5", ย   max_tokens=4096, ย   thinking={"type": "adaptive"}, ย   output_config={"effort": "low"}, ย   messages=[{"role": "user", "content": ย  ย  ย   "Classify as billing, technical, or account: cannot log in."}],)for block in response.content: ย   if block.type == "text": ย  ย  ย   print(block.text)

Developers can disable thinking at low, medium and high effort, but not at xhigh or max. Thinking consumes billed output tokens and the max_tokens budget even when its text is hidden. An empty thinking field can still carry a signature; it does not prove that no reasoning occurred. Keep returned content blocks unchanged when continuing tool conversations.

Example request with thinking disabled

Send a native request body with thinking disabled.

{  "model": "claude-haiku-5-5",  "max_tokens": 1024,  "thinking": {"type": "disabled"},  "output_config": {"effort": "low"},  "messages": [ ย   {"role": "user", "content": "Return sentiment only: positive, neutral, negative."}  ]}

Streaming Claude Haiku 5.5 responses

Streaming delivers visible text incrementally and improves responsiveness in interactive applications. Treat streamed text as provisional until the request completes successfully.

Stream visible text with the Anthropic Python SDK.

import osimport anthropicโ€‹anthropic_client = anthropic.Anthropic( ย   api_key=os.environ["ANTHROPIC_API_KEY"],)with anthropic_client.messages.stream( ย   model="claude-haiku-5-5", ย   max_tokens=4096, ย   output_config={"effort": "low"}, ย   messages=[{"role": "user", "content": "Explain API caching."}],) as stream: ย   for text in stream.text_stream: ย  ย  ย   print(text, end="", flush=True)

Claude Haiku 5.5 API Pricing

Compare token rates and prompt-length bands

Rates checked October 8, 2026. The table compares Anthropic Standard pricing with published CometAPI rates, in USD per million tokens. The band is determined by prompt length: up to 100,000 tokens versus more than 100,000. Longer requests use the applicable higher-band rates, not merely a surcharge on the excess tokens. Account billing, cached usage and optional tools must be reconciled separately.

Token categoryAnthropic โ‰ค100KCometAPI โ‰ค100KAnthropic >100KCometAPI >100K
Input$0.10$0.08$0.50$0.40
Output$0.50$0.40$2.50$2.00
Cache read$0.01$0.008$0.05$0.04
5-min cache write$0.125$0.10$0.625$0.50
1-hour cache write$0.20$0.16$1.00$0.80

The original articleโ€™s rate snapshot is retained above. Check the current CometAPI listing and account billing before purchase or deployment; the listingโ€™s starting input rate is not a complete quotation for every prompt-length band, cache operation or optional tool.

Example: Cost of 100,000 support requests

Assume 100,000 monthly requests, each with 2,000 input tokens and 300 billed output tokens, including any generated thinking. Every prompt qualifies for the up-to-100K band. Exclude caching, tool charges and retries.

ComponentUsageAnthropicCometAPI
Input200M tokens$20.00$16.00
Output30M tokens$15.00$12.00
Total100,000 requests$35.00$28.00

Under these assumptions, the illustrative difference is $7 per month, or 20%. Production costs depend on request lengths, hidden reasoning, caching and retry behavior.

Use the modelโ€™s current token counts in the calculation. Cost = input tokens ร— applicable input rate / 1,000,000 + billed output tokens ร— applicable output rate / 1,000,000, plus separately applicable cache, tool and retry charges. Thinking tokens are part of billed output.

Reduce cost without losing accepted-task quality

OptimizationActionBenefit
Effort tuningUse low where quality allowsLower reasoning overhead
Prompt reductionRemove redundant conversation historyFewer input tokens
Prompt cachingKeep reused context stableLower repeated input cost
StreamingRender tokens as they arriveBetter perceived responsiveness
Schema validationReject malformed records earlyLess downstream rework
Model routingEscalate complex tasksBetter cost-quality balance
Batch processingBatch eligible offline requestsPotential additional savings

Do not automatically select the lowest effort. A cheaper individual call may increase total costs if it produces more retries. Evaluate accuracy, p50/p95 latency, token use and cost per successful task.

For prompt caching, reuse a stable prefix and inspect cache creation and cache read usage rather than assuming a hit. Haiku 5.5 has a 512-token minimum cacheable prompt on the documented Claude API. Changing effort can invalidate cached prefixes; keep settings stable within a conversation. A 1M context window is a capacity ceiling, not a recommendation to send every document on every turn.

Keep a stable reusable prefix for caching, shorten unnecessary inputs and benchmark effort changes. Streaming improves perceived responsiveness rather than automatically reducing billed usage. Compare cost per accepted task across the complete workflow, including retries and tool calls.

Claude Haiku 4.5 to Haiku 5.5 Migration Guide

Update the integration and request format

AreaMigration action
Model IDReplace with claude-haiku-5-5
ThinkingReplace manual budget_tokens with adaptive thinking
EffortUse output_config.effort if needed
SamplingOmit temperature, top_p and top_k; any top_k value is unsupported
Response parserHandle multiple content-block types
Output budgetAllow sufficient room for reasoning and final output
PrefillRemove unsupported assistant prefilling
CachingRetest prompts when changing effort
Tool integrationsVerify supported tool versions

For the CometAPI route used here, retain COMETAPI_KEY and the native base URL while changing the model ID and request configuration. Replace legacy enabled thinking with a budget by adaptive thinking and effort. Remove assistant prefilling and omit sampling parameters. Preserve the distinction between Claudeโ€™s native content array and the compatible routeโ€™s choices response.

Preserve state and recalculate budgets

The official Haiku 5.5 migration guide reports approximately 30% more input tokens for identical text than Haiku 4.5. Recount representative prompts and retune limits before moving traffic. Returned thinking blocks work only in the account that produced them or a linked account; an account change can silently drop those blocks. Preserve an append-only conversation prefix rather than modifying earlier messages after generation. Inspect stop_reason, including max_tokens and refusal, and handle them explicitly before accepting output.

Recount representative prompts with the target model and retune max_tokens, latency limits and request-cost estimates. Test stored conversations against the account and gateway route that will replay them; do not assume signed thinking blocks are portable across unrelated accounts or rewritten conversation prefixes.

Conclusion

Claude Haiku 5.5 represents a genuine inflection point for cost-efficient, high-throughput AI. By delivering dramatically better agentic and computer-use performance at roughly one-tenth the previous Haiku price on the majority of requests, Anthropic has made large-scale deployment of capable small models practical. Whether you call the official API, route through Amazon Bedrock, or use a unified gateway such as CometAPI for further savings and operational simplicity, the combination of speed, capability, and price opens new product possibilities that were previously uneconomical.

Continue learning

Connect this article to the next decision.

View all topics
Published on Oct 10, 2026
Last updated Oct 10, 2026
0 views
Reviewed for clarity, source attribution and current API terminology.

Read More