claude-haiku-5-5
Featured Anthropic Haiku model for low-latency, high-throughput, and bounded production workflows.
Overview
- Anthropic positions Claude Haiku 5.5 as its fastest, cheapest, and most capable small model, and says it costs about 75% less to run on average than Claude Haiku 4.5. This is provider positioning; validate actual behavior on a fixed workload.
- It is a practical candidate for frequent, bounded requests where response time and throughput matter. Use a stronger model for difficult long-context work, multi-step planning, or quality-sensitive final output.
- The public GPT88 model square has not synced Haiku 5.5 yet, so this page is backed by a local catalog patch. Confirm the exact API model ID, access, pricing, and route with the console and GET /v1/models.
Best Use Cases
- Use it for high-frequency support, ticketing, classification, extraction, and review assistance
- Process rewriting, summarization, routing, and structured text in batches
- Assign retrieval, judgment, formatting, or other bounded steps to it inside an agent
- Compare speed, quality, errors, and actual usage with Haiku 4.5 and your current default
Integration Notes
- Use the GPT88 Base URL and set model to claude-haiku-5-5 in an OpenAI-compatible request after confirming access for the current API key.
- Claude- or Anthropic-style tools should use the same GPT88 Base URL and the native request format required by that client.
- Validate non-streaming, streaming, structured output, and tools separately before adding batch or long-context workloads.
- Keep a tested Sonnet, Opus, or other lightweight fallback and record first-token latency, end-to-end latency, success rate, and usage.
Usage Notes
- The public announcement does not define GPT88 route pricing, context limits, rate limits, tools, or vision support; those fields are controlled by the current console configuration.
- Faster, cheaper, and more capable are release positioning claims, not guarantees for every prompt, concurrency level, or route. Test representative workloads.
- Canary a new model and keep a validated fallback until quality, latency, errors, and usage are established in production-like traffic.
Capabilities
Recommended Scenarios
- High-frequency chat
- Classification and extraction
- Support and ticket routing
- Lightweight agent subtasks
Endpoint Path
The endpoint path is determined by the model category. Chat, image, video, and audio models use different endpoints.
Send requests with Authorization: Bearer <API_KEY>. See Chat Completions API for the full parameter reference.
API Docs
Headers / Auth
| Field | Type | Required | Description |
|---|---|---|---|
Authorization | string | Required | Send Bearer <GPT88_API_KEY>. Create the API key in the gpt88.cc console. |
Content-Type | string | Required | Request body format. Audio transcription uploads use multipart/form-data.Default: application/json |
Accept | string | Optional | Non-streaming requests return JSON; streaming chat requests return SSE chunks. Default: application/json or text/event-stream |
AuthorizationstringRequiredSendBearer <GPT88_API_KEY>. Create the API key in the gpt88.cc console.Content-TypestringRequiredRequest body format. Audio transcription uploads usemultipart/form-data.Default:application/jsonAcceptstringNon-streaming requests return JSON; streaming chat requests return SSE chunks.Default:application/json or text/event-stream
Chat Protocol Differences
For chat models, the protocol used by the client matters as much as the model name. Claude Code / Anthropic SDK and OpenAI SDK build different request paths.
- Base URL
- https://api.gpt88.cc
- Endpoint
- POST /v1/chat/completions
- Body
- { "model": "claude-haiku-5-5", "messages": [...] }
- Base URL
- https://api.gpt88.cc
- Endpoint
- POST /v1/messages
- Body
- { "model": "claude-haiku-5-5", "messages": [...] }
This model is in the Claude family. Claude Code, Anthropic SDK, OpenClaw, and similar tools should prefer the Anthropic / Claude path.
Request Parameters
The fields below are generated by model category. Pricing, context limits, rate limits, and available routes should be confirmed in the gpt88.cc console.
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Required | Current model ID: claude-haiku-5-5. Preserve the exact casing when copying. |
messages | array<Message> | Required | Conversation history. Each message includes role and content. |
stream | boolean | Optional | Enable SSE streaming. Useful for chat interfaces and long responses. Default: false |
temperature | number | Optional | Sampling temperature. Higher values produce more variation; production workloads usually start with a fixed value. Default: 1 |
max_tokens | integer | Optional | Maximum output tokens for this request. Real context limits should be confirmed in the console. |
response_format | object | Optional | Structured output config, for example { "type": "json_object" }. |
tools | array<Tool> | Optional | Function-calling tool definitions. Availability depends on the model and account configuration. |
modelstringRequiredCurrent model ID:claude-haiku-5-5. Preserve the exact casing when copying.messagesarray<Message>RequiredConversation history. Each message includesroleandcontent.streambooleanEnable SSE streaming. Useful for chat interfaces and long responses.Default:falsetemperaturenumberSampling temperature. Higher values produce more variation; production workloads usually start with a fixed value.Default:1max_tokensintegerMaximum output tokens for this request. Real context limits should be confirmed in the console.response_formatobjectStructured output config, for example{ "type": "json_object" }.toolsarray<Tool>Function-calling tool definitions. Availability depends on the model and account configuration.
Response Fields
Response bodies generally follow OpenAI-compatible shapes. Media models may return async task IDs or resource URLs.
| Field | Type | Required | Description |
|---|---|---|---|
id | string | Required | Unique completion ID for tracing and debugging. |
object | string | Required | Usually chat.completion or a streaming chunk type. |
model | string | Required | Actual model ID that handled inference. |
choices | array<Choice> | Required | Generated choices including message, delta, and finish_reason. |
usage | object | Optional | Token usage statistics. Streaming calls may return this only at the end or in non-streaming mode. |
idstringRequiredUnique completion ID for tracing and debugging.objectstringRequiredUsuallychat.completionor a streaming chunk type.modelstringRequiredActual model ID that handled inference.choicesarray<Choice>RequiredGenerated choices includingmessage,delta, andfinish_reason.usageobjectToken usage statistics. Streaming calls may return this only at the end or in non-streaming mode.
Status Codes
| Field | Type | Required | Description |
|---|---|---|---|
200 | OK | Required | Request succeeded. See the response field table above for the response body shape. |
400 | Bad Request | Optional | Invalid request fields, such as missing required fields, unsupported image size format, or unsupported file type. |
401 | Unauthorized | Optional | API key is missing, invalid, or malformed. |
404 | Not Found | Optional | Endpoint or model not found. Confirm the path is /v1/chat/completions and the model ID is claude-haiku-5-5. |
429 | Rate Limited | Optional | Rate limit, concurrency cap, or insufficient balance. Check the console for the current allowance. |
5xx | Upstream Error | Optional | Upstream or request failure. Retry with the same Base URL and keep the request ID for debugging. |
200OKRequiredRequest succeeded. See the response field table above for the response body shape.400Bad RequestInvalid request fields, such as missing required fields, unsupported image size format, or unsupported file type.401UnauthorizedAPI key is missing, invalid, or malformed.404Not FoundEndpoint or model not found. Confirm the path is/v1/chat/completionsand the model ID isclaude-haiku-5-5.429Rate LimitedRate limit, concurrency cap, or insufficient balance. Check the console for the current allowance.5xxUpstream ErrorUpstream or request failure. Retry with the same Base URL and keep the request ID for debugging.
Errors and Troubleshooting
401: verifyAuthorizationisBearer <API_KEY>and the key is still valid.404: confirm the endpoint is/v1/chat/completionsand the model IDclaude-haiku-5-5is available in the console.429: reduce concurrency, shorten requests, or check balance and quota in the console.5xx: retry with the same Base URL and keep the request ID for debugging.
Request Examples
curl https://api.gpt88.cc/v1/chat/completions \
-H "Authorization: Bearer $GPT88_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"messages": [
{"role": "user", "content": "Introduce gpt88.cc in one sentence."}
]
}'Expected response (trimmed example):
{
"id": "chatcmpl-claude-h",
"object": "chat.completion",
"created": 1730000000,
"model": "claude-haiku-5-5",
"choices": [
{
"index": 0,
"finish_reason": "stop",
"message": { "role": "assistant", "content": "..." }
}
],
"usage": { "prompt_tokens": 24, "completion_tokens": 64, "total_tokens": 88 }
}