ChatOpenAIFeatured1 vendors

gpt-6-luna

New GPT88 featured GPT-6 Luna route for frequent, focused, and cost-sensitive model calls.

Overview

  • GPT-6 Luna is the lightweight-work route in OpenAI’s latest GPT-6 family, intended for bounded tasks where latency and cost control matter.
  • It is a practical candidate for everyday chat, classification, extraction, content processing, and bounded agent steps. Compare GPT-6 Astra or GPT-6 Sol for deeper reasoning and long-running execution.
  • This page does not hardcode dynamic context, limits, tools, vision support, or pricing. Confirm the current API key access with GET /v1/models and the GPT88 console.

Best Use Cases

  • Use it for frequent, latency-sensitive chat or support requests
  • Run bounded classification, extraction, rewriting, routing, and summarization tasks
  • Assign lightweight coding assistance or preprocessing to it as an agent step
  • Balance quality, latency, and cost on a representative production workload
  • Compare it with gpt-6-astra, gpt-6-sol, or your current default model on a fixed workload

Integration Notes

  • Use the GPT88 OpenAI-compatible base URL and set model to gpt-6-luna.
  • Call GET /v1/models to confirm that the current API key can access gpt-6-luna before sending a minimal non-streaming request.
  • Validate non-streaming and streaming first, then enable structured output, tools, and batch workloads as needed.
  • Record first-token latency, end-to-end latency, success rate, retries, and actual usage instead of inferring speed or cost from the model name.

Usage Notes

  • gpt-6-luna is a newly launched model and the current local backend catalog has not synced it yet. Access, groups, quota, pricing, limits, and routes are determined by the GPT88 console and API key.
  • A lightweight positioning does not guarantee lower latency or cost for every prompt, concurrency level, or route. Test representative workloads.
  • New-model behavior and parameter support can change after launch. Start with a staged rollout before changing all production traffic.
  • Payment, permissions, production writes, and other high-risk actions still require independent checks, isolation, and human authorization.

Capabilities

GPT-6 LunaFast response positioningFocused tasksHigh-frequency useAgent subtasks

Recommended Scenarios

  • High-frequency chat
  • Classification and extraction
  • Lightweight coding help
  • Content rewriting
  • Agent subtasks

Endpoint Path

POSThttps://api.gpt88.cc/v1/chat/completions

The endpoint path is determined by the model category. Chat, image, video, and audio models use different endpoints.

Send requests with Authorization: Bearer <API_KEY>. See Chat Completions API for the full parameter reference.

API Docs

Endpoint
POST /v1/chat/completions
Model ID
gpt-6-luna
Content-Type
application/json
Protocol
OpenAI-compatible chat completions endpoint with streaming and non-streaming output.
Base URL
Use https://api.gpt88.cc for OpenAI SDK, Cursor, and cURL.

Headers / Auth

  • AuthorizationstringRequired
    Send Bearer <GPT88_API_KEY>. Create the API key in the gpt88.cc console.
  • Content-TypestringRequired
    Request body format. Audio transcription uploads use multipart/form-data.
    Default:application/json
  • Acceptstring
    Non-streaming requests return JSON; streaming chat requests return SSE chunks.
    Default:application/json or text/event-stream

Chat Protocol Differences

For chat models, the protocol used by the client matters as much as the model name. Claude Code / Anthropic SDK and OpenAI SDK build different request paths.

OpenAI CompatibleRecommended
Base URL
https://api.gpt88.cc
Endpoint
POST /v1/chat/completions
Body
{ "model": "gpt-6-luna", "messages": [...] }
Anthropic / Claude
Base URL
https://api.gpt88.cc
Endpoint
POST /v1/messages
Body
{ "model": "gpt-6-luna", "messages": [...] }

This model is usually best accessed through the OpenAI-compatible protocol. Only use the Anthropic-style path when the target tool explicitly requires it.

Request Parameters

The fields below are generated by model category. Pricing, context limits, rate limits, and available routes should be confirmed in the gpt88.cc console.

  • modelstringRequired
    Current model ID: gpt-6-luna. Preserve the exact casing when copying.
  • messagesarray<Message>Required
    Conversation history. Each message includes role and content.
  • streamboolean
    Enable SSE streaming. Useful for chat interfaces and long responses.
    Default:false
  • temperaturenumber
    Sampling temperature. Higher values produce more variation; production workloads usually start with a fixed value.
    Default:1
  • max_tokensinteger
    Maximum output tokens for this request. Real context limits should be confirmed in the console.
  • response_formatobject
    Structured output config, for example { "type": "json_object" }.
  • toolsarray<Tool>
    Function-calling tool definitions. Availability depends on the model and account configuration.

Response Fields

Response bodies generally follow OpenAI-compatible shapes. Media models may return async task IDs or resource URLs.

  • idstringRequired
    Unique completion ID for tracing and debugging.
  • objectstringRequired
    Usually chat.completion or a streaming chunk type.
  • modelstringRequired
    Actual model ID that handled inference.
  • choicesarray<Choice>Required
    Generated choices including message, delta, and finish_reason.
  • usageobject
    Token usage statistics. Streaming calls may return this only at the end or in non-streaming mode.

Status Codes

  • 200OKRequired
    Request succeeded. See the response field table above for the response body shape.
  • 400Bad Request
    Invalid request fields, such as missing required fields, unsupported image size format, or unsupported file type.
  • 401Unauthorized
    API key is missing, invalid, or malformed.
  • 404Not Found
    Endpoint or model not found. Confirm the path is /v1/chat/completions and the model ID is gpt-6-luna.
  • 429Rate Limited
    Rate limit, concurrency cap, or insufficient balance. Check the console for the current allowance.
  • 5xxUpstream Error
    Upstream or request failure. Retry with the same Base URL and keep the request ID for debugging.

Errors and Troubleshooting

  • 401: verify Authorization is Bearer <API_KEY> and the key is still valid.
  • 404: confirm the endpoint is /v1/chat/completions and the model ID gpt-6-luna is available in the console.
  • 429: reduce concurrency, shorten requests, or check balance and quota in the console.
  • 5xx: retry with the same Base URL and keep the request ID for debugging.

Request Examples

curl https://api.gpt88.cc/v1/chat/completions \
  -H "Authorization: Bearer $GPT88_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "messages": [
      {"role": "user", "content": "Introduce gpt88.cc in one sentence."}
    ]
  }'

Expected response (trimmed example):

200 OKjson
{
  "id": "chatcmpl-gpt-6-lu",
  "object": "chat.completion",
  "created": 1730000000,
  "model": "gpt-6-luna",
  "choices": [
    {
      "index": 0,
      "finish_reason": "stop",
      "message": { "role": "assistant", "content": "..." }
    }
  ],
  "usage": { "prompt_tokens": 24, "completion_tokens": 64, "total_tokens": 88 }
}