Reliable generation and recovery
The gateway supports OpenAI-compatible Chat Completions and a documented Responses
subset across enabled providers. Check /v1/models for the exact model contract.
Protocol compatibility does not imply every OpenAI hosted tool or stateful feature
is available on every provider.
Output outcomes
Chat output that reaches its token limit returns HTTP 200, its available content,
and finish_reason: "length". Responses returns status: "incomplete" with
incomplete_details.reason: "max_output_tokens". This replaces the previous 422
behavior. Refusals and filtered outcomes remain explicit. Applications must check
the finish reason/status before parsing JSON or executing generated tools.
JSON and JSON Schema can be requested with streaming when the contract advertises
structured_output.streaming. Deltas are provisional: the complete value is
validated before successful termination. Invalid completed structured output
returns an upstream error (502 before response headers, or an error event after
streaming starts). A missing terminal event is a failure, never successful EOF.
Anthropic JSON Schema uses native constrained output. Generic JSON-object mode uses
a closed schema envelope that the gateway unwraps and validates. It is buffered;
the contract declares structured_output.delivery.json_object: "buffered".
Caller schemas are preferable when an application knows the required shape.
Claude effort is exposed only for exact documented model IDs. Send Chat
reasoning_effort, or Responses reasoning.effort, using the model's advertised
values. Usage preserves provider-reported reasoning details. If the provider does
not report a separate reasoning token count, it is omitted; the gateway does not
estimate a split or subtract reasoning tokens from completion totals.
Recoverable long work
Use the standard Responses background workflow for work that must survive the client disconnecting:
from openai import OpenAI
client = OpenAI(base_url="https://YOUR_GATEWAY/v1", api_key="YOUR_KEY")
response = client.responses.create(
model="YOUR_MODEL", input="Your task", background=True,
extra_headers={"Idempotency-Key": "unique-operation-key"},
)
response = client.responses.retrieve(response.id)To stream and resume, create with background=True, stream=True. Save the Response
ID from response.created and the latest processed sequence_number. Reconnect
using GET /v1/responses/{id}?stream=true&starting_after=N, or Last-Event-ID: N.
SSE IDs equal event sequence numbers. Deduplicate events by sequence number when
reconnecting. Replay retrieves the original generation; it does not regenerate it.
POST /v1/responses/{id}/cancel cancels queued work immediately or requests
cancellation of running work. A running request may finish before cancellation is
observed. Cancellation does not guarantee the upstream vendor stopped billing.
Responses retrieval and replay require the creating API key and organization.
store defaults to true, with the project's configured content retention.
Foreground store: false prevents retained response content. Background
store: false retains recovery data for ten minutes. Expired retained results
return 410. Stream events and response snapshots are encrypted at rest.
Workers renew an ownership lease every 15 seconds. The lease expires after 60 seconds and maintenance checks every 15 seconds. A lost worker produces a failed result with committed partial output; an uncertain dispatch is never retried automatically. This bound assumes a healthy database and maintenance worker. Gateway-owned background work survives subscriber or API replica loss. An upstream stream failure cannot be turned into a completed answer that the upstream never made available.
Foreground Chat streams also retain accumulated output for operation lookup at
GET /v1/requests/{X-The-Code-Request-ID}. Failed lookups include available
error.partial_response. Foreground connections do not provide durable execution;
use background Responses for that guarantee.
Compatibility and limits
- Reuse an idempotency key only for the identical request. Different bodies return 409. Retrying an uncertain operation returns its recorded failure rather than creating another generation.
- Responses streaming supports function tools. Native hosted tools can use the
supported non-streaming provider route.
previous_response_idis not supported; send explicit conversation input and retain required tool metadata. - Default stream no-progress timeout is 300 seconds, and total stream execution timeout is 1,800 seconds. Deployment settings can adjust both. SSE heartbeats do not count as model progress. Upstream HTTP read limits apply independently.
- SSE events are bounded to 8 MiB and accumulated wire output to 32 MiB. Oversized, malformed, or unterminated provider events fail explicitly.
- Project background queues default to 256 waiting requests. Overload returns 429
with
Retry-After. Provider HTTP connection pools are isolated by provider; saturation of one provider's pool does not consume another provider's slots. - Each API/worker replica admits at most 64 concurrent requests per provider credential. Five consecutive upstream availability failures open that credential's circuit for 30 seconds, followed by one recovery probe. These deployment settings are configurable. Existing project/key limits still apply.
X-Request-IDidentifies an HTTP request (legacy caller-supplied values remain accepted).X-Client-Request-IDsupports caller correlation.X-The-Code-Request-IDidentifies the logical operation. These are CORS-exposed.
Deployment diagnostics
GET /v1/meta returns build_commit, build_time, API version, and contract profile.
Image builds embed their immutable commit. Development builds may say unknown.
Use /health/live, /health/startup, and /health/ready for their respective probes.
This release requires the durable-response database migration and running workers. After deployment, refresh/certify provider models so live contracts reflect current adapter behavior. Validate the expected commit and a background create/retrieve/ resume/cancel smoke test through the public edge before promoting the release.