# Reliable generation and recovery

The gateway supports OpenAI-compatible Chat Completions and a documented Responses
subset across enabled providers. Check `/v1/models` for the exact model contract.
Protocol compatibility does not imply every OpenAI hosted tool or stateful feature
is available on every provider.

## Output outcomes

Chat output that reaches its token limit returns HTTP 200, its available content,
and `finish_reason: "length"`. Responses returns `status: "incomplete"` with
`incomplete_details.reason: "max_output_tokens"`. This replaces the previous 422
behavior. Refusals and filtered outcomes remain explicit. Applications must check
the finish reason/status before parsing JSON or executing generated tools.

JSON and JSON Schema can be requested with streaming when the contract advertises
`structured_output.streaming`. Deltas are provisional: the complete value is
validated before successful termination. Invalid completed structured output
returns an upstream error (502 before response headers, or an error event after
streaming starts). A missing terminal event is a failure, never successful EOF.

Anthropic JSON Schema uses native constrained output. Generic JSON-object mode uses
a closed schema envelope that the gateway unwraps and validates. It is buffered;
the contract declares `structured_output.delivery.json_object: "buffered"`.
Caller schemas are preferable when an application knows the required shape.

Claude effort is exposed only for exact documented model IDs. Send Chat
`reasoning_effort`, or Responses `reasoning.effort`, using the model's advertised
values. Usage preserves provider-reported reasoning details. If the provider does
not report a separate reasoning token count, it is omitted; the gateway does not
estimate a split or subtract reasoning tokens from completion totals.

## Recoverable long work

Use the standard Responses background workflow for work that must survive the
client disconnecting:

```python
from openai import OpenAI
client = OpenAI(base_url="https://YOUR_GATEWAY/v1", api_key="YOUR_KEY")
response = client.responses.create(
    model="YOUR_MODEL", input="Your task", background=True,
    extra_headers={"Idempotency-Key": "unique-operation-key"},
)
response = client.responses.retrieve(response.id)
```

To stream and resume, create with `background=True, stream=True`. Save the Response
ID from `response.created` and the latest processed `sequence_number`. Reconnect
using `GET /v1/responses/{id}?stream=true&starting_after=N`, or `Last-Event-ID: N`.
SSE IDs equal event sequence numbers. Deduplicate events by sequence number when
reconnecting. Replay retrieves the original generation; it does not regenerate it.

`POST /v1/responses/{id}/cancel` cancels queued work immediately or requests
cancellation of running work. A running request may finish before cancellation is
observed. Cancellation does not guarantee the upstream vendor stopped billing.

Responses retrieval and replay require the creating API key and organization.
`store` defaults to true, with the project's configured content retention.
Foreground `store: false` prevents retained response content. Background
`store: false` retains recovery data for ten minutes. Expired retained results
return 410. Stream events and response snapshots are encrypted at rest.

Workers renew an ownership lease every 15 seconds. The lease expires after 60
seconds and maintenance checks every 15 seconds. A lost worker produces a failed
result with committed partial output; an uncertain dispatch is never retried
automatically. This bound assumes a healthy database and maintenance worker.
Gateway-owned background work survives subscriber or API replica loss. An
upstream stream failure cannot be turned into a completed answer that the upstream
never made available.

Foreground Chat streams also retain accumulated output for operation lookup at
`GET /v1/requests/{X-The-Code-Request-ID}`. Failed lookups include available
`error.partial_response`. Foreground connections do not provide durable execution;
use background Responses for that guarantee.

## Compatibility and limits

- Reuse an idempotency key only for the identical request. Different bodies return
  409. Retrying an uncertain operation returns its recorded failure rather than
  creating another generation.
- Responses streaming supports function tools. Native hosted tools can use the
  supported non-streaming provider route. `previous_response_id` is not supported;
  send explicit conversation input and retain required tool metadata.
- Default stream no-progress timeout is 300 seconds, and total stream execution
  timeout is 1,800 seconds. Deployment settings can adjust both. SSE heartbeats do
  not count as model progress. Upstream HTTP read limits apply independently.
- SSE events are bounded to 8 MiB and accumulated wire output to 32 MiB. Oversized,
  malformed, or unterminated provider events fail explicitly.
- Project background queues default to 256 waiting requests. Overload returns 429
  with `Retry-After`. Provider HTTP connection pools are isolated by provider;
  saturation of one provider's pool does not consume another provider's slots.
- Each API/worker replica admits at most 64 concurrent requests per provider
  credential. Five consecutive upstream availability failures open that
  credential's circuit for 30 seconds, followed by one recovery probe. These
  deployment settings are configurable. Existing project/key limits still apply.
- `X-Request-ID` identifies an HTTP request (legacy caller-supplied values remain
  accepted). `X-Client-Request-ID` supports caller correlation.
  `X-The-Code-Request-ID` identifies the logical operation. These are CORS-exposed.

## Deployment diagnostics

`GET /v1/meta` returns `build_commit`, `build_time`, API version, and contract profile.
Image builds embed their immutable commit. Development builds may say `unknown`.
Use `/health/live`, `/health/startup`, and `/health/ready` for their respective probes.

This release requires the durable-response database migration and running workers.
After deployment, refresh/certify provider models so live contracts reflect current
adapter behavior. Validate the expected commit and a background create/retrieve/
resume/cancel smoke test through the public edge before promoting the release.
