# Realtime voice over WebSocket

Agents and application backends open a WebSocket directly. No REST session-creation request is needed.

| Connection setting | Value |
| --- | --- |
| Production example | `wss://api.the-code.ai/v1/realtime?model=openai/gpt-realtime-2.1` |
| Configured deployment | Replace `https://` in `https://api.thecodeapi.com` with `wss://`, then append `/v1/realtime?model=<encoded-model-id>` |
| Authentication header | `Authorization: Bearer <THE_CODE_API_KEY>` |
| Protocol | OpenAI Realtime GA JSON events over text WebSocket frames |
| First provider event | `session.created` |
| Optional session budget | `X-The-Code-Max-Credits: 100` |

Use a **The-Code application API key**, kept on your backend. A portal login, MCP OAuth token, or OpenAI provider key is not a gateway application key. Keys in query strings and WebSocket subprotocols are not accepted. Omit the obsolete `OpenAI-Beta: realtime=v1` header.

First call authenticated `GET https://api.thecodeapi.com/v1/models`. Select an exact model or alias whose `the_code.request_contract.endpoints.realtime` is `true`. The example model must be enabled for your key. The gateway translates the gateway model ID to the provider model and supplies the centrally managed provider credential.

## Minimal agent connection

Install the Node.js `ws` package, set `THE_CODE_API_KEY` and `THE_CODE_MODEL` in your server environment, and run:

```javascript
import WebSocket from "ws";

const url = new URL("https://api.thecodeapi.com/v1/realtime");
url.protocol = url.protocol === "https:" ? "wss:" : "ws:";
url.searchParams.set("model", process.env.THE_CODE_MODEL);
const ws = new WebSocket(url, {
  headers: { Authorization: `Bearer ${process.env.THE_CODE_API_KEY}` },
});
ws.on("message", (raw) => {
  const event = JSON.parse(raw.toString());
  if (event.type === "session.created") {
    ws.send(JSON.stringify({
      type: "session.update",
      session: { type: "realtime", instructions: "Be concise and helpful." },
    }));
    ws.send(JSON.stringify({
      type: "conversation.item.create",
      item: { type: "message", role: "user",
        content: [{ type: "input_text", text: "Say hello." }] },
    }));
    ws.send(JSON.stringify({ type: "response.create" }));
  }
  // Handle audio, transcripts, function calls, response.done and error events here.
  // Never log keys, raw audio, transcripts or complete event payloads.
});
ws.on("error", () => console.error("Realtime connection failed"));
```

For microphone input, send `input_audio_buffer.append` events containing base64 audio in the configured audio format. Use the provider's GA session audio configuration and turn detection; for manual turns, commit the buffer and send `response.create`. Audio, output transcripts, function-call arguments, cancellation and conversation events travel through the socket. See the [OpenAI WebSocket guide](https://developers.openai.com/api/docs/guides/voice-websockets) and [conversation event guide](https://developers.openai.com/api/docs/guides/realtime-conversations).

## Supported scope and limits

This endpoint supports server-to-server text/audio conversations and function tools. Browser ephemeral keys, WebRTC, separate input transcription, image input and remote MCP tools are not supported by this gateway route. Keep browser microphone handling behind an authenticated application backend; do not embed a permanent key in browser code.

`POST /v1/realtime/sessions`, `/v1/realtime/client_secrets` and `/v1/realtime/calls` return `501 unsupported_transport` with the supported WebSocket alternative. An ordinary HTTP GET returns `426 websocket_upgrade_required`. WebSocket endpoints are separate from HTTP/SSE chat streaming.

Default gateway limits are 30 minutes, 32,768 total reported tokens per session, 4,096 output tokens per response, and 1 MiB per event. A key's stricter output ceiling also applies. Admission reserves the session token allowance against the key/project token limit, and holds a concurrency slot until disconnect. Time, token and default output limits can be configured by the operator; MCP `get_connection_info` returns the deployed connection limits.

Paid sessions reserve 100 credits by default, or the positive amount supplied in `X-The-Code-Max-Credits`. The key/account must have sufficient spendable credits for this reservation. Text, audio and cached usage are accounted separately from `response.done` events; the session closes when reported token usage or credits reach its allowance. Usage arrives at response boundaries, so the last in-flight response can exceed an allowance. The customer charge is capped at the reserved credits and any shortfall or unknown final usage is marked for reconciliation. Unused reserved credits are returned on disconnect. Provider voice/text payloads are not retained by this relay.

Approved prices must include input/output audio rates as well as text rates. Incomplete pricing returns `409` before provider connection. The initial connection and every 30 seconds of a session recheck key/account access. `X-The-Code-Request-ID` on the successful upgrade identifies the session for `/v1/requests/{request_id}/cost`.

## Connection errors and diagnosis

- `401`: missing, invalid, expired or revoked application key.
- `403`: key/model/IP policy or a configured direct-origin restriction.
- `404`: model or alias unavailable to this key.
- `409`: approved realtime pricing is incomplete.
- `402` / `429`: credit, rate, token or concurrency allowance prevents admission.
- `502`: provider connection, model access or usage protocol failure.

Handshake failures return a JSON error when the ASGI server supports denial responses. After upgrade, gateway failures arrive as an `error` event followed by socket closure. Do not automatically retry a disconnected paid conversation; inspect recorded usage first. Sessions cannot be resumed by idempotency key.

Earlier deployments had no WebSocket handler: the ASGI application rejected both `/v1/realtime` and invented WebSocket paths with an empty HTTP 403. That symptom alone does not establish a WAF block. Check `/v1/meta` for `features.realtime_websocket: true`, then test the real route with a missing key (expected 401) and an authorized key (expected 101 and `session.created`). Unknown WebSocket paths may still return 403.

For connector discovery, call MCP `get_connection_info` and read its `realtime` object, `list_models` with `endpoint: "/v1/realtime"`, or `the-code://docs/realtime`. MCP documents the connection and does not open a paid voice session.
