# Audio transcription

Use `POST https://api.thecodeapi.com/v1/audio/transcriptions` with a server-side The-Code API key. Send multipart form data containing a binary `file` and `model`. Do not base64-encode the file or set the multipart boundary yourself.

```bash
curl "https://api.thecodeapi.com/v1/audio/transcriptions" \
  -H "Authorization: Bearer $THE_CODE_API_KEY" \
  -F "file=@assessment.wav" \
  -F "model=openai/gpt-4o-transcribe" \
  -F "language=en"
```

```python
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["THE_CODE_API_KEY"], base_url="https://api.thecodeapi.com/v1")
with open("assessment.wav", "rb") as audio:
    transcript = client.audio.transcriptions.create(
        model="openai/gpt-4o-transcribe", file=audio, language="en"
    )
# Save transcript.text in your application's assessment record.
```

The example model must be enabled for the caller. Discover models using authenticated `GET /v1/models`; choose `the_code.request_contract.endpoints.audio_transcription=true`. Bare IDs such as `gpt-4o-transcribe` and prefixed IDs such as `openai/gpt-4o-transcribe` are accepted. The gateway strips the provider prefix before dispatch.

## Formats and fields

| Model family | File output formats | Billing usage |
| --- | --- | --- |
| `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, supported mini snapshot | `json` | Reported text/audio input and text output tokens |
| `whisper-1` | `json`, `text`, `verbose_json`, `srt`, `vtt` | Reported audio seconds |
| `gpt-4o-transcribe-diarize` | `json`, `text`, `diarized_json` | Provider-reported usage |
| `gpt-transcribe` | `json` | Reported audio seconds |

The returned model contract, including `transcription.response_formats`, is authoritative. Model discovery and provider access can change. `gpt-live-transcribe` and `gpt-realtime-whisper` are live WebSocket models; they do not accept file transcription requests.

Files must be nonempty and at most 25,000,000 bytes. Provider-supported formats include MP3, MP4, MPEG, MPGA, M4A, WAV and WebM. The multipart request limit is at least 26 MiB, including form overhead. Model-specific optional fields are forwarded: `language`, `prompt`, `temperature`, `response_format`, `stream`, `timestamp_granularities[]`, `include[]`, `chunking_strategy`, `known_speaker_names[]`, `known_speaker_references[]`, `languages[]` and `keywords[]`. Unknown fields and duplicate scalar fields are rejected.

Use ISO language hints. The newer `gpt-transcribe` uses `languages[]` and supports keyword hints. Legacy transcription uses `language`. Diarization supports speaker annotations with `diarized_json`; use `chunking_strategy=auto` for recordings longer than 30 seconds. It does not support prompts, logprobs or timestamp-granularity controls. Optional fields remain subject to OpenAI's model-specific validation.

With `stream=true`, supported models return SSE transcript events, including `transcript.text.delta`, `transcript.text.segment` where applicable, and `transcript.text.done`. Preserve the events and wait for completion. Whisper ignores streaming and returns a normal response. Streaming requires a structured response format. Text and subtitle outputs are rendered from one structured upstream transcription so usage can be billed without a second paid request.

## Accounting and retention

The gateway reserves 100 credits by default; set a positive `X-The-Code-Max-Credits` header to choose a different reservation. Key, project, account and provider limits apply. Unused credits are returned at completion. Charges use the transcription model's approved price and the provider's reported usage, never transcript length or a guessed audio duration. A final charge above the reservation, or missing final usage, requires reconciliation.

Read `X-The-Code-Request-ID`, then call `/v1/requests/{request_id}/cost` with the same key. `line_items` contains the transcription model, measured quantities and accrued credits. `charged_credits` is the actual settled total and `cost_status` distinguishes final from uncertain costs. A failed provider request with no reported usage is not charged as a successful transcription.

Audio and transcript contents are not retained by the gateway. Store the transcript in your own application. These requests cannot be replayed by idempotency key. After a timeout or interrupted SSE stream, inspect request costs before deciding to retry.

For transcription during a spoken conversation or a transcription-only live session, read the [Realtime guide](https://thecodeapi.com/docs/realtime).

## MCP discovery

Use `list_models(endpoint="/v1/audio/transcriptions")`, `get_endpoint_schema(endpoint="/v1/audio/transcriptions")`, and `the-code://docs/audio`. `validate_request` accepts synthetic file metadata such as `{"model":"openai/gpt-4o-transcribe","file":{"filename":"assessment.wav","size_bytes":240000}}`; it checks metadata and model capabilities without uploading audio, opening a session or spending credits. Account MCP additionally checks routing and approved pricing. The actual API takes a binary multipart file, not this metadata object.
