Skip to content
THE-CODE API

Audio transcription

Use POST https://api.thecodeapi.com/v1/audio/transcriptions with a server-side The-Code API key. Send multipart form data containing a binary file and model. Do not base64-encode the file or set the multipart boundary yourself.

bash
curl "https://api.thecodeapi.com/v1/audio/transcriptions" \
  -H "Authorization: Bearer $THE_CODE_API_KEY" \
  -F "file=@assessment.wav" \
  -F "model=openai/gpt-4o-transcribe" \
  -F "language=en"
python
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["THE_CODE_API_KEY"], base_url="https://api.thecodeapi.com/v1")
with open("assessment.wav", "rb") as audio:
    transcript = client.audio.transcriptions.create(
        model="openai/gpt-4o-transcribe", file=audio, language="en"
    )
# Save transcript.text in your application's assessment record.

The example model must be enabled for the caller. Discover models using authenticated GET /v1/models; choose the_code.request_contract.endpoints.audio_transcription=true. Bare IDs such as gpt-4o-transcribe and prefixed IDs such as openai/gpt-4o-transcribe are accepted. The gateway strips the provider prefix before dispatch.

Formats and fields

Model familyFile output formatsBilling usage
gpt-4o-transcribe, gpt-4o-mini-transcribe, supported mini snapshotjsonReported text/audio input and text output tokens
whisper-1json, text, verbose_json, srt, vttReported audio seconds
gpt-4o-transcribe-diarizejson, text, diarized_jsonProvider-reported usage
gpt-transcribejsonReported audio seconds

The returned model contract, including transcription.response_formats, is authoritative. Model discovery and provider access can change. gpt-live-transcribe and gpt-realtime-whisper are live WebSocket models; they do not accept file transcription requests.

Files must be nonempty and at most 25,000,000 bytes. Provider-supported formats include MP3, MP4, MPEG, MPGA, M4A, WAV and WebM. The multipart request limit is at least 26 MiB, including form overhead. Model-specific optional fields are forwarded: language, prompt, temperature, response_format, stream, timestamp_granularities[], include[], chunking_strategy, known_speaker_names[], known_speaker_references[], languages[] and keywords[]. Unknown fields and duplicate scalar fields are rejected.

Use ISO language hints. The newer gpt-transcribe uses languages[] and supports keyword hints. Legacy transcription uses language. Diarization supports speaker annotations with diarized_json; use chunking_strategy=auto for recordings longer than 30 seconds. It does not support prompts, logprobs or timestamp-granularity controls. Optional fields remain subject to OpenAI's model-specific validation.

With stream=true, supported models return SSE transcript events, including transcript.text.delta, transcript.text.segment where applicable, and transcript.text.done. Preserve the events and wait for completion. Whisper ignores streaming and returns a normal response. Streaming requires a structured response format. Text and subtitle outputs are rendered from one structured upstream transcription so usage can be billed without a second paid request.

Accounting and retention

The gateway reserves 100 credits by default; set a positive X-The-Code-Max-Credits header to choose a different reservation. Key, project, account and provider limits apply. Unused credits are returned at completion. Charges use the transcription model's approved price and the provider's reported usage, never transcript length or a guessed audio duration. A final charge above the reservation, or missing final usage, requires reconciliation.

Read X-The-Code-Request-ID, then call /v1/requests/{request_id}/cost with the same key. line_items contains the transcription model, measured quantities and accrued credits. charged_credits is the actual settled total and cost_status distinguishes final from uncertain costs. A failed provider request with no reported usage is not charged as a successful transcription.

Audio and transcript contents are not retained by the gateway. Store the transcript in your own application. These requests cannot be replayed by idempotency key. After a timeout or interrupted SSE stream, inspect request costs before deciding to retry.

For transcription during a spoken conversation or a transcription-only live session, read the Realtime guide.

MCP discovery

Use list_models(endpoint="/v1/audio/transcriptions"), get_endpoint_schema(endpoint="/v1/audio/transcriptions"), and the-code://docs/audio. validate_request accepts synthetic file metadata such as {"model":"openai/gpt-4o-transcribe","file":{"filename":"assessment.wav","size_bytes":240000}}; it checks metadata and model capabilities without uploading audio, opening a session or spending credits. Account MCP additionally checks routing and approved pricing. The actual API takes a binary multipart file, not this metadata object.