Skip to main content
POST
Create transcription

Authorizations

Authorization
string
header
required

API key as bearer token in Authorization header

Body

Speech-to-text request input. Accepts a JSON body with input_audio containing base64-encoded audio or a URL the provider downloads.

input_audio
object
required

Audio to transcribe: inline base64 bytes, or a URL the provider downloads directly.

Example:
model
string
required

STT model identifier

Example:

"openai/whisper-large-v3"

diarize
boolean

Label each word with the speaker who said it. Speaker labels are returned on the words array (speaker, speaker_label), so response_format must be "verbose_json" (a "json" request is rejected with a 400) and word timestamps are included even when timestamp_granularities omits "word". Only supported by some providers; the request is rejected with a 400 when the selected model cannot diarize. Providers may charge extra.

Example:

true

keyterms
string[]

Domain terms, names, or phrases to bias recognition toward. Only supported by some providers; the request is rejected with a 400 when the selected model cannot use keyterms. Providers may cap the number of terms or characters per term and may charge extra.

Maximum array length: 1000
Required string length: 1 - 100
Example:
language
string

ISO-639-1 language code (e.g., "en", "ja"). Auto-detected if omitted.

Example:

"en"

provider
object

Provider-specific passthrough configuration

response_format
enum<string>

Output format. "json" (default) returns { text, usage }. "verbose_json" additionally returns task, language, duration, and segment-level timestamps; only supported by OpenAI-compatible providers.

Available options:
json,
verbose_json
Example:

"json"

session_id
string

A unique identifier for grouping related requests (e.g., a conversation or agent workflow). Used for observability grouping in Broadcast and private logging; never sent to the provider. If provided in both the request body and the x-session-id header, the body value takes precedence. Maximum of 256 characters.

Maximum string length: 256
Example:

"session-1234"

temperature
number<double>

Sampling temperature for transcription

Example:

0

timestamp_granularities
enum<string>[]

Timestamp detail levels to include when response_format is "verbose_json". "segment" returns segment-level timestamps; "word" additionally returns word-level timestamps in the words array. Ignored unless response_format is "verbose_json".

A timestamp detail level for verbose_json transcription responses.

Available options:
word,
segment
Example:
trace
object

Metadata for observability and tracing. Known keys (trace_id, trace_name, span_name, generation_name, parent_span_id) have special handling. Additional keys are passed through as custom metadata to configured broadcast destinations.

Example:
user
string

A unique identifier representing your end-user. Forwarded to Broadcast and private logging as the end-user id; never sent to the provider.

Maximum string length: 256
Example:

"user-1234"

Response

Transcription result

STT response containing transcribed text and optional usage statistics

text
string
required

The transcribed text

Example:

"Hello, this is a test of OpenAI speech-to-text transcription. The weather is sunny today and the temperature is around 72 degrees."

confidence
number<double>

Provider confidence for the whole transcript from 0 to 1, present when response_format is verbose_json and the provider scores the full transcript

Example:

0.94

duration
number<double>

Duration of the input audio in seconds, present when response_format is verbose_json

Example:

9.2

entities
object[]

Detected entities with character offsets into text, present when the provider runs entity detection

language
string

Detected or forced language, present when response_format is verbose_json

Example:

"english"

language_confidence
number<double>

Provider confidence in the detected language from 0 to 1, present when response_format is verbose_json and the provider scores language detection

Example:

0.98

segments
object[]

Timestamped transcript segments, present when response_format is verbose_json

task
string

The task performed, present when response_format is verbose_json

Example:

"transcribe"

usage
object

Aggregated usage statistics for the request

Example:
words
object[]

Timestamped words, present when the provider returns word-level timestamps