Create Speech
Synthesize speech audio from text or SSML. Returns the complete audio
file plus billing and speech-mark metadata in a single JSON response.
For low-latency playback or long-form text, use POST /v1/audio/stream.
Set output_format for explicit sample-rate/bitrate control (e.g.
pcm_16000 or ulaw_8000 for telephony).
Authentication
Enter your API key with the Bearer prefix, e.g. 'Bearer sk_...'.
Headers
Request
Plain text or SSML to be synthesized to speech. Refer to https://docs.speechify.ai/docs/api-limits for the input size limits. Emotion, Pitch and Speed Rate are configured in the ssml input, please refer to the ssml documentation for more information: https://docs.speechify.ai/docs/ssml#prosody
Id of the voice to be used for synthesizing speech. Refer to /v1/voices endpoint for available voices
Language of the input. Follow the format of an ISO 639-1 language code and an ISO 3166-1 region code, separated by a hyphen, e.g. en-US. Please refer to the list of the supported languages and recommendations regarding this parameter: https://docs.speechify.ai/docs/language-support.
Model used for audio synthesis. Defaults to simba-3.0, which is streaming-native and multilingual: it officially supports English plus de-DE, es-ES, es-MX, fr-FR, it-IT and pt-BR, and routes each request to its English or its multilingual training based on language (falling back to the voice's locale when language is omitted). simba-3.2 is the streaming-native model with the lowest TTFB and richest expressivity, and the recommended Simba 3 model; it is English only, so a non-English voice returns 400.
The legacy Simba 1.6 models simba-english and simba-multilingual are retired from API version 2026-09-21: naming one returns 400 model_retired. Pinning your API version to a date before 2026-09-21 keeps them on their Simba 1.6 training until 2026-11-21; from then both ids are served by our current models on every API version that can still name them. Migrate to simba-3.2 (English) or simba-3.0 before then; call GET /v1/audio/models to see the set your workspace can select today.
The output audio format as a codec_sampleRate_bitrate string. Takes precedence over audio_format when set.
Response headers
Unique identifier for this request, present on every response (2xx and
non-2xx alike). If the caller sends a Speechify-Request-Id request
header the server echoes it back (sanitized and length-capped) so one
logical request can be traced end-to-end; otherwise the server generates
a fresh value. Log it on every response and quote it in support requests
- it is the stable handle that ties your observation to Speechify's
server-side logs, and it matches the
request_idfield in the error envelope.
The legacy alias X-Request-ID carries the same value and is still
accepted on requests, until 2027-07-24. Prefer the un-prefixed name
(RFC 6648).
Request-rate budget: the maximum number of requests in the current
window (the bucket capacity). The IETF-draft un-prefixed name; the
legacy alias X-RateLimit-Limit carries the same value. Rides every
response.
Request-rate budget: requests left in the current window. Legacy
alias: X-RateLimit-Remaining.
Request-rate budget: integer delta-seconds until the window fully
refills (same unit as Retry-After). Legacy alias:
X-RateLimit-Reset.
How much of the wait for this response was Speechify's, as W3C Server
Timing metrics in milliseconds. ttfb runs from the request reaching
the API to the first response byte, which on a streaming response is the
first audio. model is the model's own time to first audio (the whole
synthesis on a batch response) and is omitted when it was not measured.
Whatever your client measures beyond ttfb is network distance and your
own stack. Readable from browsers.
Response
Synthesized speech audio, Base64-encoded
The full codec_sampleRate_bitrate format the audio was encoded in, returned when the request set output_format. It is the requested value unless the request named a bitrate above the mp3 ceiling, in which case it reports the bitrate actually delivered.