Stream Speech From Text Input
Send text as your application produces it, for example a few tokens at a time from an LLM, and receive speech as each sentence completes.
Authenticate the upgrade with Authorization: Bearer and pass the
synthesis settings as query parameters; text never rides the URL. A
refusal (a setting the route does not accept, an unknown voice, an
exhausted balance, a full rate or concurrency limit) is an ordinary HTTP
error with the standard error envelope, before the upgrade. A request
that is not a WebSocket upgrade answers 426 with
Sec-WebSocket-Version: 13, the protocol version the channel speaks.
Send input.text as text arrives, input.flush to have everything sent
so far spoken now, and input.close when you are done. The server speaks
a sentence once the next one starts, and every sentence that completes
while one is speaking joins the next, so a fast source produces fewer,
longer segments. Each segment is one synthesis: billed, rate limited and
moderated as one POST /v1/audio/stream request. The voice’s prosody
starts afresh at each segment.
You receive speech.chunk events carrying Base64 audio (and word marks
when speech_marks is true), speech.flushed after the audio for each
flush, and exactly one terminal event: speech.done with the reason the
session ended, or speech.error carrying the error envelope and the
handshake’s request id. Marks index the session’s whole text on one
timeline. Ignore event types you do not recognize.
A session holds one concurrent request for its life. It ends after 60
seconds with nothing to speak and no message (idle), and after 5
minutes it stops taking text and speaks what it was sent
(session_limit); open a new session to continue. A server restart ends
it with shutdown. The socket then closes with code 1000, or 1001 for a
shutdown; a speech.error closes it with 1008 for a refusal,
1003 for a binary frame, 1009 for a message or text over the limits, and
1011 for a server or upstream fault.
Handshake
Authentication
Enter your API key with the Bearer prefix, e.g. 'Bearer sk_...'.
Headers
Query parameters
Id of the voice to speak with. Refer to GET /v1/voices for the available voices.
Model used for every segment, as for POST /v1/audio/stream. simba-3.2 is English only.
Audio format of every speech.chunk, as on POST /v1/audio/stream. Defaults to MP3 at 24 kHz.
Language of the input, as an ISO 639-1 language code and an ISO 3166-1 region code separated by a hyphen, e.g. en-US.
When true, each speech.chunk carries the word marks that became final with it, indexing the session's whole text.
A stable identifier for your end user, as on POST /v1/audio/stream.