Skip to navigation

Stream Speech From Text Input

Send text as your application produces it, for example a few tokens at a time from an LLM, and receive speech as each sentence completes.

Authenticate the upgrade with Authorization: Bearer and pass the synthesis settings as query parameters; text never rides the URL. A refusal (a setting the route does not accept, an unknown voice, an exhausted balance, a full rate or concurrency limit) is an ordinary HTTP error with the standard error envelope, before the upgrade. A request that is not a WebSocket upgrade answers 426 with Sec-WebSocket-Version: 13, the protocol version the channel speaks.

Send input.text as text arrives, input.flush to have everything sent so far spoken now, and input.close when you are done. The server speaks a sentence once the next one starts, and every sentence that completes while one is speaking joins the next, so a fast source produces fewer, longer segments. Each segment is one synthesis: billed, rate limited and moderated as one POST /v1/audio/stream request. The voice’s prosody starts afresh at each segment.

You receive speech.chunk events carrying Base64 audio (and word marks when speech_marks is true), speech.flushed after the audio for each flush, and exactly one terminal event: speech.done with the reason the session ended, or speech.error carrying the error envelope and the handshake’s request id. Marks index the session’s whole text on one timeline. Ignore event types you do not recognize.

A session holds one concurrent request for its life. It ends after 60 seconds with nothing to speak and no message (idle), and after 5 minutes it stops taking text and speaks what it was sent (session_limit); open a new session to continue. A server restart ends it with shutdown. The socket then closes with code 1000, or 1001 for a shutdown; a speech.error closes it with 1008 for a refusal, 1003 for a binary frame, 1009 for a message or text over the limits, and 1011 for a server or upstream fault.

Handshake

WSS
wss://api.speechify.ai/v1/audio/stream/ws

Authentication

AuthorizationBearer

Enter your API key with the Bearer prefix, e.g. 'Bearer sk_...'.

Headers

Speechify-VersionstringOptional

Query parameters

voice_idstringRequired

Id of the voice to speak with. Refer to GET /v1/voices for the available voices.

modelenumOptionalDefaults to simba-3.0

Model used for every segment, as for POST /v1/audio/stream. simba-3.2 is English only.

Allowed values:
output_formatenumOptional

Audio format of every speech.chunk, as on POST /v1/audio/stream. Defaults to MP3 at 24 kHz.

languagestringOptional

Language of the input, as an ISO 639-1 language code and an ISO 3166-1 region code separated by a hyphen, e.g. en-US.

speech_marksenumOptionalDefaults to false

When true, each speech.chunk carries the word marks that became final with it, indexing the session's whole text.

Allowed values:
safety_identifierstringOptional

A stable identifier for your end user, as on POST /v1/audio/stream.

Send

input.textobjectRequired
OR
input.flushobjectRequired
OR
input.closeobjectRequired

Receive

speech.chunkobjectRequired
OR
speech.flushedobjectRequired
OR
speech.doneobjectRequired
OR
speech.errorobjectRequired