Streaming Text Input
Overview
GET /v1/audio/stream/ws opens a WebSocket session for text that arrives in pieces, such as a reply an LLM streams a few tokens at a time.
You send the text as you receive it, and the API speaks each sentence once it is complete, so you do not have to split the text into requests yourself.
If you already have the whole text, use POST /v1/audio/stream: its first audio is just as fast.
The session speaks your text one segment at a time, usually a sentence or a few, and each segment is a separate synthesis. The voice’s prosody (its intonation and pacing) starts afresh at each segment, so a pause or a change of pitch can fall between two sentences that a single request would have joined smoothly.
Connect
Open a WebSocket to wss://api.speechify.ai/v1/audio/stream/ws with your API key in the Authorization header, and pass the synthesis settings as query parameters.
The settings apply to the whole session.
Text never goes in the URL, and neither does your API key.
A browser cannot set the Authorization header on a WebSocket, so connect from your server, where your LLM runs.
A setting the API does not accept, an unknown voice, an exhausted balance or a full rate or concurrency limit is refused before the upgrade, as an ordinary HTTP response with the standard error envelope.
Every response, the upgrade included, carries a Speechify-Request-Id header.
Messages
Every message, in both directions, is one JSON text frame with a type field.
You send:
The API sends:
Exactly one of speech.done or speech.error ends every session, and nothing follows it.
There is no [DONE] sentinel.
Ignore event types you do not recognize: new ones may be added.
How text becomes speech
- A sentence is spoken once the next one starts.
A sentence ends at
.,!,?or…followed by whitespace, or at a line break, so keep the spaces and newlines your source produces. The space after the last word of a reply never arrives, so sendinput.flushorinput.closeto have it spoken. - Sentences that complete while one is speaking join the next segment. A fast source produces fewer, longer segments, which sound smoother.
- A sentence that runs past 1,000 characters without ending is cut at its last word boundary, so unpunctuated text is still spoken.
- Text is plain text. SSML is read aloud, not interpreted.
- Segments are spoken one at a time, in order.
Concatenate the audio of every
speech.chunkin the order it arrives. Withpcm_*orulaw_8000the result is one continuous stream. With an MP3, Ogg or AAC format each segment starts a new stream in that format, so play the result with a decoder that accepts concatenated streams, or choose PCM for real-time playback. - Speech marks share one timeline.
startandendindex all the text the session received, counted from the first character of the firstinput.text, andstart_timeandend_timeare milliseconds from the start of the session’s audio. See Speech marks for the fields.
Limits and billing
- Each segment is billed as one
POST /v1/audio/streamrequest for the characters it speaks, and moderated and logged as one. Text you sent but the session never spoke is not billed. - A session uses one of your plan’s concurrent requests for as long as it is open, whether or not it is speaking.
- Each segment after the first uses one request from your rate limit. When the limit is reached the session waits instead of refusing, and the sentences that arrive meanwhile join the waiting segment.
- At most 20,000 characters can wait to be spoken at once.
More ends the session with
payload_too_large. - A session ends after 60 seconds with nothing to speak and no message from you, with
reasonidle. - A session stops taking text after 5 minutes and speaks what it was sent, then ends with
reasonsession_limit. Open a new session to continue. - A server restart ends the session with
reasonshutdown. Reconnect to continue; a session cannot be resumed.
How a session ends
A segment that fails partway through is billed for what it delivered. If you close the socket while a segment is speaking, that segment is billed in full.
Examples
Each example sends a reply in pieces, as an LLM would, then closes the session and saves the audio to speech.pcm.
Play it with ffplay -f s16le -ar 24000 -ac 1 speech.pcm.