Skip to navigation

Streaming Text Input

Send text over a WebSocket as your application produces it and receive speech sentence by sentence

Overview

GET /v1/audio/stream/ws opens a WebSocket session for text that arrives in pieces, such as a reply an LLM streams a few tokens at a time. You send the text as you receive it, and the API speaks each sentence once it is complete, so you do not have to split the text into requests yourself.

POST /v1/audio/streamGET /v1/audio/stream/ws
TextAll of it, in one requestIn pieces, as your application produces it
AudioRaw bytes over HTTPBase64 in JSON events over a WebSocket
Word timestampsPOST /v1/audio/stream/with-timestampsspeech_marks=true
BillingOne requestOne request per segment the session speaks

If you already have the whole text, use POST /v1/audio/stream: its first audio is just as fast.

The session speaks your text one segment at a time, usually a sentence or a few, and each segment is a separate synthesis. The voice’s prosody (its intonation and pacing) starts afresh at each segment, so a pause or a change of pitch can fall between two sentences that a single request would have joined smoothly.

Connect

Open a WebSocket to wss://api.speechify.ai/v1/audio/stream/ws with your API key in the Authorization header, and pass the synthesis settings as query parameters. The settings apply to the whole session.

Query parameterRequiredDescription
voice_idYesThe voice to speak with. See GET /v1/voices.
modelNosimba-3.0 (the default) or simba-3.2, which is English only.
output_formatNoAny streaming format, such as pcm_24000 or mp3_24000_64. The default is MP3 at 24 kHz.
languageNoThe language of the text, such as en-US.
speech_marksNotrue to receive word timestamps with the audio. The default is false.
safety_identifierNoA stable identifier for your end user, as on POST /v1/audio/stream.

Text never goes in the URL, and neither does your API key. A browser cannot set the Authorization header on a WebSocket, so connect from your server, where your LLM runs.

A setting the API does not accept, an unknown voice, an exhausted balance or a full rate or concurrency limit is refused before the upgrade, as an ordinary HTTP response with the standard error envelope. Every response, the upgrade included, carries a Speechify-Request-Id header.

Messages

Every message, in both directions, is one JSON text frame with a type field.

You send:

MessageMeaning
{"type": "input.text", "text": "..."}Text to add to what the session will speak. Send it as you receive it; a few tokens at a time is fine.
{"type": "input.flush"}Speak everything sent so far now, without waiting for the sentence to end.
{"type": "input.close"}Speak everything sent so far, then end the session.

The API sends:

EventMeaning
speech.chunkBase64 audio in the session’s format, the speech_marks that became final with it, or both.
speech.flushedAll the audio for the text sent before an input.flush has been sent. One arrives for each flush, in order.
speech.doneThe session ended. reason says why, and billable_characters_count and audio_duration_ms cover the whole session.
speech.errorThe session ended on an error. error carries the error envelope’s code and message, and request_id the upgrade’s request id.

Exactly one of speech.done or speech.error ends every session, and nothing follows it. There is no [DONE] sentinel. Ignore event types you do not recognize: new ones may be added.

{"type":"speech.chunk","audio":"//uQxAAAAAAAAAAAAAAA...","speech_marks":[{"type":"word","value":"Hello","start":0,"end":5,"start_time":125,"end_time":410}]}
{"type":"speech.flushed"}
{"type":"speech.done","billable_characters_count":63,"audio_duration_ms":4210,"reason":"closed"}

How text becomes speech

  • A sentence is spoken once the next one starts. A sentence ends at ., !, ? or … followed by whitespace, or at a line break, so keep the spaces and newlines your source produces. The space after the last word of a reply never arrives, so send input.flush or input.close to have it spoken.
  • Sentences that complete while one is speaking join the next segment. A fast source produces fewer, longer segments, which sound smoother.
  • A sentence that runs past 1,000 characters without ending is cut at its last word boundary, so unpunctuated text is still spoken.
  • Text is plain text. SSML is read aloud, not interpreted.
  • Segments are spoken one at a time, in order. Concatenate the audio of every speech.chunk in the order it arrives. With pcm_* or ulaw_8000 the result is one continuous stream. With an MP3, Ogg or AAC format each segment starts a new stream in that format, so play the result with a decoder that accepts concatenated streams, or choose PCM for real-time playback.
  • Speech marks share one timeline. start and end index all the text the session received, counted from the first character of the first input.text, and start_time and end_time are milliseconds from the start of the session’s audio. See Speech marks for the fields.

Limits and billing

  • Each segment is billed as one POST /v1/audio/stream request for the characters it speaks, and moderated and logged as one. Text you sent but the session never spoke is not billed.
  • A session uses one of your plan’s concurrent requests for as long as it is open, whether or not it is speaking.
  • Each segment after the first uses one request from your rate limit. When the limit is reached the session waits instead of refusing, and the sentences that arrive meanwhile join the waiting segment.
  • At most 20,000 characters can wait to be spoken at once. More ends the session with payload_too_large.
  • A session ends after 60 seconds with nothing to speak and no message from you, with reason idle.
  • A session stops taking text after 5 minutes and speaks what it was sent, then ends with reason session_limit. Open a new session to continue.
  • A server restart ends the session with reason shutdown. Reconnect to continue; a session cannot be resumed.

How a session ends

Eventreason or codeClose code
speech.doneclosed: you sent input.close and everything was spoken1000
speech.doneidle or session_limit1000
speech.doneshutdown1001
speech.errorA refusal, such as payment_required, content_policy_violation, or bad_request for a malformed message1008
speech.errorA binary frame1003
speech.errorpayload_too_large: a message over 256 KiB, or more than 20,000 characters waiting1009
speech.errorupstream_failure or internal_error1011

A segment that fails partway through is billed for what it delivered. If you close the socket while a segment is speaking, that segment is billed in full.

Examples

Each example sends a reply in pieces, as an LLM would, then closes the session and saves the audio to speech.pcm. Play it with ffplay -f s16le -ar 24000 -ac 1 speech.pcm.

# pip install websockets
import asyncio
import base64
import json
import os
from urllib.parse import urlencode
from websockets.asyncio.client import connect
URL = "wss://api.speechify.ai/v1/audio/stream/ws?" + urlencode(
{"voice_id": "geffen_32", "model": "simba-3.2", "output_format": "pcm_24000", "speech_marks": "true"}
)
async def llm_reply():
# Stands in for your LLM's streamed reply.
for token in ["Your", " order", " shipped", " this", " morning", ". It", " arrives", " on", " Thursday", "."]:
yield token
await asyncio.sleep(0.05)
async def send_reply(ws):
async for token in llm_reply():
await ws.send(json.dumps({"type": "input.text", "text": token}))
await ws.send(json.dumps({"type": "input.close"}))
async def main():
headers = {"Authorization": f"Bearer {os.environ['SPEECHIFY_API_KEY']}"}
async with connect(URL, additional_headers=headers) as ws:
sender = asyncio.create_task(send_reply(ws))
with open("speech.pcm", "wb") as audio:
async for frame in ws:
event = json.loads(frame)
if event["type"] == "speech.chunk":
if "audio" in event:
audio.write(base64.b64decode(event["audio"]))
for mark in event.get("speech_marks", []):
print(f"{mark['start_time']:>6} ms {mark['value']}")
elif event["type"] == "speech.done":
print(f"done ({event['reason']}): {event['billable_characters_count']} characters")
break
elif event["type"] == "speech.error":
error = event["error"]
print(f"error {error['code']}: {error['message']} (request {event['request_id']})")
break
sender.cancel()
asyncio.run(main())