> This page is for Build.

> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# API: stream text into speech over a WebSocket

`GET /v1/audio/stream/ws` is a new WebSocket endpoint for text that arrives in pieces, such as a reply an LLM streams token by token.
Send the text as you receive it, and the API speaks each sentence once it is complete, returning Base64 audio and, with `speech_marks=true`, word timestamps on one timeline.

- **Authenticate the upgrade with `Authorization: Bearer`.**
  The settings ride the query string under the `POST /v1/audio/stream` names, and a refused setting, voice or plan limit is an ordinary HTTP error before the upgrade.
- **Each segment is one synthesis.**
  It is billed, rate limited and moderated as one `POST /v1/audio/stream` request, and the voice's prosody starts afresh at each segment.
- **A session holds one concurrent request.**
  It ends after 60 seconds with nothing to speak, and stops taking text after 5 minutes.
- **Server-side only for now.**
  A browser cannot set the `Authorization` header on a WebSocket.

No API version change is needed.
See [Streaming Text Input](/build/guides/text-to-speech/streaming-text-input).