> This page is for Build.

> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Streaming Text Input

> Stream text into SpeechifyAI Build TTS over a WebSocket at GET /v1/audio/stream/ws: send LLM tokens as they arrive and receive Base64 audio and word timestamps as each sentence completes.

## Overview

`GET /v1/audio/stream/ws` opens a WebSocket session for text that arrives in pieces, such as a reply an LLM streams a few tokens at a time.
You send the text as you receive it, and the API speaks each sentence once it is complete, so you do not have to split the text into requests yourself.

|                 | `POST /v1/audio/stream`                 | `GET /v1/audio/stream/ws`                  |
| --------------- | --------------------------------------- | ------------------------------------------ |
| Text            | All of it, in one request               | In pieces, as your application produces it |
| Audio           | Raw bytes over HTTP                     | Base64 in JSON events over a WebSocket     |
| Word timestamps | `POST /v1/audio/stream/with-timestamps` | `speech_marks=true`                        |
| Billing         | One request                             | One request per segment the session speaks |

If you already have the whole text, use [`POST /v1/audio/stream`](/build/guides/text-to-speech/streaming): its first audio is just as fast.

> **Warning**
>
> The session speaks your text one segment at a time, usually a sentence or a few, and each segment is a separate synthesis.
> The voice's prosody (its intonation and pacing) starts afresh at each segment, so a pause or a change of pitch can fall between two sentences that a single request would have joined smoothly.

## Connect

Open a WebSocket to `wss://api.speechify.ai/v1/audio/stream/ws` with your API key in the `Authorization` header, and pass the synthesis settings as query parameters.
The settings apply to the whole session.

| Query parameter     | Required | Description                                                                                                                        |
| ------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| `voice_id`          | Yes      | The voice to speak with. See [`GET /v1/voices`](/build/api-reference/v1/voices/get).                                               |
| `model`             | No       | `simba-3.0` (the default) or `simba-3.2`, which is English only.                                                                   |
| `output_format`     | No       | Any [streaming format](/build/guides/concepts/audio-formats), such as `pcm_24000` or `mp3_24000_64`. The default is MP3 at 24 kHz. |
| `language`          | No       | The language of the text, such as `en-US`.                                                                                         |
| `speech_marks`      | No       | `true` to receive word timestamps with the audio. The default is `false`.                                                          |
| `safety_identifier` | No       | A stable identifier for your end user, as on `POST /v1/audio/stream`.                                                              |

> **Note**
>
> Text never goes in the URL, and neither does your API key.
> A browser cannot set the `Authorization` header on a WebSocket, so connect from your server, where your LLM runs.

A setting the API does not accept, an unknown voice, an exhausted balance or a full rate or concurrency limit is refused before the upgrade, as an ordinary HTTP response with the [standard error envelope](/build/guides/concepts/error-handling).
Every response, the upgrade included, carries a `Speechify-Request-Id` header.

## Messages

Every message, in both directions, is one JSON text frame with a `type` field.

You send:

| Message                                 | Meaning                                                                                                |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------ |
| `{"type": "input.text", "text": "..."}` | Text to add to what the session will speak. Send it as you receive it; a few tokens at a time is fine. |
| `{"type": "input.flush"}`               | Speak everything sent so far now, without waiting for the sentence to end.                             |
| `{"type": "input.close"}`               | Speak everything sent so far, then end the session.                                                    |

The API sends:

| Event            | Meaning                                                                                                                              |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| `speech.chunk`   | Base64 `audio` in the session's format, the `speech_marks` that became final with it, or both.                                       |
| `speech.flushed` | All the audio for the text sent before an `input.flush` has been sent. One arrives for each flush, in order.                         |
| `speech.done`    | The session ended. `reason` says why, and `billable_characters_count` and `audio_duration_ms` cover the whole session.               |
| `speech.error`   | The session ended on an error. `error` carries the error envelope's `code` and `message`, and `request_id` the upgrade's request id. |

Exactly one of `speech.done` or `speech.error` ends every session, and nothing follows it.
There is no `[DONE]` sentinel.
Ignore event types you do not recognize: new ones may be added.

```json
{"type":"speech.chunk","audio":"//uQxAAAAAAAAAAAAAAA...","speech_marks":[{"type":"word","value":"Hello","start":0,"end":5,"start_time":125,"end_time":410}]}
{"type":"speech.flushed"}
{"type":"speech.done","billable_characters_count":63,"audio_duration_ms":4210,"reason":"closed"}
```

## How text becomes speech

* **A sentence is spoken once the next one starts.**
  A sentence ends at `.`, `!`, `?` or `…` followed by whitespace, or at a line break, so keep the spaces and newlines your source produces.
  The space after the last word of a reply never arrives, so send `input.flush` or `input.close` to have it spoken.
* **Sentences that complete while one is speaking join the next segment.**
  A fast source produces fewer, longer segments, which sound smoother.
* **A sentence that runs past 1,000 characters without ending** is cut at its last word boundary, so unpunctuated text is still spoken.
* **Text is plain text.**
  SSML is read aloud, not interpreted.
* **Segments are spoken one at a time, in order.**
  Concatenate the audio of every `speech.chunk` in the order it arrives.
  With `pcm_*` or `ulaw_8000` the result is one continuous stream.
  With an MP3, Ogg or AAC format each segment starts a new stream in that format, so play the result with a decoder that accepts concatenated streams, or choose PCM for real-time playback.
* **Speech marks share one timeline.**
  `start` and `end` index all the text the session received, counted from the first character of the first `input.text`, and `start_time` and `end_time` are milliseconds from the start of the session's audio.
  See [Speech marks](/build/guides/text-to-speech/speech-marks) for the fields.

## Limits and billing

* **Each segment is billed as one `POST /v1/audio/stream` request** for the characters it speaks, and moderated and logged as one.
  Text you sent but the session never spoke is not billed.
* **A session uses one of your plan's [concurrent requests](/build/guides/concepts/api-limits#concurrency-limits)** for as long as it is open, whether or not it is speaking.
* **Each segment after the first uses one request from your rate limit.**
  When the limit is reached the session waits instead of refusing, and the sentences that arrive meanwhile join the waiting segment.
* **At most 20,000 characters can wait to be spoken at once.**
  More ends the session with `payload_too_large`.
* **A session ends after 60 seconds with nothing to speak and no message from you**, with `reason` `idle`.
* **A session stops taking text after 5 minutes** and speaks what it was sent, then ends with `reason` `session_limit`.
  Open a new session to continue.
* **A server restart ends the session** with `reason` `shutdown`.
  Reconnect to continue; a session cannot be resumed.

## How a session ends

| Event          | `reason` or `code`                                                                                          | Close code |
| -------------- | ----------------------------------------------------------------------------------------------------------- | ---------- |
| `speech.done`  | `closed`: you sent `input.close` and everything was spoken                                                  | 1000       |
| `speech.done`  | `idle` or `session_limit`                                                                                   | 1000       |
| `speech.done`  | `shutdown`                                                                                                  | 1001       |
| `speech.error` | A refusal, such as `payment_required`, `content_policy_violation`, or `bad_request` for a malformed message | 1008       |
| `speech.error` | A binary frame                                                                                              | 1003       |
| `speech.error` | `payload_too_large`: a message over 256 KiB, or more than 20,000 characters waiting                         | 1009       |
| `speech.error` | `upstream_failure` or `internal_error`                                                                      | 1011       |

A segment that fails partway through is billed for what it delivered.
If you close the socket while a segment is speaking, that segment is billed in full.

## Examples

Each example sends a reply in pieces, as an LLM would, then closes the session and saves the audio to `speech.pcm`.
Play it with `ffplay -f s16le -ar 24000 -ac 1 speech.pcm`.

#### Python

```python
# pip install websockets
import asyncio
import base64
import json
import os
from urllib.parse import urlencode

from websockets.asyncio.client import connect

URL = "wss://api.speechify.ai/v1/audio/stream/ws?" + urlencode(
    {"voice_id": "geffen_32", "model": "simba-3.2", "output_format": "pcm_24000", "speech_marks": "true"}
)


async def llm_reply():
    # Stands in for your LLM's streamed reply.
    for token in ["Your", " order", " shipped", " this", " morning", ". It", " arrives", " on", " Thursday", "."]:
        yield token
        await asyncio.sleep(0.05)


async def send_reply(ws):
    async for token in llm_reply():
        await ws.send(json.dumps({"type": "input.text", "text": token}))
    await ws.send(json.dumps({"type": "input.close"}))


async def main():
    headers = {"Authorization": f"Bearer {os.environ['SPEECHIFY_API_KEY']}"}
    async with connect(URL, additional_headers=headers) as ws:
        sender = asyncio.create_task(send_reply(ws))
        with open("speech.pcm", "wb") as audio:
            async for frame in ws:
                event = json.loads(frame)
                if event["type"] == "speech.chunk":
                    if "audio" in event:
                        audio.write(base64.b64decode(event["audio"]))
                    for mark in event.get("speech_marks", []):
                        print(f"{mark['start_time']:>6} ms  {mark['value']}")
                elif event["type"] == "speech.done":
                    print(f"done ({event['reason']}): {event['billable_characters_count']} characters")
                    break
                elif event["type"] == "speech.error":
                    error = event["error"]
                    print(f"error {error['code']}: {error['message']} (request {event['request_id']})")
                    break
        sender.cancel()


asyncio.run(main())
```

#### TypeScript

```typescript
// npm install ws
import { createWriteStream } from "node:fs";
import WebSocket from "ws";

const params = new URLSearchParams({
  voice_id: "geffen_32",
  model: "simba-3.2",
  output_format: "pcm_24000",
  speech_marks: "true",
});
const ws = new WebSocket(`wss://api.speechify.ai/v1/audio/stream/ws?${params}`, {
  headers: { Authorization: `Bearer ${process.env.SPEECHIFY_API_KEY}` },
});
const audio = createWriteStream("speech.pcm");

// Stands in for your LLM's streamed reply.
async function* llmReply() {
  for (const token of ["Your", " order", " shipped", " this", " morning", ". It", " arrives", " on", " Thursday", "."]) {
    yield token;
    await new Promise((resolve) => setTimeout(resolve, 50));
  }
}

ws.on("open", async () => {
  for await (const token of llmReply()) {
    ws.send(JSON.stringify({ type: "input.text", text: token }));
  }
  ws.send(JSON.stringify({ type: "input.close" }));
});

ws.on("message", (data) => {
  const event = JSON.parse(data.toString());
  switch (event.type) {
    case "speech.chunk":
      if (event.audio) audio.write(Buffer.from(event.audio, "base64"));
      for (const mark of event.speech_marks ?? []) console.log(`${mark.start_time} ms  ${mark.value}`);
      break;
    case "speech.done":
      console.log(`done (${event.reason}): ${event.billable_characters_count} characters`);
      break;
    case "speech.error":
      console.error(`error ${event.error.code}: ${event.error.message} (request ${event.request_id})`);
      break;
  }
});

// A refused upgrade is an HTTP response carrying the error envelope.
ws.on("unexpected-response", (request, response) => {
  let body = "";
  response.on("data", (chunk) => (body += chunk));
  response.on("end", () => {
    console.error(`HTTP ${response.statusCode}: ${body}`);
    request.destroy();
  });
});

ws.on("close", () => audio.end());
```

#### websocat

```bash
printf '%s\n' \
  '{"type":"input.text","text":"Your order shipped this morning. "}' \
  '{"type":"input.text","text":"It arrives on Thursday."}' \
  '{"type":"input.close"}' |
websocat --text --no-close \
  "wss://api.speechify.ai/v1/audio/stream/ws?voice_id=geffen_32&model=simba-3.2&output_format=pcm_24000" \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" |
tee events.jsonl |
jq -r 'select(.type == "speech.chunk" and .audio) | .audio' |
while read -r chunk; do printf '%s' "$chunk" | base64 -d; done > speech.pcm

jq -c 'select(.type != "speech.chunk")' events.jsonl
```

`--no-close` keeps the socket open after the last message, until the session ends.
Without it, websocat closes the socket as soon as it has sent the text, and the session ends before it has spoken.