> This page is for Build.

> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Latency

> SpeechifyAI text-to-speech latency: Simba 3.2 first byte 56 ms p50 and 102 ms p90, measured 15 Sep 2026 on the production US East streaming path. How it is measured, what adds latency, and code to measure time to first audio yourself.

On 15 Sep 2026, SpeechifyAI's `simba-3.2` model returned its first audio byte in 56 ms at the median (p50) and 102 ms at p90 on our production streaming path in US East. That is the part of the wait spent inside our API. What a listener hears also depends on the network between you and US East, how your client connects, the audio format and your player. This page shows how each of those adds up, and how to measure it from where your code runs.

## Measured numbers

| Metric          | Value      | What it measures                                                                | Measured    | Path                                        |
| --------------- | ---------- | ------------------------------------------------------------------------------- | ----------- | ------------------------------------------- |
| First byte, p50 | **56 ms**  | `simba-3.2`: from our API admitting the request to writing its first audio byte | 15 Sep 2026 | Production `POST /v1/audio/stream`, US East |
| First byte, p90 | **102 ms** | Same metric; 9 in 10 requests were at or under this                             | 15 Sep 2026 | Production `POST /v1/audio/stream`, US East |
| First byte, p50 | 161 ms     | Same metric on `simba-3.0`                                                      | 15 Sep 2026 | Production `POST /v1/audio/stream`, US East |

These are measured percentiles from production traffic, not a guarantee. They are not time to audible audio, and they do not include the network between you and US East.

## First byte, first audible audio, total generation time

| Term                        | Definition                                                                                                                         |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Time to first byte          | From sending the request to receiving the first byte of audio. The numbers above are the part of it spent inside our API.          |
| Time to first audible audio | Time to first byte, plus any silence at the start of the audio, plus the time your decoder and player take before sound comes out. |
| Total generation time       | From sending the request to receiving the last byte of audio. It grows with the length of the text.                                |

Generated speech can begin with a short silence, and its length varies with the voice and the text. On `POST /v1/audio/speech`, time to first byte and total generation time are the same, because the response is sent once the whole clip exists.

## How we measure

* **What is timed.** Every production request to `POST /v1/audio/stream` records, inside our API, the time from admitting the request (after authentication and rate limits) to writing the first audio byte to the response.
* **Which requests.** Successful streaming requests from real customer traffic, not a synthetic test, split by model. The figures above were read from that traffic on 15 Sep 2026.
* **Which statistic.** Percentiles, not averages. p50 is the median request; p90 is the request slower than 9 in 10. A few slow outliers move an average, not a median.
* **Where.** Our US East serving path, which served every request to `api.speechify.ai` in September 2026. The clock stops when the first byte leaves our API, so the network between you and US East, the handshakes of a new connection, silence at the start of the audio and your player's buffer are all outside it.
* **Cross-checked from outside.** We also time requests from the client side: a probe sends varied text to the streaming endpoint from a machine in US East and from other locations, timestamps DNS, TCP, TLS and every chunk, and finds the first audible 10 ms window in the decoded audio. The [examples below](#measure-it-yourself) do the same, without the audio analysis.

## Independent measurements

Two benchmarks we do not run also time `simba-3.2`, from their own runners to the first audible sample.
[Coval](https://benchmarks.coval.ai/tts) measured a 106 ms median time to first audio over the 24 hours to 24 Sep 2026, and [Voice Arena](https://voicearena.com/tts-leaderboard/us-english) read 123 ms p50 on its US English board the same day.
Both sit above the 56 ms first byte because they include the network between their runner and US East and any silence at the start of the audio, which our first byte does not count.
Each is that benchmark's own reading on that date, not our measurement, and the boards change as they re-run.
Our [write-up of both readings](https://speechify.ai/blog/simba-3-2-first-audio-latency-coval) links the reviewed, dated snapshot behind each figure.

## What adds latency on your side, and how to cut it

| Factor                       | Effect                                                                                                                                                                                                                    | What to do                                                                                                                                                                                                                          |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Batch instead of streaming   | `POST /v1/audio/speech` responds once the whole clip is generated, so the first audio waits for the last.                                                                                                                 | Use [`POST /v1/audio/stream`](/build/guides/text-to-speech/streaming) for anything played while it generates. Keep the speech endpoint for files you store.                                                                         |
| Model                        | Median first byte was 56 ms on `simba-3.2` and 161 ms on `simba-3.0`, same path, same day.                                                                                                                                | Set `model: "simba-3.2"` for English. The API uses `simba-3.0` when `model` is omitted. See [Models](/build/guides/concepts/models).                                                                                                |
| Distance to US East          | Every request pays at least one network round trip to US East. In our September 2026 measurements that round trip was about 35 ms from the central US and about 100 ms from central Europe.                               | If the code that calls the API runs in a cloud, run it in or near US East.                                                                                                                                                          |
| A new connection per request | A DNS lookup and the TCP and TLS handshakes run before the request is sent.                                                                                                                                               | Create one HTTP client and reuse it. Open the connection at startup so a user's first request does not pay for it. HTTP/2 lets concurrent streams share one connection.                                                             |
| Compressed output            | MP3, Ogg/Opus and AAC are encoded as the audio is generated and decoded by your player, and decoded MP3 starts with a short encoder delay. `audio/pcm` at 24 kHz is the model's native output and needs no encoding step. | Use `Accept: audio/pcm` when you play or process raw samples. Use a compressed format when bandwidth or file size matters more. See [Audio formats](/build/guides/concepts/audio-formats).                                          |
| Waiting for the full text    | Synthesis cannot start before the text arrives. An LLM reply sent in one request waits for the LLM's last token.                                                                                                          | Send each complete sentence as soon as it exists and play the streams in order. Split only at sentence ends: splitting inside a sentence changes how it sounds. Text you already have goes in one request, up to 20,000 characters. |
| Player start-up              | A player that fills a buffer before it starts adds that buffer to time to first audible audio.                                                                                                                            | Start playback on the first chunk. With PCM, write samples to the audio device as they arrive.                                                                                                                                      |

## Measure it yourself

Time the first chunk of a streaming response from where your code runs, and read the `Server-Timing` header on the same response. Its `ttfb` is our share of that request's wait, from the request reaching our API to the first audio byte, and `model` is the model's own time to first audio. Your first-chunk time minus `ttfb` is the network and your own stack. `ttfb` starts slightly earlier than the published figures, so it also counts authentication and rate limiting.

For a fair number:

* Discard the first request or two. They include opening the connection.
* Vary the text between requests, as your real traffic does.
* Send requests one after another. A request beyond your plan's [concurrency limit](/build/guides/concepts/api-limits#concurrency-limits) returns `429`, not a slower response.
* Take at least 20 requests and report p50 and p90, not the average or the single best run.
* Measure with PCM. With Ogg, the first bytes are header pages, not audio.

#### Python

```python
# pip install httpx
import os
import re
import statistics
import time

import httpx

# One client for every request, so the connection is reused.
client = httpx.Client(
    base_url="https://api.speechify.ai",
    headers={"Authorization": f"Bearer {os.environ['SPEECHIFY_API_KEY']}"},
    timeout=30.0,
)


def measure(text: str) -> tuple[float, float, str]:
    start = time.perf_counter()
    first_chunk_ms = None
    with client.stream(
        "POST",
        "/v1/audio/stream",
        headers={"Accept": "audio/pcm"},
        json={"input": text, "voice_id": "geffen_32", "model": "simba-3.2"},
    ) as response:
        response.raise_for_status()
        for _ in response.iter_raw():
            if first_chunk_ms is None:
                first_chunk_ms = (time.perf_counter() - start) * 1000
    total_ms = (time.perf_counter() - start) * 1000
    match = re.search(r"ttfb;dur=([\d.]+)", response.headers.get("server-timing", ""))
    if first_chunk_ms is None:
        first_chunk_ms = total_ms
    return first_chunk_ms, total_ms, match.group(1) if match else "n/a"


for n in range(2):
    measure(f"Opening the connection, request {n}.")  # not counted

first_chunks = []
for n in range(20):
    first_chunk_ms, total_ms, server_ttfb = measure(
        f"Your order {1000 + n} shipped this morning and arrives on Thursday."
    )
    first_chunks.append(first_chunk_ms)
    print(f"first chunk {first_chunk_ms:.0f} ms, server ttfb {server_ttfb} ms, total {total_ms:.0f} ms")

deciles = statistics.quantiles(first_chunks, n=10)
print(f"first chunk p50 {deciles[4]:.0f} ms, p90 {deciles[8]:.0f} ms")
```

#### TypeScript

```typescript
// Node 18+. The built-in fetch reuses connections to the same origin.
const apiKey = process.env.SPEECHIFY_API_KEY!;

async function measure(text: string) {
  const start = performance.now();
  const response = await fetch("https://api.speechify.ai/v1/audio/stream", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      Accept: "audio/pcm",
    },
    body: JSON.stringify({ input: text, voice_id: "geffen_32", model: "simba-3.2" }),
  });
  if (!response.ok || !response.body) {
    throw new Error(`HTTP ${response.status}: ${await response.text()}`);
  }
  const reader = response.body.getReader();
  let firstChunkMs: number | undefined;
  for (;;) {
    const { done } = await reader.read();
    if (done) break;
    firstChunkMs ??= performance.now() - start;
  }
  const totalMs = performance.now() - start;
  const serverTtfb =
    /ttfb;dur=([\d.]+)/.exec(response.headers.get("server-timing") ?? "")?.[1] ?? "n/a";
  return { firstChunkMs: firstChunkMs ?? totalMs, totalMs, serverTtfb };
}

for (let n = 0; n < 2; n++) {
  await measure(`Opening the connection, request ${n}.`); // not counted
}

const firstChunks: number[] = [];
for (let n = 0; n < 20; n++) {
  const { firstChunkMs, totalMs, serverTtfb } = await measure(
    `Your order ${1000 + n} shipped this morning and arrives on Thursday.`,
  );
  firstChunks.push(firstChunkMs);
  console.log(
    `first chunk ${firstChunkMs.toFixed(0)} ms, server ttfb ${serverTtfb} ms, total ${totalMs.toFixed(0)} ms`,
  );
}

firstChunks.sort((a, b) => a - b);
const percentile = (p: number) => firstChunks[Math.ceil((p / 100) * firstChunks.length) - 1];
console.log(`first chunk p50 ${percentile(50).toFixed(0)} ms, p90 ${percentile(90).toFixed(0)} ms`);
```

#### cURL

```bash
curl -sS -o speech.pcm -D headers.txt \
  -X POST "https://api.speechify.ai/v1/audio/stream" \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: audio/pcm" \
  -d '{"input":"Your order shipped this morning and arrives on Thursday.","voice_id":"geffen_32","model":"simba-3.2"}' \
  -w 'dns        %{time_namelookup}s\nconnect    %{time_connect}s\ntls        %{time_appconnect}s\nfirst byte %{time_starttransfer}s\ntotal      %{time_total}s\n'

grep -i '^server-timing' headers.txt
```

curl's timings are cumulative from the start of the run. Each run opens a new connection, so `first byte` minus `tls` is the request itself on an open connection. The response headers are sent with the first audio chunk, so `first byte` is the first audio.

## FAQ

#### Is 56 ms a guarantee?

No. It is the median first byte measured on production streaming traffic on 15 Sep 2026, with p90 at 102 ms, not a service-level commitment. The `Server-Timing` header on each of your own responses reports the same kind of measurement for that request.

#### Why do I measure more than 56 ms?

Your measurement includes the network between you and US East, the handshakes of any new connection, and, depending on the tool, decoding and playback. Compare your first-chunk time with `ttfb` in the `Server-Timing` header: the difference is everything outside our API.

#### Which endpoint, model and format give the lowest time to first audio?

`POST /v1/audio/stream` with `model: "simba-3.2"` and `Accept: audio/pcm`, from a client that reuses its connection and runs near US East. `simba-3.2` is English only; use `simba-3.0` for the [other languages it supports](/build/guides/text-to-speech/language-support#simba-30-languages).

#### Does longer text delay the first audio?

On `POST /v1/audio/stream`, audio starts arriving before the whole input has been synthesized, so send text you already have in one request; total generation time grows with length. On `POST /v1/audio/speech`, the response waits for the whole clip, so the first audio grows with length.

#### Where is the API served from?

As of September 2026, every request to `api.speechify.ai` is served from US East, and the measured numbers on this page are for that path.

## Related

#### [Streaming](/build/guides/text-to-speech/streaming)

The streaming endpoint, its formats and mid-stream errors.

#### [Streaming TTS Guide](/build/streaming-tts-guide)

Step by step: stream audio and start playback early.

#### [Models](/build/guides/concepts/models)

Every live model and how to choose one.

#### [Audio formats](/build/guides/concepts/audio-formats)

Every `output_format` value and how it relates to `Accept`.