> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Changelog

## API: `pcm_16000` is 16 kHz audio from version `2026-09-30`

From API version `2026-09-30`, `output_format: pcm_16000` returns 16 kHz audio on every synthesis route: `POST /v1/audio/stream`, `POST /v1/audio/stream/with-timestamps` and `POST /v1/audio/speech`.

**What was wrong.** On the Simba 3 models, `pcm_16000` returned 24 kHz audio while the `Content-Type` said `rate=16000`. Played at the labelled rate, speech ran 1.5x slow and pitched down.

**What changes for you.** Nothing, until you move to `2026-09-30`. A workspace or request pinned to an earlier version keeps receiving exactly the bytes it receives today, so an integration that already plays `pcm_16000` at 24 kHz keeps working.

When you move your pin:

- If you feed `pcm_16000` into a 16 kHz pipeline (telephony, LiveKit SIP, Twilio), it now plays correctly with no change on your side.
- If you worked around the old audio by playing it at 24 kHz, switch that step to 16 kHz, or request `pcm_24000` to keep 24 kHz audio.

New workspaces start on `2026-09-30`.

See the [API Versioning guide](/build/guides/concepts/api-versioning) for how to read and set your workspace's pinned version.

## API: projects, webhook endpoints and workspace entitlements are in the API reference and SDKs

Three workspace surfaces that Build integrations already rely on are now in the API Reference and the official SDKs:

- **Projects** (`/v1/projects`): model each of your customers as a project, cap its spend with `monthly_budget`, pin API keys and members to it, and archive, purge or restore it.
- **Webhook endpoints** (`/v1/webhooks/endpoints`): subscribe to the `workspace.spend_budget.*`, `project.spend_budget.*` and `api_key.spend_cap.*` events, rotate the signing secret, and read delivery attempts.
- **Workspace entitlements** (`GET /v1/workspaces/current/entitlements`): what your plan allows, including the TTS request-rate and concurrency limits.

Nothing changes on the wire. These endpoints already answered API-key requests; this release documents them and adds them to the SDKs.

See [Spend Limits](/build/guides/concepts/spend-limits) and [Projects](/build/guides/concepts/projects) for how they fit together.

## API: voice cloning requires verified consent on every API version

**The `consent` form field on `POST /v1/voices` is switched off, on every API version, whatever your workspace is pinned to.** Creating a cloned voice now requires proof that the speaker agreed: a consent challenge, and a recording of the speaker reading the phrase it issues. This is the switch-off the [2026-08-13 entry](/build/changelog/2026/8/13) announced and said would be dated here.

**What changes.** A `POST /v1/voices` that sends `consent` and no `consent_challenge_id` returns `400` with the error code `consent_verification_required`; the message links the migration guide. Requests that already send a challenge are unaffected, and so is every existing cloned voice and every synthesis call.

**What to do.** Create a challenge with `POST /v1/voices/consent-challenges`, show the speaker the phrase it returns, record them reading it, and send `consent_challenge_id` and `consent_recording` on the create. Re-pin `Speechify-Version: 2026-09-13` while you are there. The [migration guide](/build/migrating-voice-cloning-consent) walks the two calls, and the [consent guide](/build/guides/voice-cloning/consent) covers the flow end to end.

**Why the window was shorter than 12 months.** Our default sunset is a year, and this one was not, exactly as the 2026-08-13 entry said it would be. An endpoint that clones a voice without checking the speaker agreed is a safety liability rather than an old request shape.

**If this catches you mid-migration, [contact support](mailto:support@speechify.ai)** with your workspace and what you are moving. We keep the old flow open for individual workspaces with an end date while they finish.

## API: Simba 1.6 is retired at version `2026-09-21` and served by current models from 2026-11-21

`simba-english` and `simba-multilingual` - the Simba 1.6 pair - are being withdrawn in two steps:

- **From API version `2026-09-21`** they are no longer selectable. Naming either returns `400` with the error code `model_retired`.
- **On 2026-11-21** both ids are served by our current models instead of their Simba 1.6 training. They keep answering on every API version that can still name them; what changes is the model behind the id, not your integration.

| If you send | From 2026-11-21 it is served by |
|---|---|
| `simba-multilingual` | our current multilingual model, in every language you send it today |
| `simba-english`, English text | `simba-3.2` |
| `simba-english`, any other language | our current multilingual model |

**Nothing breaks on that date, but the audio changes.** A voice you tuned on Simba 1.6 is rendered by a different model, and SSML `<prosody pitch>`, `<prosody volume>`, `<emphasis>` and `<speechify:style emotion>` are not applied by the current models, though breaks and `<prosody rate>` are. The request, the response, the voice IDs and the price stay the same.

**To choose the model yourself, migrate before then** - for most integrations that is a one-line change. [Migrating off Simba 1.6](/build/migrating-off-simba-1-6) walks it step by step. Pinning your workspace's API version to a date before `2026-09-21` keeps the Simba 1.6 training in the meantime with no code change at all.

**Why the shorter notice.** Our usual sunset window is 12 months. Simba 1.6 is a previous-generation family: it cannot serve `/v1/audio/stream/with-timestamps` at all and runs at roughly 2.5x the time-to-first-byte of the streaming-native models. Consolidating onto the Simba 3 fleet is what lets us keep improving latency and quality for everyone, and we did not want to spend a year running two stacks to do it.

**Where to go if you migrate yourself.**

| From | To | Notes |
|---|---|---|
| `simba-english` | `simba-3.2` | English. Lower time-to-first-byte, richer expressivity, streaming-native. Serves every English voice in the catalog, your own cloned voices included. |
| `simba-english` | `simba-3.0` | English, and the API default. Pick it if you also need German, Spanish, French, Italian or Brazilian Portuguese. Accepts every catalog voice. |
| `simba-multilingual` | `simba-3.0` | English, `de-DE`, `es-ES`, `es-MX`, `fr-FR`, `it-IT`, `pt-BR`. Streaming-native, and a cloned voice speaks all of them from one voice ID. |

An omitted `model` is unaffected: it already resolves to `simba-3.0`.

**If you synthesize outside Simba 3.0's seven locales**, `simba-3.0` on its own is not a like-for-like replacement, and we are not asking you to drop those languages. Broader multilingual coverage on the current model generation is what serves `simba-multilingual` from 2026-11-21, and it reaches the full set the pair covers today. [Talk to us](https://speechify.ai/talk-to-sales) if you want to test it before the date.

**What changed at this version.** `POST /v1/audio/speech` and both `/v1/audio/stream` routes reject the two ids. `GET /v1/audio/models` returns only the models you can actually call, and each voice's `models` array in `GET /v1/voices` does the same - so a picker driven off either endpoint stays correct without special-casing.

**Read the dates off the API.** While your workspace is pinned below `2026-09-21`, both models still appear in `GET /v1/audio/models` carrying `retired_at: "2026-09-21"` and `sunset_at: "2026-11-21"`. `sunset_at` is the day the model behind the id changes; surface it in your own tooling if you have a deadline to track.

See the [API Versioning guide](/build/guides/concepts/api-versioning) for how to read and set your workspace's pinned version.

## API: simba-3.2 serves every English voice

`simba-3.2` no longer requires a voice to be registered for it. Every English voice in `GET /v1/voices` now synthesises on it, your workspace's own cloned voices included, and each voice's `models` array lists it. Before today it served eight registered voices out of a catalogue of nearly a thousand, and any other voice returned `400`.

Nothing about the request changes, and nothing you already send breaks: the eight voices built for the model are served by exactly the training they were before, and every other voice is served by the model's zero-shot training — the same one that has served cloned voices on `simba-3.2` since 2026-08-26. Which training runs is an internal routing detail; the request, the response, the latency class and the audio format are identical either way.

This is what makes the Simba 1.6 retirement on 2026-09-21 a one-line change. `simba-english` and `simba-multilingual` are served by our current models from 2026-11-21, and the 953 catalogue voices registered for them can now move straight to `simba-3.2` — an English voice — or to `simba-3.0` for any other language, with no voice substitution.

**`curated_voices` on `GET /v1/audio/models` is deprecated and is now `false` for every model.** No model restricts synthesis to a registered voice set any more, so a picker that filters on the flag simply stops filtering. The field stays on the response; read `english_voices_only` instead, which is the one voice constraint left — a model with no multilingual training returns `400` for a non-English voice.

## API: watermark verification with no credential

A new endpoint answers whether a clip carries a Speechify watermark without any credential. `POST /v1/audio/watermark/verify` takes the same audio upload as `POST /v1/audio/watermark/detect` and returns a bare `{"watermarked": true|false}` — no key, no session, nothing to authenticate.

It is the API half of the public tool at [speechify.ai/detect](https://speechify.ai/detect), and it exists because California's AI Transparency Act (BPC 22757.2) requires a detection tool that is publicly accessible and invokable without visiting a website.

`verify` answers; `detect` measures. The new verb deliberately omits the confidence score its sibling returns — a public score is a gradient to optimise against. Keep using `POST /v1/audio/watermark/detect` with an API key when you want the score.

Because `verify` takes no credential, it is rate-limited per client address and draws on a shared platform budget, so expect `429` under sustained automated use. Nothing changes for an existing integration: `POST /v1/audio/watermark/detect`'s path, request and response are untouched.

## API: watermark detection with your API key

`POST /v1/audio/watermark/detect` now answers, for any credentialed workspace, whether a clip carries the watermark Speechify seals into the audio it generates. This check was previously operator-only; it is now on the API and in the console. Upload the clip as `audio` (`multipart/form-data`, at most 25MB) and get back `WatermarkDetectionResponse` — `{ watermarked, confidence }`, with `confidence` a score in `[0, 1]`. Nothing about the upload is stored, and no voice is read or written.

Read the answer in one direction only. `watermarked: true` is positive evidence the audio came from Speechify synthesis. `watermarked: false` is the **absence** of that evidence, not proof of a negative: only models redeployed since the watermark shipped mark their output, the detector needs at least three seconds of clear speech, and re-encoding or changing the speed of a clip degrades the mark. Confidence scores are comparable only between checks against the same detector version.

Detection is capped at 20 requests per hour per workspace — a forensic question, not a data-plane call.

- **`422 watermark_audio_unusable`** — the clip could not be read (an undecodable container, or too little audio to judge). Deliberately not a negative verdict: "we could not tell" and "this is not ours" are different answers.
- **`502 watermark_detection_unavailable`** — the detector could not answer. Nothing about the request is wrong; retry the same bytes.

The credential-free sibling `POST /v1/audio/watermark/verify` — a bare boolean with no confidence score — follows the next day (see 2026-08-29); use `detect` with an API key when you want the score.

## API: `simba-3.2` voice cloning is now self-serve for every workspace

Cloned (personal) voices synthesize on `simba-3.2` for every workspace, with no enablement step. The limited release announced on 2026-08-06 is over and the per-workspace allow-list behind it is gone; you no longer need to contact us.

Nothing else changes. The request and response are identical to a stock-voice call, and `simba-3.2` remains English only, so a cloned voice with a non-English locale still returns `400` — use `simba-3.0` for those. `GET /v1/voices` now names `simba-3.2` on your cloned voices without any per-workspace condition, and driving a picker off each voice's `models` array remains the right pattern. Cloning on `simba-3.0`, `simba-english` and `simba-multilingual` is unchanged.

## API: Projects — group resources, scope credentials, and attribute spend

The `/v1/projects` endpoints are now in the public API reference. A **project** groups the resources you create inside a workspace and the spend you incur from them, so one workspace can run several environments or several end customers without splitting into separate accounts.

Every workspace has an implicit **Default project**: any resource with no project lives there, and nothing you already send changes — the surface is additive and opt-in.

What a project groups:

| Kind | Belongs to a project |
|---|---|
| Audio assets | Yes, and can be moved later |
| API keys and service accounts | Yes — a pin fixed when the credential is created |
| Vault credentials and webhook endpoints | One project, or workspace-wide |
| Usage and spend | Attributed through the calling credential's pin |
| Cloned voices | From the pin on the creating credential — except a consent-verified clone, which is always workspace-wide |
| The public voice and model catalog | No — workspace-wide |

Manage the lifecycle with `POST/GET/PATCH/DELETE /v1/projects` and `.../{project_id}`, plus `archive`, `unarchive`, `restore`, `teardown`, `stats`, `audit`, `promote`, and the `members` sub-tree (`grant`/`revoke` access).

Filtering: every list endpoint that takes a project accepts a `project_id` query parameter. Omit it to get everything you can reach, pass a `proj_...` id for one project, or pass the literal `default` for the implicit Default project. On lists whose rows can be workspace-wide — credentials, webhook endpoints and cloned voices — that literal is `shared` instead, because an absent project there means workspace-wide rather than Default.

Names are unique per workspace, case-insensitively; a workspace holds at most 100 live projects, and at the cap the create is refused with `409 project_limit_reached`.

A project is a filter and a grouping, **not a security boundary** — it scopes what a credential may reach, what a scoped member sees, and where spend lands. If one team must be unable to see another's data at all, use a separate workspace. See [Projects](/build/guides/concepts/projects).

## Per-project capacity ceilings

A project can now carry a capacity ceiling alongside its spend limit: `max_requests_per_minute`, the most API requests per minute from credentials pinned to the project, across every surface. It is read on the Project, set via `PATCH` merge-patch, and cleared by sending `null`. Setting it requires the `billing.manage` permission.

It must sit at or below the workspace's own plan cap; a value above it is refused with `400 validation_failed` naming the field and the ceiling. A project can only narrow the workspace's capacity, never raise it. It is checked **after** the workspace's own caps: a request over the rate ceiling is refused with `429 rate_limited` from a per-project bucket — the same code the workspace limit already answers. Sibling projects keep their own headroom; console sessions and unpinned keys carry no project and are subject to neither.

## Error codes

- **`404 project_not_found`** — no project with that id, or one your credential cannot reach.
- **`409 project_limit_reached`** — the workspace already holds the maximum live projects its plan allows (up to 100); delete an unused one to free a slot.
- **`409 project_archived`** — work would start or bill inside an archived project (synthesis on a credential pinned to it, for example). Data stays readable and configuration stays editable; unarchive to resume — no balance or ceiling change clears it. The project consulted is the one the work is attributed to, so a workspace-wide key can be refused too.
- **`409 project_not_archived`** — a destructive teardown was asked for on a live project. Archive is the reversible pending-deletion state; archive first, then purge from there.
- **`402 project_spend_limit_exceeded`** — the project's Orb-rated spend this calendar month reached its limit, the middle ceiling between the workspace budget and a key's cap. Raise that project's limit, move the work, or wait for the monthly reset. The project charged is the one the spend is attributed to, so a workspace-wide key can trip it.
- **`409 project_has_scoped_credentials`** — a delete was refused because API keys or service accounts are still pinned to the project; the listed credentials ride `error.details`. Revoke or re-mint them elsewhere first — a pin is never silently widened to the whole workspace.
- **`409 cross_project_reference`** — a move was refused because live references tie the resource to its current project; the blockers ride `error.details.referrers` so you can move them together. Detaching them always unblocks.
- **`409 project_too_large_to_promote`** — the project holds more resources than one synchronous `promote` may copy.
- **`409 project_restore_window_expired`** — a `restore` arrived after the purge retention window closed. Distinct from `project_not_found` on purpose: the project is gone for good, not an id you mistyped.
- **`409 project_not_purged`** — a `restore` was asked for on a project that has not been purged. Not a 404: the project exists and you can still see it.
- **`409 project_name_taken`** — a `restore` cannot reinstate the project because another has taken its name since the purge (a purge frees the name immediately). Rename the holder, then restore again.

## Filtering request logs by project

The request-log and analytics `project_id` filter adds a third literal beyond the `default` and `shared` used on the resource lists above: `unattributed` narrows to workspace-level traffic that names no project — an unpinned key, or a console session. That is **not** the Default project, and a project filter cannot select Default, because a key can only ever be pinned to a project you created.

## API: default MP3 output is now 128 kbps

Requests that ask for MP3 without naming a bitrate now receive **128 kbps** instead of 64 kbps. That is `audio_format: "mp3"` on `POST /v1/audio/speech` and `POST /v1/audio/dialogue`, and `Accept: audio/mpeg` on `POST /v1/audio/stream` and `POST /v1/audio/stream/with-timestamps`. The sample rate, channel count, duration and container are unchanged: 24 kHz mono MP3, as before.

At 24 kHz mono, 64 kbps left compression noise only about 12-14 dB below the signal across the 2-11 kHz band, which is audible on sibilants and on the quiet tail of a phrase. 128 kbps puts that noise about 30 dB down, which is where the quality curve stops paying for itself - 160 kbps measures roughly 1 dB better again for 25% more bytes.

**What this means for your integration:**

- Responses are about **twice as large** for the same text. Time-to-first-byte is unchanged: the encoder emits its first frame after a fixed amount of audio regardless of bitrate.
- Nothing in the response shape changes. `audio_format` still reads `mp3`, and `output_format` is still echoed only when you set it.
- **On `/v1/audio/speech` and the two stream routes, keep the previous output by sending `output_format: "mp3_24000_64"`.** Any request that already names an `output_format` is unaffected - this changes only what an unspecified bitrate resolves to.
- **`POST /v1/audio/dialogue` has no `output_format` field**, so it has no per-request bitrate control and never did: `audio_format: "mp3"` there simply encodes at 128 kbps now instead of 64. Use `wav` or `pcm` if you need the unencoded audio.

The set of `output_format` values you can ask for is unchanged. As the `audio_format` field has always documented, the resolved default is not part of the contract: name the `output_format` you want if your pipeline depends on it.

## API: maximum-fidelity mp3, and a fix for the 22.05 kHz formats

`output_format` gains `mp3_22050_160` and `mp3_24000_160`, the highest-fidelity mp3 the format can carry: 160 kbps is the ceiling at these sample rates, so there is no higher mp3 to ask for. A request for `mp3_22050_192` or `mp3_24000_192` was always encoded at 160 kbps, and the response now reports the `mp3_*_160` it actually delivered rather than echoing the 192 back. Nothing about those two requests changes on the wire except the reported value, and no existing integration has to move.

Every `mp3_22050_*` format also plays at the right speed now. They were being encoded about 8% fast and close to a semitone high; if you compensated for that in your own pipeline, remove the correction. `mp3_24000_*`, the default for `audio_format: "mp3"`, was never affected.

Behind both: any sample-rate conversion we do is now a high-quality resample from the model's native 24 kHz rather than the encoder default, which keeps the audio flat to about 10.7 kHz at 22.05 kHz instead of rolling off from 10 kHz. That applies to `pcm_22050`, `pcm_44100`, `pcm_48000`, `pcm_8000` and `ulaw_8000` as well.

The two `mp3_*_160` formats are produced by the Simba 3 models; pairing one with a legacy Simba 1.6 model returns `400`.

## API: consent recordings are matched against the voice being cloned

Consent verification now checks who is speaking, not only what they said. A create whose `consent_recording` reads the phrase correctly but in a different voice from your `sample` is refused with `422 consent_speaker_mismatch` and no voice is created.

This closes the gap the flow was always meant to close: the person consenting has to be the person being cloned, so permission relayed by somebody else - an account holder reading the phrase on a speaker's behalf - does not pass.

`consent_speaker_mismatch` is not a new code, and clients that already branch on it need no change. What is new is that it fires. If you tested your integration by reading the phrase yourself against someone else's sample, that path now returns a 422; put the challenge in front of the speaker instead. Re-reading the same phrase does not help, which is why this is a distinct code from `consent_phrase_mismatch`. See [Consent](/build/guides/voice-cloning/consent).

## API: voice cloning now verifies the speaker's consent

Creating a cloned voice now requires proof that the speaker agreed to it, in place of the `consent` object you used to send.

The flow adds one call in front of your existing create. `POST /v1/voices/consent-challenges` with the speaker's `full_name` returns a `phrase` and an `id`. Show the phrase to the speaker exactly as it comes back, record them reading it aloud, and send that recording as `consent_recording` with `consent_challenge_id` on `POST /v1/voices`. Speechify transcribes the recording, checks it against the phrase it issued, and keeps it as the consent record for that voice.

A challenge is single use, bound to your workspace, and short-lived, so create it when your speaker is ready to record rather than at the start of your flow. If it expires, create another one and record again.

**Update, 2026-09-23: that date arrived.** The unverified flow is switched off on every API version - see the [2026-09-23 entry](/build/changelog/2026/9/23) for the error it now returns and the two calls that replace it.

**The unverified flow is deprecated and will be switched off.** The date will be announced in this changelog and to affected workspaces ahead of time; plan for a window deliberately shorter than the [standard 12-month sunset](/build/guides/concepts/api-versioning), because an endpoint that clones a voice without checking the speaker agreed is a safety liability, not just an old shape. The new shape is `Speechify-Version: 2026-09-13`; it is callable now by pinning that version, and it becomes the default for new workspaces on that date. Until the switch-off, workspaces pinned to earlier versions keep the old `consent` object, and a pinned default does not move on its own: migrating means re-pinning `2026-09-13`. Existing cloned voices are unaffected and keep working, and synthesis endpoints are unchanged.

If you cannot migrate ahead of the switch-off, contact [support](mailto:support@speechify.ai) and we will work out an extension for your workspace.

One thing to know if you use an SDK: each release sends its own build date as the default version, so **upgrading to an SDK published on or after 2026-09-13 moves you to the new flow** even though your workspace is pinned to the old one. That is a deliberate break rather than a silent one - `consent_challenge_id` and `consent_recording` are required arguments on the new `create`, and the `consent` argument is gone, so the call stops building rather than failing at runtime. To upgrade the SDK without migrating yet, pass the version explicitly:

<Tabs>
  <Tab title="Python">
    ```python
    client = Speechify(token=os.environ["SPEECHIFY_API_KEY"], version="2026-08-07")
    ```
  </Tab>
  <Tab title="TypeScript">
    ```typescript
    const client = new SpeechifyClient({ token: process.env.SPEECHIFY_API_KEY, version: "2026-08-07" });
    ```
  </Tab>
  <Tab title="cURL">
    ```bash
    curl -H "Speechify-Version: 2026-08-07" ...
    ```
  </Tab>
</Tabs>

Migrating: `full_name` moves from the `consent` object onto the challenge call, `email` is dropped and nothing replaces it, and `consent_challenge_id` plus `consent_recording` become required.

The challenge itself can be refused before the recording is ever assessed. These do not share the 422 status, so branch on the code here too:

- **`404 consent_challenge_not_found`** — the `consent_challenge_id` does not resolve. A challenge minted for a different workspace answers with this same code and message, so a challenge id is never confirmed to exist by probing.
- **`409 consent_challenge_expired`** — the challenge is past its `expires_at`. Mint a new challenge and record again; the sample, the name, and the rest of the create are still good.
- **`409 consent_challenge_already_used`** — the challenge was already consumed. A challenge is single use, so this is also what a retry of a create whose response you never received returns: after a success, mint a fresh challenge rather than replaying the old one.
- **`502 consent_verification_unavailable`** — the verification backend could not answer. Nothing about the request is wrong, so this is transient — retry the same recording shortly.

Three of the new error codes share HTTP 422 and mean different things, so branch on the code rather than the status: `consent_phrase_mismatch` (the phrase was misread - read it again), `consent_speaker_mismatch` (the person in the recording is not the person in the sample - the speaker consenting has to be the speaker being cloned), and `consent_recording_unusable` (silence, too little speech, or an unreadable file - record it again). See [Consent](/build/guides/voice-cloning/consent) and the [Voice Cloning API](/build/voice-cloning-api).

## API: text is screened before synthesis

Text you send for synthesis is now screened before Speechify produces audio from it. A request whose content is not permitted returns `400 content_policy_violation` with no audio, and is not billed. This applies to `/v1/audio/speech`, `/v1/audio/stream` and `/v1/audio/stream/with-timestamps`.

Most published work is unaffected - fiction, journalism, true crime and court reporting routinely describe or quote violence, and depicting that material is treated differently from producing it.

One thing worth handling in your client: `content_policy_violation` is a persistent error, so retrying the same text will be refused again. On the streaming endpoints the decision is always made before the first audio byte, so a `200` means the request passed and a refusal is always a JSON `400`, never a truncated or empty audio file. See [Content Policy](/build/guides/concepts/content-policy).

## API: `simba-3.2` voice cloning enters limited release

Cloned (personal) voices now synthesize on `simba-3.2` for workspaces enabled for it, with no per-voice step. The earlier per-voice-key approval is gone: enablement is per workspace, and once yours is on, every clone you own works there. [Contact us](https://speechify.ai/talk-to-sales) to be enabled.

Nothing else changes. The request and response are identical to a stock-voice call, and `simba-3.2` remains English only, so a cloned voice with a non-English locale still returns `400` — use `simba-3.0` for those. `GET /v1/voices` names `simba-3.2` on your cloned voices once your workspace is enabled, so drive a picker off each voice's `models` array rather than assuming. Cloning on `simba-3.0`, `simba-english` and `simba-multilingual` is unchanged.

## API: `simba-3.0` becomes the default; `simba-3.2` is the newest model

<Note>
**`simba-3.0` is the default model — the one you get when you don't provide one — but the newest model is `simba-3.2`.** We recommend `simba-3.2` for English integrations; `simba-3.0` stays the default for the compatibility reasons below. For the current list of every text-to-speech model live today, see [Models](https://docs.speechify.ai/build/guides/concepts/models).
</Note>

`POST /v1/audio/speech`, `POST /v1/audio/stream` and `POST /v1/audio/stream/with-timestamps` now resolve a request that omits `model` to **`simba-3.0`** instead of the legacy `simba-english`. `GET /v1/audio/models` marks the change on its `default` flag.

**Nothing you already send changes shape, and nothing that worked starts failing.** A request that names a model explicitly is untouched - `model: "simba-english"` keeps getting Simba 1.6 English, and that model stays fully supported with nothing scheduled for removal.

What changes if you omit `model`:

| | Before (`simba-english`) | Now (`simba-3.0`) |
| --- | --- | --- |
| Voices accepted | any voice in `GET /v1/voices` | unchanged - any voice, cloned voices self-serve |
| Non-English voices | synthesized by the English 1.6 training | routed to the Simba 3.0 multilingual training |
| `POST /v1/audio/stream/with-timestamps` | `400 speech_marks_unsupported` | supported |
| Audio | Simba 1.6 | Simba 3.0 - streaming-native, lower time-to-first-byte |

`simba-3.0` was chosen over the recommended `simba-3.2` precisely because it accepts everything the old default did: `simba-3.2` serves a curated voice set and rejects non-English voices, so making *it* the default would have turned working calls into `400`s.

The rendered audio does change if you omit `model`. **Pin the old behaviour by sending `model: "simba-english"` explicitly** if your integration depends on the Simba 1.6 output. For new English work we still recommend `model: "simba-3.2"` - see [Models](https://docs.speechify.ai/build/guides/concepts/models).

## Models list now reports per-model endpoints and curated-voice flag

`GET /v1/audio/models` now returns two new fields for each model:

- **`endpoints`** — the synthesis routes this model may be passed to. Passing a model to an endpoint its `endpoints` list omits returns 400.
- **`curated_voices`** — when `true`, only voices that explicitly name this model in their `models` array are accepted. When `false`, every catalogue voice works, including workspace clones.

Both fields are surfaced to help a model picker show only valid combinations and reject unsuitable voice selections before the request reaches the server.

## API: Free tier gets a burst allowance on `/v1/audio/*`

The Free plan's TTS rate limit no longer caps its burst bucket at the sustained rate. Previously Free was 1 request/second with a bucket capacity of 1, so a second request issued in the same second was rejected with `429`. Free now gets a burst capacity of 10, so a normal opening burst of requests (a quickstart script, a first integration test) no longer trips the limiter.

| Plan | Sustained requests/second | Burst |
| --- | --- | --- |
| Free | 1 | 10 |
| Starter | 20 | 60 |
| Pro | 40 | 120 |
| Scale | 80 | 240 |
| Enterprise | 150 | 450 |

**Sustained throughput is unchanged on every tier** — burst only smooths the first second of traffic; it does not raise your steady-state rate. Burst allowances were first published with the [rate and concurrency limits](https://docs.speechify.ai/build/changelog/2026/7/19).

See the [API limits reference](https://docs.speechify.ai/build/guides/concepts/api-limits) for the full per-plan table.

## Audio 404s now say what wasn't found

`POST /v1/audio/speech`, `POST /v1/audio/stream`, and `POST /v1/audio/stream/with-timestamps` previously answered an unknown `voice_id` or `model` with an opaque, passed-through upstream `404` - indistinguishable from any other not-found. All three now classify the cause and return an actionable error instead:

| Cause | Code | Message |
| --- | --- | --- |
| Unknown `voice_id` | `voice_not_found` | Voice not found. List the voices available to your workspace with `GET /v1/voices`. |
| Unknown `model` | `not_found` | Model not found. List the available models with `GET /v1/audio/models`. |

`voice_not_found` is already part of the public `ErrorCode` enum; this is the first time these three endpoints return it instead of a generic upstream passthrough. No other error responses on these endpoints changed.

## API: canonical `Speechify-*` header names (legacy `X-` aliases still work)

Every header in Speechify's public API surface now has one canonical, un-prefixed `Speechify-*` name (RFC 6648). The pre-2026 `X-`-prefixed spellings are unaffected today - they're still accepted on requests and still emitted on responses - and will keep working until **2027-07-24**.

| Legacy (still works until 2027-07-24) | Canonical |
| --- | --- |
| `X-Request-ID` | `Speechify-Request-Id` |
| `X-Speechify-Audio-Content-Type` | `Speechify-Audio-Content-Type` |
| `X-RateLimit-Limit` / `-Remaining` / `-Reset` | `RateLimit-Limit` / `-Remaining` / `-Reset` |
| `X-Speechify-SDK` / `X-Speechify-SDK-Version` | `Speechify-SDK` / `Speechify-SDK-Version` |
| `X-Tenant-ID` | `Speechify-Tenant-Id` |

The last two are request headers - if your integration sends them, either spelling is read the same way. `X-Speechify-Billable-Characters-Count` (an internal usage-accounting header on the audio-streaming responses) is unaffected by this change and has no canonical form.

Nothing to change today: whenever a response carries one of these headers, it carries both spellings with the same value, so an integration reading either name keeps working unmodified. `Speechify-Audio-Content-Type` is only present on the audio-streaming endpoints that already emitted its `X-` predecessor - it doesn't appear on responses that never had it, like `POST /v1/audio/speech`. New integrations should read the canonical name.

```
Speechify-Request-Id: 7f3a2c1b4d5e6f7a
X-Request-ID: 7f3a2c1b4d5e6f7a
RateLimit-Limit: 20
X-RateLimit-Limit: 20
```

(`RateLimit-Limit` reflects your plan's request-rate budget - the value above is illustrative, not a specific plan's limit. See the [API limits reference](https://docs.speechify.ai/build/guides/concepts/api-limits) for the per-plan table.)

If you send a request-correlation id yourself, send it as `Speechify-Request-Id` (the legacy `X-Request-ID` request header is still echoed back the same way). If both are sent, the canonical name wins.

Full header reference: [API limits & rate limiting](https://docs.speechify.ai/build/guides/concepts/api-limits).

## API: stream speech marks with `POST /v1/audio/stream/with-timestamps`

A new endpoint streams word-level speech marks alongside audio, so text highlighting, captions, and audio-text sync no longer need the batch `POST /v1/audio/speech` round trip. `POST /v1/audio/stream` is unchanged - same request body, still plain audio.

```bash
curl -N -X POST https://api.speechify.ai/v1/audio/stream/with-timestamps \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Hello world.",
    "voice_id": "oliver",
    "model": "simba-3.2"
  }'
```

The response is a Server-Sent Events stream:

```
event: speech.chunk
data: {"audio":"<base64>","speech_marks":[{"type":"word","start":0,"end":5,"start_time":0,"end_time":512,"value":"Hello"}]}

event: speech.done
data: {"billable_characters_count":60,"audio_duration_ms":5350}
```

- `speech.chunk` carries a Base64-encoded run of audio, the speech marks that became final with it, or both - either field may be absent, and the last chunk of a stream is often marks-only.
- `speech.done` is terminal; there is no `[DONE]` sentinel.
- `speech.error` carries the standard error envelope when a failure happens after the stream has started (the status code is already committed by then).
- Ignore any event type you do not recognize, so new event types never break your integration.

Speech-mark times are absolute milliseconds from the start of the synthesis - concatenate the audio chunks into one stream and apply the marks against that single timeline. Which chunk a mark arrives on is a delivery detail with no meaning of its own, and times stay correct across every `output_format`.

Marks are produced by the streaming-native models: use `simba-3.0` or `simba-3.2`. `simba-english` and `simba-multilingual` return `400 speech_marks_unsupported` on this endpoint - use `POST /v1/audio/speech` for a non-streamed response with marks on any model.

_Showing the 20 most recent of 64 entries. Append `/llms.txt` to the changelog URL for the complete index._