> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/SpeechifyInc — GitHub org. `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family (1.6 multilingual, 3.0 streaming multilingual, 3.2 streaming English), not the brand. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Models

> Compare Speechify's streaming-native TTS models. Simba 3.2 is recommended for English; Simba 3.0 is the API default with six languages across seven locales. Both support voice cloning on paid plans.

## Available models

Below is every Speechify text-to-speech model live today. For the authoritative list at runtime, call [`GET /v1/audio/models`](#listing-models-via-the-api) — it always returns exactly the models your workspace can use.

| Model                                      | ID                   | Languages                                                                                        | Voice Cloning           | Best for                                                                                                                          |
| ------------------------------------------ | -------------------- | ------------------------------------------------------------------------------------------------ | ----------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| Simba 3.2                                  | `simba-3.2`          | English only                                                                                     | Zero-shot               | **Recommended for English.** Streaming-native; lowest TTFB and richest expressivity                                               |
| Simba 3.0                                  | `simba-3.0`          | [6 languages across 7 locales](/build/guides/text-to-speech/language-support#simba-30-languages) | Zero-shot               | **Current API default.** Streaming-native synthesis beyond English                                                                |
| Simba Multilingual *(upgraded 2026-11-21)* | `simba-multilingual` | [30+ languages](/build/guides/text-to-speech/language-support)                                   | Zero-shot + fine-tuning | Retired from API version `2026-09-21`, [served by our current multilingual model from 2026-11-21](#simba-16-and-your-api-version) |
| Simba English *(upgraded 2026-11-21)*      | `simba-english`      | English only                                                                                     | Zero-shot + fine-tuning | Retired from API version `2026-09-21`, [served by our current models from 2026-11-21](#simba-16-and-your-api-version)             |

Pass the model ID as the `model` parameter in your API calls. If omitted, the API defaults to `simba-3.0`; we recommend explicitly setting `model: "simba-3.2"` on English-only integrations for the lowest TTFB and richest expressivity.

What changed in Simba 3.2, which voices it serves and how to move an existing `simba-3.0` integration over is the [Simba 3.2 announcement](https://speechify.ai/blog/simba-3-2-and-the-models-endpoint); this page stays the reference for what is live.

## Simba 1.6 and your API version

`simba-english` and `simba-multilingual` - the Simba 1.6 pair - are being withdrawn in two steps:

| Date                         | What happens                                                                                                                                                                                                                                                                                                                        |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **API version `2026-09-21`** | No longer selectable. Naming either returns `400` `model_retired`, and neither appears in `GET /v1/audio/models`.                                                                                                                                                                                                                   |
| **2026-11-21**               | Both ids are served by our current models instead of their Simba 1.6 training, on every API version that can still name them. `simba-multilingual` is served by our current multilingual model; `simba-english` by `simba-3.2` for English and by the multilingual model for any other language. Nothing breaks; the audio changes. |

**If you use either model, nothing breaks on 2026-11-21, but what you hear changes.** Moving to a current model yourself before then lets you choose the model and test the voice first: [Migrating off Simba 1.6](/build/migrating-off-simba-1-6) is the step-by-step guide. Pinning your workspace's API version to a date before `2026-09-21` keeps the Simba 1.6 training in the meantime, with no code change - see [API Versioning](/build/guides/concepts/api-versioning) for how to read and set it. The pin holds the old training until 2026-11-21 and no further.

New workspaces are pinned to the current version at signup, so a new integration cannot select the pair at all. Use `simba-3.2` for English and `simba-3.0` for the languages it covers.

**Check which languages you actually send before you plan anything.** Most workspaces on `simba-multilingual` synthesize only English, German, Spanish, French, Italian or Portuguese - [all covered by `simba-3.0`](/build/guides/text-to-speech/language-support#simba-30-languages), streaming-native and lower latency. If that is you, migrating is a one-line change.

If you do synthesize outside those six languages, broader multilingual coverage is coming on the current model generation before 2026-11-21, and it is the model `simba-multilingual` is served by from that date. [Talk to us](https://speechify.ai/talk-to-sales) if you want to test it earlier.

> **Note**
>
> While your workspace is pinned below `2026-09-21`, both models still appear in `GET /v1/audio/models` carrying `retired_at: "2026-09-21"` and `sunset_at: "2026-11-21"`. Read `sunset_at` for the day the model behind the id changes.

### Request

POST [https://api.speechify.ai/v1/audio/speech](https://api.speechify.ai/v1/audio/speech)

```curl
curl -X POST https://api.speechify.ai/v1/audio/speech \
     -H "Authorization: Bearer <token>" \
     -H "Content-Type: application/json" \
     -d '{
  "input": "Hello! This is the Speechify text-to-speech API.",
  "voice_id": "geffen_32",
  "audio_format": "mp3",
  "model": "simba-3.2"
}'
```

```typescript
import { SpeechifyClient } from "@speechify/api";

async function main() {
    const client = new SpeechifyClient({
        token: "YOUR_TOKEN_HERE",
    });
    await client.audio.speech({
        audioFormat: "mp3",
        input: "Hello! This is the Speechify text-to-speech API.",
        model: "simba-3.2",
        voiceId: "geffen_32",
    });
}
main();

```

```python
from speechify import Speechify

client = Speechify(
    token="YOUR_TOKEN_HERE",
)

client.audio.speech(
    audio_format="mp3",
    input="Hello! This is the Speechify text-to-speech API.",
    model="simba-3.2",
    voice_id="geffen_32",
)

```

## Listing models via the API

Fetch the current set of selectable models at runtime instead of hardcoding the table above. `GET /v1/audio/models` returns, in its `models` array, every model ID you can pass as the `model` parameter on the synthesis endpoints. It marks the default (the model used when `model` is omitted) and the recommended model, and describes each one. Drive a model picker from this response so it stays current as models are added or the recommendation changes.

### Request

GET [https://api.speechify.ai/v1/audio/models](https://api.speechify.ai/v1/audio/models)

```curl
curl https://api.speechify.ai/v1/audio/models \
     -H "Authorization: Bearer <token>"
```

```typescript
import { SpeechifyClient } from "@speechify/api";

async function main() {
    const client = new SpeechifyClient({
        token: "YOUR_TOKEN_HERE",
    });
    await client.models.list();
}
main();

```

```python
from speechify import Speechify

client = Speechify(
    token="YOUR_TOKEN_HERE",
)

client.models.list()

```

### Response (200)

```json
{
  "models": [
    {
      "id": "simba-3.0",
      "name": "Simba 3.0",
      "default": true,
      "recommended": false,
      "deprecated": false,
      "description": "Streaming-native model serving both English and multilingual synthesis under one id: English, German, Spanish, French, Italian and Portuguese, routed by the request `language`. The default when a request omits `model`.",
      "languages": [
        "en",
        "de-DE",
        "es-ES",
        "es-MX",
        "fr-FR",
        "it-IT",
        "pt-BR"
      ],
      "endpoints": [
        "/v1/audio/speech",
        "/v1/audio/stream",
        "/v1/audio/stream/with-timestamps"
      ],
      "english_voices_only": false,
      "curated_voices": false
    },
    {
      "id": "simba-3.2",
      "name": "Simba 3.2",
      "default": false,
      "recommended": true,
      "deprecated": false,
      "description": "Streaming-native model with the lowest time-to-first-byte and richest expressivity, English only today. Serves every English voice in the catalog, your workspace's own cloned voices included. Use `simba-3.0` for any other language.",
      "languages": [
        "en"
      ],
      "endpoints": [
        "/v1/audio/speech",
        "/v1/audio/stream",
        "/v1/audio/stream/with-timestamps"
      ],
      "english_voices_only": true,
      "curated_voices": false
    }
  ],
  "dialogue_models": [
    {
      "id": "simba-dialogue-1.0",
      "name": "Simba Dialogue 1.0",
      "default": true,
      "recommended": false,
      "deprecated": false,
      "description": "Multi-speaker model that renders a speaker-attributed script as one conversation with natural turn-taking.",
      "languages": [
        "en"
      ],
      "endpoints": [
        "/v1/audio/dialogue"
      ],
      "english_voices_only": true,
      "curated_voices": false
    }
  ]
}
```

Each entry carries the model `id`, a human-readable `name` and `description`, a `default` flag, a `recommended` flag (the model we suggest for new integrations, which is distinct from the `default` - the default accepts every voice in every supported language, while the recommended model may be English-only), a `deprecated` flag, the `languages` it can synthesize (BCP-47 locale strings matching the `language` parameter), the `endpoints` it is valid on, and the `curated_voices` and `english_voices_only` flags. `deprecated` marks a legacy model - a cue to de-emphasise it in a picker and steer new integrations elsewhere. A model being withdrawn carries two more fields: `retired_at`, the API version at which it stops being selectable (present only while your workspace is pinned below it, because at or after it the model is absent from this response entirely), and `sunset_at`, the date its own training stops serving the id on every version. Read them together - `retired_at` is what a pin defers, `sunset_at` is when that stops working. **The list is always exactly what your workspace can call**, so a picker driven off it never offers a model your synthesis request would reject. These values reflect current support and can change over time - a model may gain languages, for example - so read them at runtime rather than caching them: because the response shape is stable and clients ignore unknown fields, both changing values and future new fields are picked up without breaking existing integrations.

Voice cloning is supported by both Simba 3 models and requires a paid plan; see [Voice Cloning](/build/guides/voice-cloning/overview).

## Simba 3.2

Streaming-native flagship model with the lowest TTFB (time to first byte) and richest expressivity. Recommended for new English integrations.

* Optimized for real-time streaming with the lowest startup latency
* Richer expressive range than earlier Simba generations
* Full support for [SSML](/build/guides/text-to-speech/ssml) and [emotion control](/build/guides/text-to-speech/emotion-control)
* Serves every English voice in the catalog, your workspace's own clones included
* English only; a non-English voice returns `400`. Use `simba-3.0` for the other supported languages
* Zero-shot voice cloning is supported: your workspace's own cloned voices work here like any catalog voice — see [Voice Cloning](/build/guides/voice-cloning/overview)

## Simba 3.0

Streaming-native model covering six languages across seven locales: English, German, Spanish (Spain and Mexico), French, Italian, and Brazilian Portuguese. The model a request resolves to when it omits `model`.

* Officially supports `en-*`, `de-DE`, `es-ES`, `es-MX`, `fr-FR`, `it-IT` and `pt-BR`. See [Language Support](/build/guides/text-to-speech/language-support#simba-30-languages)
* Set the `language` parameter to pick the language; when omitted, the voice's own locale decides
* English and non-English are served by two separate trainings, but that routing is internal: the model ID you pass is always `simba-3.0`
* Languages outside the supported set often work but are not validated; `simba-multilingual` covers the full 30+ locale set on a [pinned API version](#simba-16-and-your-api-version), and from 2026-11-21 is served by our current multilingual model
* Prefer `simba-3.2` for English-only integrations; it has lower TTFB and richer expressivity
* Full support for [SSML](/build/guides/text-to-speech/ssml) and [emotion control](/build/guides/text-to-speech/emotion-control)
* Zero-shot voice cloning works self-serve, in every language the model supports - see [Voice Cloning](/build/guides/voice-cloning/overview)

## Simba Multilingual

> **Note**
>
> Legacy Simba 1.6 model. **Retired from API version `2026-09-21`; from 2026-11-21 served by our current models.** A workspace pinned before the retirement keeps the Simba 1.6 training until then - see [Simba 1.6 and your API version](#simba-16-and-your-api-version). Migrate to `simba-3.0` for the languages it covers.

Supports multiple languages, including mixing languages within a single sentence.

* 31 locales covering 30 distinct languages live today
* Automatic language detection when the `language` parameter is omitted
* Zero-shot voice cloning works across all supported languages
* Fine-tuned voice cloning available (contact sales)

See [Language Support](/build/guides/text-to-speech/language-support) for the full list.

## Simba English

> **Note**
>
> Legacy Simba 1.6 model. **Retired from API version `2026-09-21`; from 2026-11-21 served by our current models.** A workspace pinned before the retirement keeps the Simba 1.6 training until then - see [Simba 1.6 and your API version](#simba-16-and-your-api-version). Migrate to `simba-3.2`.

Kept for integrations that name it explicitly.

* Full support for [SSML](/build/guides/text-to-speech/ssml) and [emotion control](/build/guides/text-to-speech/emotion-control)
* Zero-shot voice cloning from short audio samples
* Fine-tuned voice cloning from hours of speaker audio (contact sales)
* Prefer `simba-3.2` for new integrations using the built-in voice catalog; a new workspace cannot select this model at all

## Voice cloning

Simba English and Simba Multilingual support two tiers of voice cloning; Simba 3.0 and Simba 3.2 both support zero-shot cloning self-serve:

| Tier       | Input                     | Quality | Availability                  |
| ---------- | ------------------------- | ------- | ----------------------------- |
| Zero-shot  | 10-30 second audio sample | Good    | Self-serve via API or Console |
| Fine-tuned | Hours of speaker audio    | Best    | Contact sales                 |

A cloned voice works on `simba-3.0` with no approval step, and on `simba-english` / `simba-multilingual` where your [API version](#simba-16-and-your-api-version) still offers them; from 2026-11-21 those ids are served by our current models, which take cloned voices too. `simba-3.2` cloning is self-serve for every workspace, with no enablement step: an English clone you own works there, and you pass the voice ID exactly as you would a stock voice. (`simba-3.2` is English-only, so a non-English clone returns `400` — use `simba-3.0` for those.)

See [Voice Cloning](/build/guides/voice-cloning/overview) for implementation details.

## FAQ

#### What TTS models are live today?

Speechify currently offers two streaming-native Simba 3 text-to-speech models: **Simba 3.2** (`simba-3.2`), recommended for English integrations, and **Simba 3.0** (`simba-3.0`), the API default (used when you omit `model`), covering six languages across seven locales. Both support voice cloning on paid plans. The legacy Simba 1.6 pair (`simba-english`, `simba-multilingual`) is [retired from API version `2026-09-21` and served by current models from 2026-11-21](#simba-16-and-your-api-version). For the authoritative, always-current list of every model your workspace can call, use `GET /v1/audio/models`.

#### Which model should I use?

Use **Simba 3.2** for most English use cases - it has the lowest startup latency and richest expressivity, and is the recommended Simba 3 model. It takes cloned voices too, self-serve, as every other model does. Use **Simba 3.0** for streaming-native synthesis in German, Spanish, French, Italian or Brazilian Portuguese, and **Simba Multilingual** for the full 30+ locale set or mixed-language content - that one needs an [API version pinned before `2026-09-21`](#simba-16-and-your-api-version) and is served by our current multilingual model from 2026-11-21. Note: the API defaults to **Simba 3.0** when `model` is omitted, so set `model: "simba-3.2"` explicitly to opt in.

#### Can I switch models without changing my code?

Yes. Just change the `model` parameter. All other parameters (voice, format, SSML) work the same across models. The one thing to check is language: Simba 3.2 is English-only, so a non-English voice returns `400` there - use Simba 3.0 for those.

#### Do all models support the same voices?

Yes, including your own cloned voices, with one language caveat: Simba 3.2 is English-only, so it takes every English voice in the catalog and returns `400` for a non-English one. Simba 3.0, Simba English and Simba Multilingual serve the full catalog in every language they support. Each voice's `models` array in `GET /v1/voices` is the authoritative per-voice answer.