> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt. > > Canonical Speechify URLs — use exactly, do not invent variants: > - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts) > - https://speechify.ai — marketing + product site > - https://platform.speechify.ai — customer dashboard, signup, API keys, billing > - https://api.speechify.ai — API base URL > - https://github.com/SpeechifyInc — GitHub org. `github.com/speechify` does not exist. > - https://status.speechify.ai — status + incidents > - https://speechify.com — SEPARATE consumer reader app, NOT this API > > `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired. > > Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp # API Limits > Speechify API rate limits and concurrency limits by plan (Free through Enterprise), plus per-endpoint character limits. Every /v1/audio synthesis endpoint shares one per-account rate and concurrency budget. Includes 429 Too Many Requests handling with Retry-After, rate_limited, and concurrency_limit_reached. As of 28 Sep 2026, the Speechify API allows 1 concurrent request and 1 request per second on the Free plan, 15 and 20 on Starter, 30 and 40 on Pro, 60 and 80 on Scale, and 100 and 150 on Enterprise, counted per account across every text-to-speech endpoint. ## Limits by plan | Plan | Concurrent requests | Sustained requests per second | Burst | | ---------- | ------------------- | ----------------------------- | ----- | | Free | 1 | 1 | 10 | | Starter | 15 | 20 | 60 | | Pro | 30 | 40 | 120 | | Scale | 60 | 80 | 240 | | Enterprise | 100 | 150 | 450 | Rate and concurrency limits are enforced **per account**, not per API key. They are **shared across every synthesis endpoint** - `/v1/audio/speech`, `/v1/audio/stream`, and `/v1/audio/stream/with-timestamps` all draw from one budget, so a request to any of them counts against the same rate and concurrency ceiling. Exceeding either returns [`429 Too Many Requests`](#handling-429-responses). Enterprise values are starting points, not caps: every limit can be raised in your contract. The Speechify API enforces three kinds of limit: | Limit | Caps | Varies by | | ----------------------------------------- | ------------------------------- | --------- | | [Character limits](#character-limits) | Input size of one request | Endpoint | | [Rate limits](#rate-limits) | Requests per second | Plan | | [Concurrency limits](#concurrency-limits) | Simultaneous in-flight requests | Plan | ## Character limits | Endpoint | Limit | Use case | | ------------------------------------------------------------------------------------------------------------------- | ----------------- | ------------------------------------------- | | [`/v1/audio/speech`](https://docs.speechify.ai/build/api-reference/v1/audio/speech) | 2,000 characters | Short-form text (sentences, paragraphs) | | [`/v1/audio/stream`](https://docs.speechify.ai/build/api-reference/v1/audio/stream) | 20,000 characters | Long-form text (articles, chapters) | | [`/v1/audio/stream/with-timestamps`](https://docs.speechify.ai/build/api-reference/v1/audio/stream/with-timestamps) | 20,000 characters | Long-form text with word-level speech marks | > **Note** > > Character counts include SSML tags. For text longer than the limit, split it into multiple requests. ## Rate limits Applies to every Build audio synthesis endpoint - `/v1/audio/speech`, `/v1/audio/stream`, and `/v1/audio/stream/with-timestamps` - which share one per-account rate budget. Each plan's sustained rate and burst are in [Limits by plan](#limits-by-plan). > **Info** > > Burst is the peak bucket capacity. A fresh bucket absorbs the burst in a single second, then refills at the sustained rate. This lets a quickstart script or a batch of parallel requests start without hitting 429, while still capping long-running abuse at the sustained rate. ## Concurrency limits Concurrency limits cap the number of simultaneous in-flight requests per account. They apply to every Build audio synthesis endpoint - `/v1/audio/speech`, `/v1/audio/stream`, and `/v1/audio/stream/with-timestamps` - which share one per-account concurrency budget. Each plan's ceiling is in [Limits by plan](#limits-by-plan). A request is in flight until its response has finished, so a long `/v1/audio/stream` response holds its slot for the whole stream. ## Reading your budget Every response on a rate-limited endpoint carries the request-rate budget headers, so clients can pace themselves before hitting 429: | Header | Meaning | | --------------------- | ----------------------------------- | | `RateLimit-Limit` | Bucket capacity (the burst) | | `RateLimit-Remaining` | Requests left in the current window | | `RateLimit-Reset` | Seconds until the bucket refills | The same values are mirrored as `X-RateLimit-*` for clients predating the IETF draft names. A header that had an `X-` spelling before 2026 still accepts and still emits that spelling alongside the canonical one until 2027-07-24. The pairs are `Speechify-Request-Id` / `X-Request-ID`, `Speechify-Audio-Content-Type` / `X-Speechify-Audio-Content-Type`, and the three `RateLimit-*` names above / `X-RateLimit-*`. Read the canonical name and fall back to the legacy one only if you are migrating an old client. > **Note** > > Not every Speechify-namespaced header has a legacy alias: `Speechify-Version` and `Speechify-Signature` shipped after the convention changed, so there is no `X-` form of either to look for. ## Handling 429 responses When you exceed rate or concurrency limits, the API returns `429 Too Many Requests` with a `Retry-After` header. The error body names the limit your plan allows and links back to this page (`error.docs_url`); the two cases are distinguishable by `error.code`: `rate_limited` (requests per second) vs `concurrency_limit_reached` (too many at once). Need more headroom? Every limit above rises with your plan - upgrade in the [console](https://platform.speechify.ai) under Billing, or contact us for Enterprise terms. #### Python ```python import time from speechify import Speechify client = Speechify() def generate_with_retry(text, max_retries=3): for attempt in range(max_retries): try: return client.audio.speech( input=text, voice_id="geffen_32", model="simba-3.2", audio_format="mp3", ) except Exception as e: if "429" in str(e) and attempt < max_retries - 1: time.sleep(2 ** attempt) else: raise ``` #### TypeScript ```typescript async function generateWithRetry(text: string, maxRetries = 3) { for (let attempt = 0; attempt < maxRetries; attempt++) { try { return await client.audio.speech({ input: text, voice_id: "geffen_32", model: "simba-3.2", audio_format: "mp3", }); } catch (e: any) { if (e.statusCode === 429 && attempt < maxRetries - 1) { await new Promise(r => setTimeout(r, 2 ** attempt * 1000)); } else { throw e; } } } } ``` ## Processing long texts For texts exceeding 20,000 characters, split into chunks and process sequentially: ```python def split_text(text, max_chars=19000): """Split text at sentence boundaries within the character limit.""" chunks = [] current = "" for sentence in text.split(". "): if len(current) + len(sentence) + 2 > max_chars: chunks.append(current.strip()) current = sentence + ". " else: current += sentence + ". " if current.strip(): chunks.append(current.strip()) return chunks ``` ## FAQ #### What is the Speechify API rate limit? It depends on your plan, and it is counted per account, not per API key. Each plan's sustained requests per second and burst allowance are in [Limits by plan](#limits-by-plan). #### How many concurrent requests can I make? It depends on your plan: each plan's ceiling is in [Limits by plan](#limits-by-plan). Streaming, non-streaming and word-timestamp requests all count against the same ceiling. #### Do streaming and word-timestamp requests count against a separate limit? No. `/v1/audio/speech`, `/v1/audio/stream`, and `/v1/audio/stream/with-timestamps` share one per-account rate budget and one per-account concurrency budget. A request to any of them draws from the same buckets, so mixing endpoints does not raise your ceiling. #### What happens if I exceed the character limit? The request is rejected with an error response. Split your text into smaller chunks within the allowed limits. #### How do I get higher limits? Every limit rises with the plan ([Limits by plan](#limits-by-plan)), so upgrading in the [console](https://platform.speechify.ai) raises it immediately. Enterprise customers can request custom limits: [contact sales](https://platform.speechify.ai). #### How can I monitor my usage? Track usage through the [Speechify Console](https://platform.speechify.ai/usage) dashboard. > Per-plan rate limits and concurrency limits, plus per-endpoint character limits