> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt. > > Canonical Speechify URLs — use exactly, do not invent variants: > - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts) > - https://speechify.ai — marketing + product site > - https://platform.speechify.ai — customer dashboard, signup, API keys, billing > - https://api.speechify.ai — API base URL > - https://github.com/SpeechifyInc — GitHub org. `github.com/speechify` does not exist. > - https://status.speechify.ai — status + incidents > - https://speechify.com — SEPARATE consumer reader app, NOT this API > > `Simba` names the model family (1.6 multilingual, 3.0 streaming multilingual, 3.2 streaming English), not the brand. `SimbaVoice` / `simbavoice.ai` are retired. > > Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp # Create Speech POST https://api.speechify.ai/v1/audio/speech Content-Type: application/json Synthesize speech audio from text or SSML. Returns the complete audio file plus billing and speech-mark metadata in a single JSON response. For low-latency playback or long-form text, use POST /v1/audio/stream. Set `output_format` for explicit sample-rate/bitrate control (e.g. `pcm_16000` or `ulaw_8000` for telephony). Reference: https://docs.speechify.ai/build/api-reference/v1/audio/speech ## Authentication - `Authorization` header (bearer token, required) — Enter your API key with the `Bearer` prefix, e.g. 'Bearer sk_...'. ## Request ### Body (application/json) This endpoint expects a GetSpeechRequest. - `input` (string, required) — Plain text or SSML to be synthesized to speech. Refer to https://docs.speechify.ai/docs/api-limits for the input size limits. Emotion, Pitch and Speed Rate are configured in the ssml input, please refer to the ssml documentation for more information: https://docs.speechify.ai/docs/ssml#prosody - `voice_id` (string, required) — Id of the voice to be used for synthesizing speech. Refer to /v1/voices endpoint for available voices - `audio_format` (enum, optional, default: wav) — The format for the output audio. Note, that the current default is "wav", but there's no guarantee it will not change in the future. We recommend always passing the specific param you expect. - Allowed values: `wav`, `mp3`, `ogg`, `aac`, `pcm` - `language` (string, optional) — Language of the input. Follow the format of an ISO 639-1 language code and an ISO 3166-1 region code, separated by a hyphen, e.g. en-US. Please refer to the list of the supported languages and recommendations regarding this parameter: https://docs.speechify.ai/docs/language-support. - `model` (enum, optional, default: simba-3.0) — Model used for audio synthesis. Defaults to `simba-3.0`, which is streaming-native and multilingual: it officially supports English plus `de-DE`, `es-ES`, `es-MX`, `fr-FR`, `it-IT` and `pt-BR`, and routes each request to its English or its multilingual training based on `language` (falling back to the voice's locale when `language` is omitted). `simba-3.2` is the streaming-native model with the lowest TTFB and richest expressivity, and the recommended Simba 3 model; it is English only, so a non-English voice returns 400. The legacy Simba 1.6 models `simba-english` and `simba-multilingual` are retired from API version `2026-09-21`: naming one returns 400 `model_retired`. Pinning your API version to a date before `2026-09-21` keeps them on their Simba 1.6 training until **2026-11-21**; from then both ids are served by our current models on every API version that can still name them. Migrate to `simba-3.2` (English) or `simba-3.0` before then; call GET /v1/audio/models to see the set your workspace can select today. - Allowed values: `simba-3.0`, `simba-3.2` - `options` (GetSpeechOptionsRequest, optional) — GetSpeechOptionsRequest is the wrapper for request parameters to the client - `output_format` (enum, optional) — The output audio format as a `codec_sampleRate_bitrate` string. Takes precedence over `audio_format` when set. - Allowed values: `pcm_8000`, `pcm_16000`, `pcm_22050`, `pcm_24000`, `pcm_44100`, `pcm_48000`, `mp3_22050_32`, `mp3_22050_64`, `mp3_22050_96`, `mp3_22050_128`, `mp3_22050_160`, `mp3_22050_192`, `mp3_24000_32`, `mp3_24000_64`, `mp3_24000_96`, `mp3_24000_128`, `mp3_24000_160`, `mp3_24000_192`, `wav_24000`, `wav_48000`, `ulaw_8000`, `ogg_24000`, `aac_24000` ## Response ### 200 Synthesized speech audio for the requested input. - `audio_data` (string, required) — Synthesized speech audio, Base64-encoded - `audio_format` (enum, required) — The codec of the audio data - Allowed values: `wav`, `mp3`, `ogg`, `aac`, `pcm`, `ulaw` - `billable_characters_count` (long, required) — The number of billable characters processed in the request. - `speech_marks` (SpeechMarks, required) — It is used to annotate the audio data with metadata about the synthesis process, like word timing or phoneme details. - `output_format` (enum, optional) — The full `codec_sampleRate_bitrate` format the audio was encoded in, returned when the request set `output_format`. It is the requested value unless the request named a bitrate above the mp3 ceiling, in which case it reports the bitrate actually delivered. - Allowed values: `pcm_8000`, `pcm_16000`, `pcm_22050`, `pcm_24000`, `pcm_44100`, `pcm_48000`, `mp3_22050_32`, `mp3_22050_64`, `mp3_22050_96`, `mp3_22050_128`, `mp3_22050_160`, `mp3_22050_192`, `mp3_24000_32`, `mp3_24000_64`, `mp3_24000_96`, `mp3_24000_128`, `mp3_24000_160`, `mp3_24000_192`, `wav_24000`, `wav_48000`, `ulaw_8000`, `ogg_24000`, `aac_24000` ## Errors ### 400 Bad Request Error The request was malformed or failed validation. The response body is the standard `Error` envelope; for validation failures `error.fields` enumerates the offending fields as a `path -> message` map (code = `validation_failed`). - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 401 Unauthorized Error Authentication is missing or invalid. The request did not carry a recognised credential (console session token, API key, or worker JWT). - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 402 Payment Required Error The workspace has insufficient credits, or the request needs a plan tier the workspace is not on (e.g. voice cloning). Distinct from `Forbidden` so SDK consumers can drive upgrade UX. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 403 Forbidden Error The credential authenticated, but is not authorised for this resource - typically a workspace-role gate (owner / admin required) or a cross-tenant access attempt. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 404 Not Found Error The referenced resource does not exist or is not visible to the caller's workspace. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 409 Conflict Error The request conflicts with the current resource state - e.g. duplicate, optimistic-concurrency mismatch, or last-owner guard. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 429 Too Many Requests Error Rate limit or concurrency limit exceeded. `error.code` says which ceiling, and they need different responses: `rate_limited` is the request-rate budget (slow down), `concurrency_limit_reached` is a workspace-wide concurrency ceiling (fewer at once, or raise it), and `conversation_turn_in_progress` is contention over one named conversation (keep one message in flight on it). Every 429 carries `Retry-After` and the request-rate budget headers; the active-call cap also carries `RateLimit-Remaining-Calls: 0`. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 500 Internal Server Error An unexpected server-side error occurred. Safe to retry with exponential backoff for idempotent requests. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 502 Bad Gateway Error An upstream dependency (the TTS composer or voice-metadata service) returned a 5xx. The raw upstream detail is not forwarded - the cause is in the server log; the response is a fixed `upstream_failure` envelope. Safe to retry. - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ### 503 Service Unavailable Error A downstream dependency is degraded or the endpoint is intentionally disabled (e.g. phone-number purchase before ops setup). - `error` (ErrorDetail, required) - `request_id` (string, optional) — Server-side request identifier. Echoes the `Speechify-Request-Id` response header. Stable across the request's lifetime, written to structured logs, and useful when reporting issues. ## Types ### GetSpeechOptionsRequest GetSpeechOptionsRequest is the wrapper for request parameters to the client - `loudness_normalization` (boolean, optional, default: false) — Determines whether to normalize the audio loudness to a standard level. When enabled, loudness normalization aligns the audio output to the following standards: Integrated loudness: -14 LUFS True peak: -2 dBTP Loudness range: 7 LU If disabled, the audio loudness will match the original loudness of the selected voice, which may vary significantly and be either too quiet or too loud. Enabling loudness normalization can increase latency due to additional processing required for audio level adjustments. - `text_normalization` (boolean, optional, default: true) — Determines whether to normalize the text. If enabled, it will transform numbers, dates, etc. into words. For example, "55" is normalized into "fifty five". This can increase latency due to additional processing required for text normalization. ### SpeechMarks It is used to annotate the audio data with metadata about the synthesis process, like word timing or phoneme details. - `chunks` (list of NestedChunk, required) — Array of NestedChunk, each providing detailed segment information within the synthesized speech. - `end` (long, required) - `end_time` (double, required) - `start` (long, required) - `start_time` (double, required) - `type` (string, required) - `value` (string, optional) ### ErrorDetail - `code` (enum, required) — Stable machine-readable error code. Additive only: codes are never renamed, only deprecated. SDKs may map each code to a typed exception class. Status-code semantics: 4xx codes describe caller-fixable issues; 5xx codes describe server-side failures and are safe to retry with backoff for idempotent requests. - Allowed values: `endpoint_moved`, `bad_request`, `validation_failed`, `unauthorized`, `payment_required`, `forbidden`, `not_found`, `method_not_allowed`, `conflict`, `idempotency_conflict`, `payload_too_large`, `unsupported_media_type`, `rate_limited`, `concurrency_limit_reached`, `invalid_api_version`, `internal_error`, `upstream_failure`, `service_unavailable`, `caller_not_found`, `contact_not_found`, `contact_identifier_not_found`, `contact_identifier_conflict`, `contact_resolver_not_found`, `credential_not_found`, `credential_in_use`, `agent_not_found`, `agent_in_use`, `agent_run_not_found`, `kb_not_found`, `kb_document_not_found`, `kb_folder_not_found`, `tool_not_found`, `tool_name_taken`, `channel_instance_not_found`, `team_not_found`, `trigger_not_found`, `store_not_found`, `store_document_not_found`, `hosted_api_not_found`, `api_route_not_found`, `consumer_key_not_found`, `skill_not_found`, `skill_version_not_found`, `file_not_found`, `file_path_taken`, `file_storage_limit_reached`, `store_limit_reached`, `store_document_limit_reached`, `store_bytes_limit_reached`, `store_not_configured`, `store_document_version_conflict`, `store_document_deleted`, `entitlement_override_exists`, `hosted_apis_not_in_plan`, `skills_not_in_plan`, `voice_agents_not_in_plan`, `skill_in_use`, `skill_tool_name_conflict`, `skill_limit_reached`, `agent_skill_limit_reached`, `hosted_api_slug_taken`, `api_route_conflict`, `mount_plan_changed`, `route_output_unavailable`, `route_run_timeout`, `route_run_failed`, `route_run_limit_reached`, `route_read_limit_reached`, `route_write_limit_reached`, `hosted_api_public_refused`, `route_tool_not_readable`, `route_tool_unavailable`, `route_upstream_rate_limited`, `route_upstream_error`, `hosted_mcp_not_enabled`, `hosted_api_busy`, `conversation_not_found`, `phone_number_not_found`, `sip_trunk_not_found`, `voice_not_found`, `audio_asset_not_found`, `builtin_not_found`, `batch_not_found`, `agent_test_not_found`, `workspace_not_found`, `invite_not_found`, `project_not_found`, `cross_project_reference`, `project_has_scoped_credentials`, `project_not_empty`, `project_limit_reached`, `agent_limit_reached`, `project_too_large_to_promote`, `call_not_found`, `message_not_found`, `thread_not_found`, `call_not_active`, `relay_displaces_agent`, `brain_not_found`, `brain_in_use`, `custom_model_not_found`, `custom_model_in_use`, `insufficient_scope`, `purchased_numbers_not_included`, `phone_number_quota_reached`, `batch_calls_not_included`, `voice_cloning_not_included`, `consent_challenge_not_found`, `consent_challenge_expired`, `consent_challenge_already_used`, `consent_phrase_mismatch`, `consent_speaker_mismatch`, `consent_recording_unusable`, `consent_verification_unavailable`, `consent_verification_required`, `watermark_audio_unusable`, `watermark_detection_unavailable`, `workspace_last_owner`, `workspace_last_workspace`, `account_deletion_blocked`, `workspace_free_limit`, `workspace_single_owner`, `invite_email_mismatch`, `invite_already_pending`, `service_account_limit_reached`, `service_accounts_not_in_plan`, `speech_marks_unsupported`, `model_retired`, `too_many_voices`, `content_policy_violation`, `topup_not_in_plan`, `credit_purchase_unpaid`, `credit_purchase_payment_in_progress`, `tool_config_shared`, `spend_cap_exceeded`, `spend_budget_exceeded`, `project_spend_limit_exceeded`, `project_archived`, `project_not_archived`, `project_not_purged`, `project_restore_window_expired`, `project_name_taken`, `funded_balance_required`, `agent_publish_gate_failed`, `agent_publish_gate_required`, `agent_publish_gate_unavailable`, `agent_publish_gate_tool_unreachable`, `text_channel_not_in_plan`, `channel_not_in_plan`, `text_turn_failed`, `conversation_turn_in_progress`, `conversation_closed`, `conversation_not_reachable`, `conversation_channel_bound`, `conversation_prompt_not_found`, `text_message_quota_exceeded`, `durable_runs_not_in_plan`, `tool_transport_unsupported`, `agent_config_too_large`, `agent_run_not_pending`, `agent_run_action_stale`, `share_link_not_found`, `share_link_exhausted`, `share_link_limit_reached`, `destination_not_allowed`, `international_dialing_not_enabled`, `number_not_sms_capable`, `verification_required`, `intended_use_required` - `message` (string, required) — Human-readable explanation of this specific occurrence. Safe to surface in UI banners or pass to support. The wording can change between releases; clients should match on `code`, not on the message string. - `fields` (map from string to string, optional) — Per-field validation errors as `path -> message`. Only present on 400 responses caused by request validation (typically code=`validation_failed`). Keys are field paths in dotted/bracket notation; values are short human explanations safe to inline-surface next to the offending form field. - `details` (map from string to any, optional) — Structured, endpoint-specific context beyond the flat `fields` map. Present only on the few errors that carry it (e.g. the `used_by` referrer list on a credential delete-conflict); its shape depends on the error `code`. Clients that don't recognise a `details` shape can ignore it - the `code` + `message` contract is unchanged. - `docs_url` (string, optional) — Link to the documentation that resolves this class of error, when a stable page exists. Rate and concurrency 429s link the API limits reference, which lists each plan's limits and how to raise them. ### NestedChunk It details the type of segment, its start and end points in the text, and its start and end times in the synthesized speech audio. - `end` (long, optional) - `end_time` (double, optional) - `start` (long, optional) - `start_time` (double, optional) - `type` (string, optional) - `value` (string, optional) ## Examples **Request** ```json { "input": "Hello! This is the Speechify text-to-speech API.", "voice_id": "geffen_32", "audio_format": "mp3", "model": "simba-3.2" } ``` **Response** ```json { "audio_data": "example", "audio_format": "wav", "billable_characters_count": 10, "speech_marks": { "chunks": [ {} ], "end": 1, "end_time": 1, "start": 1, "start_time": 1, "type": "example", "value": "example" } } ```