> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt. > > Canonical Speechify URLs — use exactly, do not invent variants: > - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts) > - https://speechify.ai — marketing + product site > - https://platform.speechify.ai — customer dashboard, signup, API keys, billing > - https://api.speechify.ai — API base URL > - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist. > - https://status.speechify.ai — status + incidents > - https://speechify.com — SEPARATE consumer reader app, NOT this API > > `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired. > > Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp # Speech marks > Speech marks from the Speechify API map audio timing to text, enabling word highlighting, seeking, and synchronization. See the data structure and key gotchas. ## Overview Speech marks provide a mapping between time and text. They inform the client when each word is spoken in the audio, enabling features like: * Text highlighting during playback * Precise audio seeking by text position * Usage tracking and analytics * Synchronization between text and audio Two endpoints return them: | Endpoint | Delivery | Models | | --------------------------------------- | ----------------------------------------------------- | ------ | | `POST /v1/audio/speech` | One JSON response after the whole synthesis completes | All | | `POST /v1/audio/stream/with-timestamps` | Server-Sent Events, as the audio is generated | All | Use the streaming endpoint when you want highlighting to start before the synthesis has finished. Use the batch endpoint when you want a single response. `POST /v1/audio/stream` returns raw audio only and is unchanged. ## Data structure Speech marks use the following TypeScript interfaces: ```ts type NestedChunk = { start_time: number // Time in milliseconds when this chunk starts in the audio end_time: number // Time in milliseconds when this chunk ends in the audio start: number // Character index where this chunk starts in the original text end: number // Character index where this chunk ends in the original text value: string // The text content of this chunk } type SpeechMarks = NestedChunk & { chunks: NestedChunk[] // Array of word-level chunks within this sentence/paragraph } ``` ## Important considerations * **Offsets index your text, not the markup**: `start` and `end` count characters of your input's text content, with SSML tags removed and entities decoded, and `value` is that text. `&` counts as the one character `&`. To map a mark back into raw SSML, track positions while you strip the markup; the [string tracker](https://github.com/SpeechifyInc/string-tracker) library helps. * **Substituted text**: A `` element is spoken as its alias but marked on its own text. The alias's words become one mark whose `value`, `start` and `end` cover the element's text, timed across the whole alias, and every later mark keeps its position in your text. * **Index gaps**: The `start` and `end` values of each word may have gaps. When looking for a word at a specific index, check for `start` being `>= yourIndex` rather than checking if the index is within both `start` and `end` bounds. * **Timing gaps**: Similarly, `start_time` and `end_time` of each word may have gaps. Follow the same approach as with index gaps. * **Initial silence**: The `start_time` of the first word is not necessarily `0` like the `SpeechMarks`. There can be silence at the beginning of the sentence that leads to the word starting partway through. * **Trailing silence**: The `end_time` of the last word does not necessarily correspond with the end of the `SpeechMarks`. There can be silence at the end that will make the `SpeechMarks` longer. ## Example output For the input `"Hello, welcome to Speechify"`, the response includes: ```ts const chunk: SpeechMarks = { start: 0, end: 27, start_time: 0, end_time: 1850, value: 'Hello, welcome to Speechify', chunks: [ { start: 0, end: 6, start_time: 125, end_time: 375, value: 'Hello,' }, { start: 7, end: 14, start_time: 375, end_time: 750, value: 'welcome' }, { start: 15, end: 17, start_time: 750, end_time: 875, value: 'to' }, { start: 18, end: 27, start_time: 875, end_time: 1850, value: 'Speechify' }, ], } ``` Note how `start_time` of the first word (125ms) doesn't match the `SpeechMarks` start (0ms) - there's initial silence before speech begins. ## Streaming speech marks `POST /v1/audio/stream/with-timestamps` takes the same request body as `POST /v1/audio/stream` and returns a Server-Sent Events stream instead of raw audio bytes. ```bash curl -N -X POST https://api.speechify.ai/v1/audio/stream/with-timestamps \ -H "Authorization: Bearer $SPEECHIFY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input": "Hello world. This is a speech mark test.", "voice_id": "geffen_32", "model": "simba-3.2", "output_format": "mp3_24000_64" }' ``` Each event is named. `speech.chunk` carries a Base64-encoded run of audio, the speech marks that became final with it, or both: ``` event: speech.chunk data: {"audio":"SUQzBAAAAAAA...","speech_marks":[{"type":"word","start":0,"end":5,"start_time":0,"end_time":469,"value":"Hello"}]} event: speech.chunk data: {"speech_marks":[{"type":"word","start":6,"end":12,"start_time":469,"end_time":1152,"value":"world."}]} event: speech.done data: {"billable_characters_count":40,"audio_duration_ms":4350} ``` The marks use the same fields as the batch endpoint, so one parser handles both. ### Consuming the stream * **Concatenate the audio.** Base64-decode each `audio` field and append it. The result is byte-for-byte what `POST /v1/audio/stream` would have returned. * **Apply marks against that one timeline.** `start_time` and `end_time` are absolute milliseconds from the start of the synthesis. Which event a mark arrives on is a delivery detail and carries no meaning, so do not try to align a mark to the audio chunk it arrived with. * **A chunk may carry only one of the two.** Marks lag their audio slightly, and the last chunk of a stream is often marks-only. The legacy `simba-english` and `simba-multilingual` models synthesize a sentence at a time, so their marks arrive a sentence at a time too. * **Ignore unknown event types.** New event types may be added; a client that skips names it does not recognize keeps working. * **There is no `[DONE]` sentinel.** `speech.done` is the terminal event. ### Errors Before the stream starts, failures are ordinary HTTP responses with the standard error envelope. Every model streams speech marks; if a model's serving cannot produce them, the request fails here with `400 speech_marks_unsupported`, and `POST /v1/audio/speech`, which returns speech marks for every model, is the fallback. Once the stream has started the status code is already committed, so a later failure arrives as a terminal event carrying the same envelope: ``` event: speech.error data: {"error":{"code":"upstream_failure","message":"synthesis failed mid-stream"},"request_id":"..."} ``` A [content policy](/build/guides/concepts/content-policy) refusal that lands after the stream has started ends it the same way, with code `content_policy_violation`. Edit the input rather than retrying. ### Audio format `output_format` and the `Accept` header select the codec exactly as they do on `POST /v1/audio/stream`. Because the response's own `Content-Type` is `text/event-stream`, the media type of the audio inside the events is echoed on the `Speechify-Audio-Content-Type` response header. Changing the codec or sample rate does not change the duration, so mark times stay correct for every `output_format`. > Learn how speech marks map text to audio timing for synchronization features.