> This page is for Build.

> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Vercel AI SDK

> Use @speechify/ai-sdk-provider, the Speechify-maintained provider for the Vercel AI SDK, to call Speechify text to speech from generateSpeech() with word-level speech marks.

## Overview

[`@speechify/ai-sdk-provider`](https://www.npmjs.com/package/@speechify/ai-sdk-provider) is Speechify's provider for the [Vercel AI SDK](https://ai-sdk.dev).
It implements the AI SDK's speech model interface, so `generateSpeech()` calls [`POST /v1/audio/speech`](/build/api-reference/v1/audio/speech) and returns the audio together with its speech marks.

Speechify maintains the package, and its source is [Speechify-AI/ai-sdk-provider](https://github.com/Speechify-AI/ai-sdk-provider).

## Prerequisites

* Speechify API key: sign up at [platform.speechify.ai](https://platform.speechify.ai/signup?sfy_source=docs\&sfy_campaign=vercel-ai-sdk\&sfy_placement=prerequisites\&sfy_product=build)
* AI SDK 7 (`ai@^7`) on Node.js 22 or later

## Install

```bash
npm install @speechify/ai-sdk-provider ai
```

Put your key in the environment of the server that runs the AI SDK:

```bash
SPEECHIFY_API_KEY=your_speechify_api_key
```

## Generate speech

```ts
import { writeFile } from "node:fs/promises";
import { generateSpeech } from "ai";
import { speechify } from "@speechify/ai-sdk-provider";

const { audio, providerMetadata, warnings } = await generateSpeech({
  model: speechify.speech("simba-3.2"),
  text: "Hello from Speechify and the AI SDK.",
  voice: "harper_32",
});

await writeFile("hello.mp3", audio.uint8Array);
```

`speechify` reads `SPEECHIFY_API_KEY` when the request is made.
To pass a key or a proxy URL explicitly, create your own instance with `createSpeechify({ apiKey, baseURL, headers, fetch })`.

> **Warning**
>
> Run `generateSpeech()` on the server, for example in a route handler. An API key shipped to a browser is readable by anyone who opens the page.

## How AI SDK options map to the API

| `generateSpeech()` option   | Speechify request                 | Notes                                                                                                                                                |
| --------------------------- | --------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| `model`                     | `model`                           | `simba-3.2` for English, `simba-3.0` for the other [supported languages](/build/guides/text-to-speech/language-support).                             |
| `text`                      | `input`                           | Plain text or [SSML](/build/guides/text-to-speech/ssml), up to 2,000 characters.                                                                     |
| `voice`                     | `voice_id`                        | Defaults to `harper_32`. Any voice whose `models` list in [`GET /v1/voices`](/build/api-reference/v1/voices/get) includes the model.                 |
| `outputFormat`              | `audio_format` or `output_format` | Defaults to `mp3`. Accepts `mp3`, `wav`, `ogg`, `aac`, `pcm`, `mulaw` and every `codec_sampleRate_bitrate` value such as `pcm_16000` or `ulaw_8000`. |
| `speed`                     | SSML `<prosody rate>`             | 0.5 to 100. The provider wraps the text, so billing and speech-mark offsets are unchanged.                                                           |
| `language`                  | `language`                        | `de`, `de-DE` and so on. `auto` sends nothing.                                                                                                       |
| `instructions`              | Not sent                          | Returns a warning. Use SSML in `text` to shape delivery.                                                                                             |
| `providerOptions.speechify` | `options`                         | `loudnessNormalization` and `textNormalization`.                                                                                                     |

An option the API cannot honour comes back on `warnings` instead of being dropped silently.

## Speech marks

The provider puts the response's speech marks and billable character count on `providerMetadata.speechify`:

```ts
import type { SpeechifySpeechProviderMetadata } from "@speechify/ai-sdk-provider";

const { speechMarks, billableCharactersCount } =
  providerMetadata.speechify as SpeechifySpeechProviderMetadata;

for (const word of speechMarks?.chunks ?? []) {
  console.log(word.value, word.start_time, word.end_time);
}
```

Times are in milliseconds and `start` and `end` index characters of your text, as described in [Speech marks](/build/guides/text-to-speech/speech-marks).

## API version and attribution

Every request pins `Speechify-Version: 2026-09-30`, so the provider behaves the same whatever your workspace's default [API version](/build/guides/concepts/api-versioning) is.
At that version `simba-english` and `simba-multilingual` return `400` `model_retired`; use `simba-3.2` or `simba-3.0`.

Requests also carry `Speechify-Caller: vercel-ai-sdk` and the package version, so they appear under this integration in your request log. See [Build an integration](/build/guides/integrations/build-an-integration).

## When to use the streaming endpoint instead

`generateSpeech()` waits for the whole clip.
For playback that starts while audio is still being generated, such as a conversational turn, call [`POST /v1/audio/stream`](/build/guides/text-to-speech/streaming) directly.
It also accepts up to 20,000 characters per request.