Skip to navigation

Vercel AI SDK

Generate speech with generateSpeech() and Speechify’s Simba 3 models, speech marks included.

Overview

@speechify/ai-sdk-provider is Speechify’s provider for the Vercel AI SDK. It implements the AI SDK’s speech model interface, so generateSpeech() calls POST /v1/audio/speech and returns the audio together with its speech marks.

Speechify maintains the package, and its source is Speechify-AI/ai-sdk-provider.

Prerequisites

Install

npm install @speechify/ai-sdk-provider ai

Put your key in the environment of the server that runs the AI SDK:

SPEECHIFY_API_KEY=your_speechify_api_key

Generate speech

import { writeFile } from "node:fs/promises";
import { generateSpeech } from "ai";
import { speechify } from "@speechify/ai-sdk-provider";
const { audio, providerMetadata, warnings } = await generateSpeech({
model: speechify.speech("simba-3.2"),
text: "Hello from Speechify and the AI SDK.",
voice: "harper_32",
});
await writeFile("hello.mp3", audio.uint8Array);

speechify reads SPEECHIFY_API_KEY when the request is made. To pass a key or a proxy URL explicitly, create your own instance with createSpeechify({ apiKey, baseURL, headers, fetch }).

Run generateSpeech() on the server, for example in a route handler. An API key shipped to a browser is readable by anyone who opens the page.

How AI SDK options map to the API

generateSpeech() optionSpeechify requestNotes
modelmodelsimba-3.2 for English, simba-3.0 for the other supported languages.
textinputPlain text or SSML, up to 2,000 characters.
voicevoice_idDefaults to harper_32. Any voice whose models list in GET /v1/voices includes the model.
outputFormataudio_format or output_formatDefaults to mp3. Accepts mp3, wav, ogg, aac, pcm, mulaw and every codec_sampleRate_bitrate value such as pcm_16000 or ulaw_8000.
speedSSML <prosody rate>0.5 to 100. The provider wraps the text, so billing and speech-mark offsets are unchanged.
languagelanguagede, de-DE and so on. auto sends nothing.
instructionsNot sentReturns a warning. Use SSML in text to shape delivery.
providerOptions.speechifyoptionsloudnessNormalization and textNormalization.

An option the API cannot honour comes back on warnings instead of being dropped silently.

Speech marks

The provider puts the response’s speech marks and billable character count on providerMetadata.speechify:

import type { SpeechifySpeechProviderMetadata } from "@speechify/ai-sdk-provider";
const { speechMarks, billableCharactersCount } =
providerMetadata.speechify as SpeechifySpeechProviderMetadata;
for (const word of speechMarks?.chunks ?? []) {
console.log(word.value, word.start_time, word.end_time);
}

Times are in milliseconds and start and end index characters of your text, as described in Speech marks.

API version and attribution

Every request pins Speechify-Version: 2026-09-30, so the provider behaves the same whatever your workspace’s default API version is. At that version simba-english and simba-multilingual return 400 model_retired; use simba-3.2 or simba-3.0.

Requests also carry Speechify-Caller: vercel-ai-sdk and the package version, so they appear under this integration in your request log. See Build an integration.

When to use the streaming endpoint instead

generateSpeech() waits for the whole clip. For playback that starts while audio is still being generated, such as a conversational turn, call POST /v1/audio/stream directly. It also accepts up to 20,000 characters per request.