Vercel AI SDK
Generate speech with generateSpeech() and Speechify’s Simba 3 models, speech marks included.
Overview
@speechify/ai-sdk-provider is Speechify’s provider for the Vercel AI SDK.
It implements the AI SDK’s speech model interface, so generateSpeech() calls POST /v1/audio/speech and returns the audio together with its speech marks.
Speechify maintains the package, and its source is Speechify-AI/ai-sdk-provider.
Prerequisites
- Speechify API key: sign up at platform.speechify.ai
- AI SDK 7 (
ai@^7) on Node.js 22 or later
Install
Put your key in the environment of the server that runs the AI SDK:
Generate speech
speechify reads SPEECHIFY_API_KEY when the request is made.
To pass a key or a proxy URL explicitly, create your own instance with createSpeechify({ apiKey, baseURL, headers, fetch }).
Run generateSpeech() on the server, for example in a route handler. An API key shipped to a browser is readable by anyone who opens the page.
How AI SDK options map to the API
An option the API cannot honour comes back on warnings instead of being dropped silently.
Speech marks
The provider puts the response’s speech marks and billable character count on providerMetadata.speechify:
Times are in milliseconds and start and end index characters of your text, as described in Speech marks.
API version and attribution
Every request pins Speechify-Version: 2026-09-30, so the provider behaves the same whatever your workspace’s default API version is.
At that version simba-english and simba-multilingual return 400 model_retired; use simba-3.2 or simba-3.0.
Requests also carry Speechify-Caller: vercel-ai-sdk and the package version, so they appear under this integration in your request log. See Build an integration.
When to use the streaming endpoint instead
generateSpeech() waits for the whole clip.
For playback that starts while audio is still being generated, such as a conversational turn, call POST /v1/audio/stream directly.
It also accepts up to 20,000 characters per request.