> This page is for Build.

> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Vapi

> Use the tts-shims Vapi provider to answer Vapi's custom TTS voice-request with raw PCM from Speechify while keeping your API key server-side.

## Overview

Vapi custom voice calls a webhook you host and expects raw PCM back. The [`tts-shims`](https://github.com/Speechify-AI/tts-shims) repo ships a `vapi` provider — a small Go binary that answers Vapi's `POST` with `Content-Type: application/octet-stream` mono 16-bit little-endian PCM at the requested sample rate, translating the request into a Speechify stream call under the hood.

Your `SPEECHIFY_API_KEY` lives on the shim server; Vapi never sees it.

## Why a Vapi-specific shim?

Vapi custom TTS doesn't send an OpenAI-shaped body. It posts a nested message object:

```json
{
  "message": {
    "type": "voice-request",
    "text": "The text the assistant wants to say.",
    "sampleRate": 24000,
    "timestamp": 1720000000000,
    "call": {},
    "assistant": {}
  }
}
```

The response contract is equally specific: raw PCM, mono, signed 16-bit little-endian, sample rate matching `message.sampleRate`, no WAV header. No OpenAI-shaped shim can serve this — hence the dedicated `vapi` provider.

## Prerequisites

* Speechify API key
* Vapi account with a custom-voice-capable plan
* Public HTTPS URL for the shim (Vapi calls it from its cloud)
* Go toolchain locally (only to build the shim)

## Build and run the shim

```bash
git clone https://github.com/Speechify-AI/tts-shims
cd tts-shims
make vapi
```

Start it with your Speechify key in the environment:

```bash
export SPEECHIFY_API_KEY=sk_your_key_here
export SHIM_ADDR=:8772
export SHIM_DEFAULT_MODEL=simba-3.2
export VAPI_SECRET=replace_with_a_random_secret

./bin/vapi
```

The provider mounts one route: `POST /synthesize`. Defaults: voice `geffen_32`, model `simba-3.2`.

## Verify without a Vapi account

The custom TTS contract is HTTP — you can prove the whole path locally by sending the exact body Vapi sends:

```bash
curl -s -o shim-smoke.pcm -w "%{http_code} %{size_download} %{content_type}\n" \
  -X POST "http://localhost:8772/synthesize" \
  -H "Content-Type: application/json" \
  -H "X-VAPI-SECRET: replace_with_a_random_secret" \
  -d '{"message":{"type":"voice-request","text":"Vapi custom voice, now speaking with Speechify.","sampleRate":24000,"timestamp":1720000000000,"call":{},"assistant":{}}}'
```

A working shim returns:

```text
200 <byte-count> application/octet-stream
```

Wrap it as WAV locally to spot-check:

```bash
ffmpeg -f s16le -ar 24000 -ac 1 -i shim-smoke.pcm shim-smoke.wav
```

## Configure Vapi custom voice

In your Vapi assistant config:

```json
{
  "voice": {
    "provider": "custom-voice",
    "server": {
      "url": "https://your-shim-host.example/synthesize",
      "secret": "replace_with_a_random_secret"
    },
    "timeoutSeconds": 30
  }
}
```

Vapi sends the `server.secret` value as `X-VAPI-SECRET`; the shim validates that header before calling Speechify. Vapi also supports custom headers and `credentialId` flows if that's your credential-management style.

## Sample rate matters

Vapi's documented custom TTS sample rates are `8000`, `16000`, `22050`, and `24000`. The response must match `message.sampleRate` exactly. The shim validates the incoming rate and maps it to Speechify's stream output format.

> **Note**
>
> The reference test above uses `24000`. If you deploy for telephony (`8000` / `16000`), test with the actual rate before sending live calls through it — sample-rate mismatches surface as choppy or noisy audio, not clean errors.

## Never return WAV

Vapi's contract is raw PCM with `Content-Type: application/octet-stream`. Do not return a WAV file, MP3, or any container-wrapped audio. The demo saves a local `.wav` only after the response is written to disk, for humans and tools to spot-check — that's not what Vapi receives.

## Production deployment

* Put the shim behind HTTPS and a service host that holds environment variables (Cloud Run, Fly, Render, ECS — anywhere).
* Set `SPEECHIFY_API_KEY`, `SHIM_ADDR`, `SHIM_DEFAULT_MODEL`, and `VAPI_SECRET`.
* Rotate `VAPI_SECRET` and your Speechify key on a normal schedule — the shim doesn't cache either.
* Add health check monitoring on `GET /healthz`.

## Resources

#### [Demo repo](https://github.com/Speechify-AI/demos/tree/main/demos/vapi-custom-voice)

Runnable end-to-end demo with `run.sh` that clones the shim, builds it, and executes the smoke test.

#### [tts-shims repo](https://github.com/Speechify-AI/tts-shims)

Open-source Go proxy — Vapi, OpenAI, and other provider dialects under `cmd/`.

#### [Vapi custom TTS docs](https://docs.vapi.ai/customization/custom-voices/custom-tts)

Vapi's request/response contract for `custom-voice`.