Skip to navigation

Streaming

Generate and play audio in real time over chunked transfer encoding; first byte 56 ms p50 and 102 ms p90 on Simba 3.2, measured 15 Sep 2026

Overview

The streaming endpoint delivers audio chunks as they’re generated, so your application can start playback before the full audio is ready. This is ideal for long-form content and low-latency applications. How the first-byte numbers are measured, what adds latency on your side and how to measure your own are on Latency.

Speech endpointStream endpoint
Character limit2,00020,000
Response formatBase64 JSON + metadataRaw audio chunks
Playback startAfter full generationImmediately

Usage

POST
/v1/audio/stream
from speechify import Speechify
client = Speechify(
token="YOUR_TOKEN_HERE",
)
client.audio.stream(
input="Streaming long-form audio with the Speechify API.",
model="simba-3.2",
voice_id="geffen_32",
)

The endpoint returns audio chunks via HTTP chunked transfer encoding. Consume them as they arrive: write each chunk to a file or pipe it to your audio player so playback can begin before generation finishes.

Supported audio formats

The Accept header is required and selects the audio container. These default to 24 kHz mono; send output_format to choose a different sample rate or bitrate, which overrides the header.

FormatAccept headerResponse Content-TypeNotes
MP3audio/mpegaudio/mpeg128 kbps. Best compatibility. Send output_format to pick another bitrate.
Opusaudio/oggaudio/oggOgg container, Opus codec. Open format.
AACaudio/aacaudio/aacAAC-LC (ADTS-framed). Apple ecosystem.
PCMaudio/pcmaudio/L16; rate=24000; channels=1Raw 16-bit signed little-endian samples, no container or header. Lowest latency.

The audio/pcm request is mapped to the IANA-registered audio/L16 type on the response (with rate and channels parameters per RFC 4856). Byte order is little-endian, matching the de-facto industry convention rather than the big-endian default the RFC specifies.

WAV format is not available for streaming. Use the speech endpoint for WAV output.

For the full output_format list, how it relates to audio_format, and the telephony formats, see Audio Formats.

Use cases

Automated podcast generation

Transform articles or blog posts into spoken audio for distribution

Assistive technology

Convert on-screen text to spoken audio in real-time

Voice agents

Generate conversational responses with minimal latency

Audiobook production

Process full chapters without hitting the 2K character limit

Error handling

If an error occurs during synthesis after the stream has started, the connection closes without an error message. This is a limitation of HTTP chunked responses. Errors before streaming starts return standard HTTP status codes.

To handle mid-stream failures:

  • Check the total bytes received against expected audio length
  • Implement retry logic for the remaining text

Example projects

See our Examples Repository for complete browser and server-side streaming demos.