Streaming
Overview
The streaming endpoint delivers audio chunks as they’re generated, so your application can start playback before the full audio is ready. This is ideal for long-form content and low-latency applications. How the first-byte numbers are measured, what adds latency on your side and how to measure your own are on Latency.
Usage
The endpoint returns audio chunks via HTTP chunked transfer encoding. Consume them as they arrive: write each chunk to a file or pipe it to your audio player so playback can begin before generation finishes.
Supported audio formats
The Accept header is required and selects the audio container. These default to 24 kHz mono; send output_format to choose a different sample rate or bitrate, which overrides the header.
The audio/pcm request is mapped to the IANA-registered audio/L16 type on
the response (with rate and channels parameters per RFC 4856).
Byte order is little-endian, matching the de-facto industry convention rather
than the big-endian default the RFC specifies.
WAV format is not available for streaming. Use the speech endpoint for WAV output.
For the full output_format list, how it relates to audio_format, and the
telephony formats, see Audio Formats.
Use cases
Error handling
If an error occurs during synthesis after the stream has started, the connection closes without an error message. This is a limitation of HTTP chunked responses. Errors before streaming starts return standard HTTP status codes.
To handle mid-stream failures:
- Check the total bytes received against expected audio length
- Implement retry logic for the remaining text
Example projects
See our Examples Repository for complete browser and server-side streaming demos.