Skip to navigation

API: stream text into speech over a WebSocket

GET /v1/audio/stream/ws is a new WebSocket endpoint for text that arrives in pieces, such as a reply an LLM streams token by token. Send the text as you receive it, and the API speaks each sentence once it is complete, returning Base64 audio and, with speech_marks=true, word timestamps on one timeline.

  • Authenticate the upgrade with Authorization: Bearer. The settings ride the query string under the POST /v1/audio/stream names, and a refused setting, voice or plan limit is an ordinary HTTP error before the upgrade.
  • Each segment is one synthesis. It is billed, rate limited and moderated as one POST /v1/audio/stream request, and the voice’s prosody starts afresh at each segment.
  • A session holds one concurrent request. It ends after 60 seconds with nothing to speak, and stops taking text after 5 minutes.
  • Server-side only for now. A browser cannot set the Authorization header on a WebSocket.

No API version change is needed. See Streaming Text Input.