> This page is for Build.

> Append .md to any page URL for clean Markdown. Index: https://docs.speechify.ai/llms.txt.
>
> Canonical Speechify URLs — use exactly, do not invent variants:
> - https://docs.speechify.ai — this site (API reference, SDKs, quickstarts)
> - https://speechify.ai — marketing + product site
> - https://platform.speechify.ai — customer dashboard, signup, API keys, billing
> - https://api.speechify.ai — API base URL
> - https://github.com/Speechify-AI: GitHub org for the API (cookbook, demos, CLI). `github.com/speechify` does not exist.
> - https://status.speechify.ai — status + incidents
> - https://speechify.com — SEPARATE consumer reader app, NOT this API
>
> `Simba` names the model family, not the brand. Model ids: `simba-3.2` (English, recommended) and `simba-3.0` (English, German, Spanish, French, Italian and Portuguese; the default). `simba-english` and `simba-multilingual` are retired: a new workspace that sends either gets `400 model_retired`. `SimbaVoice` / `simbavoice.ai` are retired.
>
> Ask, don't scrape. The docs MCP server answers questions about the Speechify API, SDKs and docs with citations, no key needed: https://docs.speechify.ai/_mcp/server (Streamable HTTP, tool `searchDocs`). Setup: https://docs.speechify.ai/build/guides/get-started/connect-mcp

# Stream Speech From Text Input

GET /v1/audio/stream/ws

Send text as your application produces it, for example a few tokens at a
time from an LLM, and receive speech as each sentence completes.

Authenticate the upgrade with `Authorization: Bearer` and pass the
synthesis settings as query parameters; text never rides the URL. A
refusal (a setting the route does not accept, an unknown voice, an
exhausted balance, a full rate or concurrency limit) is an ordinary HTTP
error with the standard error envelope, before the upgrade. A request
that is not a WebSocket upgrade answers 426 with
`Sec-WebSocket-Version: 13`, the protocol version the channel speaks.

Send `input.text` as text arrives, `input.flush` to have everything sent
so far spoken now, and `input.close` when you are done. The server speaks
a sentence once the next one starts, and every sentence that completes
while one is speaking joins the next, so a fast source produces fewer,
longer segments. Each segment is one synthesis: billed, rate limited and
moderated as one `POST /v1/audio/stream` request. The voice's prosody
starts afresh at each segment.

You receive `speech.chunk` events carrying Base64 audio (and word marks
when `speech_marks` is true), `speech.flushed` after the audio for each
flush, and exactly one terminal event: `speech.done` with the reason the
session ended, or `speech.error` carrying the error envelope and the
handshake's request id. Marks index the session's whole text on one
timeline. Ignore event types you do not recognize.

A session holds one concurrent request for its life. It ends after 60
seconds with nothing to speak and no message (`idle`), and after 5
minutes it stops taking text and speaks what it was sent
(`session_limit`); open a new session to continue. A server restart ends
it with `shutdown`. The socket then closes with code 1000, or 1001 for a
shutdown; a `speech.error` closes it with 1008 for a refusal,
1003 for a binary frame, 1009 for a message or text over the limits, and
1011 for a server or upstream fault.

Reference: https://docs.speechify.ai/build/api-reference/v1/audio/stream/ws

## AsyncAPI Specification

```yaml
asyncapi: 2.6.0
info:
  title: Stream Speech From Text Input
  version: subpackage_audio.Stream Speech From Text Input
  description: |-
    Send text as your application produces it, for example a few tokens at a
    time from an LLM, and receive speech as each sentence completes.

    Authenticate the upgrade with `Authorization: Bearer` and pass the
    synthesis settings as query parameters; text never rides the URL. A
    refusal (a setting the route does not accept, an unknown voice, an
    exhausted balance, a full rate or concurrency limit) is an ordinary HTTP
    error with the standard error envelope, before the upgrade. A request
    that is not a WebSocket upgrade answers 426 with
    `Sec-WebSocket-Version: 13`, the protocol version the channel speaks.

    Send `input.text` as text arrives, `input.flush` to have everything sent
    so far spoken now, and `input.close` when you are done. The server speaks
    a sentence once the next one starts, and every sentence that completes
    while one is speaking joins the next, so a fast source produces fewer,
    longer segments. Each segment is one synthesis: billed, rate limited and
    moderated as one `POST /v1/audio/stream` request. The voice's prosody
    starts afresh at each segment.

    You receive `speech.chunk` events carrying Base64 audio (and word marks
    when `speech_marks` is true), `speech.flushed` after the audio for each
    flush, and exactly one terminal event: `speech.done` with the reason the
    session ended, or `speech.error` carrying the error envelope and the
    handshake's request id. Marks index the session's whole text on one
    timeline. Ignore event types you do not recognize.

    A session holds one concurrent request for its life. It ends after 60
    seconds with nothing to speak and no message (`idle`), and after 5
    minutes it stops taking text and speaks what it was sent
    (`session_limit`); open a new session to continue. A server restart ends
    it with `shutdown`. The socket then closes with code 1000, or 1001 for a
    shutdown; a `speech.error` closes it with 1008 for a refusal,
    1003 for a binary frame, 1009 for a message or text over the limits, and
    1011 for a server or upstream fault.
channels:
  /v1/audio/stream/ws:
    description: |-
      Send text as your application produces it, for example a few tokens at a
      time from an LLM, and receive speech as each sentence completes.

      Authenticate the upgrade with `Authorization: Bearer` and pass the
      synthesis settings as query parameters; text never rides the URL. A
      refusal (a setting the route does not accept, an unknown voice, an
      exhausted balance, a full rate or concurrency limit) is an ordinary HTTP
      error with the standard error envelope, before the upgrade. A request
      that is not a WebSocket upgrade answers 426 with
      `Sec-WebSocket-Version: 13`, the protocol version the channel speaks.

      Send `input.text` as text arrives, `input.flush` to have everything sent
      so far spoken now, and `input.close` when you are done. The server speaks
      a sentence once the next one starts, and every sentence that completes
      while one is speaking joins the next, so a fast source produces fewer,
      longer segments. Each segment is one synthesis: billed, rate limited and
      moderated as one `POST /v1/audio/stream` request. The voice's prosody
      starts afresh at each segment.

      You receive `speech.chunk` events carrying Base64 audio (and word marks
      when `speech_marks` is true), `speech.flushed` after the audio for each
      flush, and exactly one terminal event: `speech.done` with the reason the
      session ended, or `speech.error` carrying the error envelope and the
      handshake's request id. Marks index the session's whole text on one
      timeline. Ignore event types you do not recognize.

      A session holds one concurrent request for its life. It ends after 60
      seconds with nothing to speak and no message (`idle`), and after 5
      minutes it stops taking text and speaks what it was sent
      (`session_limit`); open a new session to continue. A server restart ends
      it with `shutdown`. The socket then closes with code 1000, or 1001 for a
      shutdown; a `speech.error` closes it with 1008 for a refusal,
      1003 for a binary frame, 1009 for a message or text over the limits, and
      1011 for a server or upstream fault.
    bindings:
      ws:
        query:
          type: object
          properties:
            voice_id:
              type: string
            model:
              $ref: '#/components/schemas/audioStreamInput_model'
              default: simba-3.0
            output_format:
              $ref: '#/components/schemas/audioStreamInput_output_format'
            language:
              type: string
            speech_marks:
              $ref: '#/components/schemas/audioStreamInput_speech_marks'
              default: 'false'
            safety_identifier:
              type: string
        headers:
          type: object
          properties:
            Speechify-Version:
              type: string
    publish:
      operationId: subpackage_audio.Stream Speech From Text Input-publish
      summary: Server messages
      message:
        oneOf:
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-server-0-speechChunk
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-server-1-speechFlushed
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-server-2-speechDone
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-server-3-speechError
    subscribe:
      operationId: subpackage_audio.Stream Speech From Text Input-subscribe
      summary: Client messages
      message:
        oneOf:
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-client-0-inputText
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-client-1-inputFlush
          - $ref: >-
              #/components/messages/subpackage_audio.Stream Speech From Text
              Input-client-2-inputClose
servers:
  https://api.speechify.ai:
    url: wss://api.speechify.ai/
    protocol: wss
    x-default: true
components:
  messages:
    subpackage_audio.Stream Speech From Text Input-server-0-speechChunk:
      name: speechChunk
      title: speech.chunk
      payload:
        $ref: '#/components/schemas/audioStreamInput_speechChunk'
    subpackage_audio.Stream Speech From Text Input-server-1-speechFlushed:
      name: speechFlushed
      title: speech.flushed
      payload:
        $ref: '#/components/schemas/audioStreamInput_speechFlushed'
    subpackage_audio.Stream Speech From Text Input-server-2-speechDone:
      name: speechDone
      title: speech.done
      payload:
        $ref: '#/components/schemas/audioStreamInput_speechDone'
    subpackage_audio.Stream Speech From Text Input-server-3-speechError:
      name: speechError
      title: speech.error
      payload:
        $ref: '#/components/schemas/audioStreamInput_speechError'
    subpackage_audio.Stream Speech From Text Input-client-0-inputText:
      name: inputText
      title: input.text
      payload:
        $ref: '#/components/schemas/audioStreamInput_inputText'
    subpackage_audio.Stream Speech From Text Input-client-1-inputFlush:
      name: inputFlush
      title: input.flush
      payload:
        $ref: '#/components/schemas/audioStreamInput_inputFlush'
    subpackage_audio.Stream Speech From Text Input-client-2-inputClose:
      name: inputClose
      title: input.close
      payload:
        $ref: '#/components/schemas/audioStreamInput_inputClose'
  schemas:
    audioStreamInput_model:
      type: string
      enum:
        - simba-3.0
        - simba-3.2
      default: simba-3.0
      description: >-
        Model used for every segment, as for `POST /v1/audio/stream`.
        `simba-3.2` is English only.
      title: audioStreamInput_model
    audioStreamInput_output_format:
      type: string
      enum:
        - pcm_8000
        - pcm_16000
        - pcm_22050
        - pcm_24000
        - pcm_44100
        - pcm_48000
        - mp3_22050_32
        - mp3_22050_64
        - mp3_22050_96
        - mp3_22050_128
        - mp3_22050_160
        - mp3_22050_192
        - mp3_24000_32
        - mp3_24000_64
        - mp3_24000_96
        - mp3_24000_128
        - mp3_24000_160
        - mp3_24000_192
        - ulaw_8000
        - ogg_24000
        - aac_24000
      description: >-
        Audio format of every `speech.chunk`, as on `POST /v1/audio/stream`.
        Defaults to MP3 at 24 kHz.
      title: audioStreamInput_output_format
    audioStreamInput_speech_marks:
      type: string
      enum:
        - 'true'
        - 'false'
      default: 'false'
      description: >-
        When true, each `speech.chunk` carries the word marks that became final
        with it, indexing the session's whole text.
      title: audioStreamInput_speech_marks
    ChannelsAudioStreamInputMessagesSpeechChunkSpeechMarksItems:
      type: object
      properties:
        end:
          type: integer
          format: int64
        end_time:
          type: number
          format: double
        start:
          type: integer
          format: int64
        start_time:
          type: number
          format: double
        type:
          type: string
        value:
          type: string
      description: >-
        It details the type of segment, its start and end points in the text,
        and its start and end times in the synthesized speech audio.
      title: ChannelsAudioStreamInputMessagesSpeechChunkSpeechMarksItems
    audioStreamInput_speechChunk:
      type: object
      properties:
        type:
          type: string
          enum:
            - speech.chunk
        audio:
          type: string
          description: |-
            A run of the synthesized audio, Base64-encoded, in the format the
            request selected (echoed on the `Speechify-Audio-Content-Type`
            response header). Absent on a marks-only chunk.
        speech_marks:
          type: array
          items:
            $ref: >-
              #/components/schemas/ChannelsAudioStreamInputMessagesSpeechChunkSpeechMarksItems
          description: |-
            Word timings addressing the original input text, with absolute
            millisecond times from the start of the synthesis. Absent when the
            chunk carries only audio.
      required:
        - type
      description: |-
        A run of synthesized audio, the speech marks that became final with it,
        or both - a chunk may carry only one of the two, and the last chunk of
        a stream is often marks-only. Mark times are absolute milliseconds from
        the start of the synthesis: concatenate the audio chunks into one
        stream and apply the marks against that single timeline. Which chunk a
        mark arrives on is a delivery detail and carries no meaning.
      title: audioStreamInput_speechChunk
    audioStreamInput_speechFlushed:
      type: object
      properties:
        type:
          type: string
          enum:
            - speech.flushed
      required:
        - type
      description: >-
        Server event on GET /v1/audio/stream/ws: the audio for everything sent
        before an input.flush has been sent. One arrives for every input.flush,
        in order.
      title: audioStreamInput_speechFlushed
    ChannelsAudioStreamInputMessagesSpeechDoneReason:
      type: string
      enum:
        - closed
        - idle
        - session_limit
        - shutdown
      description: >-
        closed: the client sent input.close and everything was spoken. idle: 60
        seconds passed with nothing to speak and no message. session_limit: the
        session reached its 5 minute limit, and everything sent before it was
        spoken. shutdown: the server is restarting; reconnect, and the next
        session lands on another server. Ignore a reason you do not recognize.
      title: ChannelsAudioStreamInputMessagesSpeechDoneReason
    audioStreamInput_speechDone:
      type: object
      properties:
        type:
          type: string
          enum:
            - speech.done
        billable_characters_count:
          type: integer
          description: Number of billable characters processed.
        audio_duration_ms:
          type: integer
          description: Duration of the synthesized audio in milliseconds.
        reason:
          $ref: >-
            #/components/schemas/ChannelsAudioStreamInputMessagesSpeechDoneReason
          description: >-
            closed: the client sent input.close and everything was spoken. idle:
            60 seconds passed with nothing to speak and no message.
            session_limit: the session reached its 5 minute limit, and
            everything sent before it was spoken. shutdown: the server is
            restarting; reconnect, and the next session lands on another server.
            Ignore a reason you do not recognize.
      required:
        - type
        - billable_characters_count
        - audio_duration_ms
        - reason
      description: >-
        What speech.done adds on GET /v1/audio/stream/ws, where it ends a whole
        session rather than one synthesis: why the session ended. Its
        billable_characters_count and audio_duration_ms total every segment the
        session spoke.
      title: audioStreamInput_speechDone
    ChannelsAudioStreamInputMessagesSpeechErrorErrorCode:
      type: string
      enum:
        - endpoint_moved
        - bad_request
        - validation_failed
        - unauthorized
        - payment_required
        - forbidden
        - not_found
        - method_not_allowed
        - conflict
        - idempotency_conflict
        - payload_too_large
        - unsupported_media_type
        - rate_limited
        - concurrency_limit_reached
        - invalid_api_version
        - internal_error
        - upstream_failure
        - service_unavailable
        - caller_not_found
        - contact_not_found
        - contact_identifier_not_found
        - contact_identifier_conflict
        - contact_resolver_not_found
        - credential_not_found
        - credential_in_use
        - agent_not_found
        - agent_in_use
        - agent_run_not_found
        - kb_not_found
        - kb_document_not_found
        - kb_folder_not_found
        - tool_not_found
        - tool_name_taken
        - channel_instance_not_found
        - team_not_found
        - trigger_not_found
        - store_not_found
        - store_document_not_found
        - hosted_api_not_found
        - api_route_not_found
        - consumer_key_not_found
        - skill_not_found
        - skill_version_not_found
        - file_not_found
        - file_path_taken
        - file_storage_limit_reached
        - store_limit_reached
        - store_document_limit_reached
        - store_bytes_limit_reached
        - store_not_configured
        - store_document_version_conflict
        - store_document_deleted
        - entitlement_override_exists
        - hosted_apis_not_in_plan
        - skills_not_in_plan
        - voice_agents_not_in_plan
        - skill_in_use
        - skill_tool_name_conflict
        - skill_limit_reached
        - agent_skill_limit_reached
        - hosted_api_slug_taken
        - api_route_conflict
        - mount_plan_changed
        - route_output_unavailable
        - route_run_timeout
        - route_run_failed
        - route_run_limit_reached
        - route_read_limit_reached
        - route_write_limit_reached
        - hosted_api_public_refused
        - route_tool_not_readable
        - route_tool_unavailable
        - route_upstream_rate_limited
        - route_upstream_error
        - hosted_mcp_not_enabled
        - hosted_api_busy
        - conversation_not_found
        - phone_number_not_found
        - sip_trunk_not_found
        - voice_not_found
        - audio_asset_not_found
        - builtin_not_found
        - batch_not_found
        - agent_test_not_found
        - workspace_not_found
        - invite_not_found
        - project_not_found
        - cross_project_reference
        - project_has_scoped_credentials
        - project_not_empty
        - project_limit_reached
        - agent_limit_reached
        - project_too_large_to_promote
        - call_not_found
        - message_not_found
        - thread_not_found
        - call_not_active
        - relay_displaces_agent
        - brain_not_found
        - brain_in_use
        - custom_model_not_found
        - custom_model_in_use
        - insufficient_scope
        - purchased_numbers_not_included
        - phone_number_quota_reached
        - batch_calls_not_included
        - voice_cloning_not_included
        - voice_cloning_unavailable_in_region
        - consent_challenge_not_found
        - consent_challenge_expired
        - consent_challenge_already_used
        - consent_phrase_mismatch
        - consent_speaker_mismatch
        - consent_recording_unusable
        - consent_verification_unavailable
        - consent_verification_required
        - watermark_audio_unusable
        - watermark_detection_unavailable
        - workspace_last_owner
        - workspace_last_workspace
        - account_deletion_blocked
        - workspace_free_limit
        - workspace_single_owner
        - invite_email_mismatch
        - invite_already_pending
        - service_account_limit_reached
        - service_accounts_not_in_plan
        - speech_marks_unsupported
        - model_retired
        - too_many_voices
        - content_policy_violation
        - safety_identifier_blocked
        - topup_not_in_plan
        - credit_purchase_unpaid
        - credit_purchase_payment_in_progress
        - tool_config_shared
        - spend_cap_exceeded
        - spend_budget_exceeded
        - project_spend_limit_exceeded
        - project_archived
        - project_not_archived
        - project_not_purged
        - project_restore_window_expired
        - project_name_taken
        - funded_balance_required
        - agent_publish_gate_failed
        - agent_publish_gate_required
        - agent_publish_gate_unavailable
        - agent_publish_gate_tool_unreachable
        - text_channel_not_in_plan
        - channel_not_in_plan
        - text_turn_failed
        - conversation_turn_in_progress
        - conversation_closed
        - conversation_not_reachable
        - conversation_channel_bound
        - conversation_prompt_not_found
        - text_message_quota_exceeded
        - durable_runs_not_in_plan
        - tool_transport_unsupported
        - agent_config_too_large
        - agent_run_not_pending
        - agent_run_action_stale
        - share_link_not_found
        - share_link_exhausted
        - share_link_limit_reached
        - destination_not_allowed
        - international_dialing_not_enabled
        - number_not_sms_capable
        - verification_required
        - intended_use_required
      description: |
        Stable machine-readable error code. Additive only: codes are
        never renamed, only deprecated. SDKs may map each code to a
        typed exception class. Status-code semantics:
        4xx codes describe caller-fixable issues; 5xx codes describe
        server-side failures and are safe to retry with backoff for
        idempotent requests.
      title: ChannelsAudioStreamInputMessagesSpeechErrorErrorCode
    ChannelsAudioStreamInputMessagesSpeechErrorError:
      type: object
      properties:
        code:
          $ref: >-
            #/components/schemas/ChannelsAudioStreamInputMessagesSpeechErrorErrorCode
          description: |
            Stable machine-readable error code. Additive only: codes are
            never renamed, only deprecated. SDKs may map each code to a
            typed exception class. Status-code semantics:
            4xx codes describe caller-fixable issues; 5xx codes describe
            server-side failures and are safe to retry with backoff for
            idempotent requests.
        message:
          type: string
          description: |
            Human-readable explanation of this specific occurrence.
            Safe to surface in UI banners or pass to support. The
            wording can change between releases; clients should
            match on `code`, not on the message string.
        fields:
          type: object
          additionalProperties:
            type: string
          description: |
            Per-field validation errors as `path -> message`. Only
            present on 400 responses caused by request validation
            (typically code=`validation_failed`). Keys are field
            paths in dotted/bracket notation; values are short
            human explanations safe to inline-surface next to the
            offending form field.
        details:
          type: object
          additionalProperties:
            description: Any type
          description: |
            Structured, endpoint-specific context beyond the flat
            `fields` map. Present only on the few errors that carry
            it (e.g. the `used_by` referrer list on a credential
            delete-conflict); its shape depends on the error `code`.
            Clients that don't recognise a `details` shape can ignore
            it - the `code` + `message` contract is unchanged.
        docs_url:
          type: string
          format: uri
          description: |
            Link to the documentation that resolves this class of
            error, when a stable page exists. Rate and concurrency
            429s link the API limits reference, which lists each
            plan's limits and how to raise them.
      required:
        - code
        - message
      title: ChannelsAudioStreamInputMessagesSpeechErrorError
    audioStreamInput_speechError:
      type: object
      properties:
        type:
          type: string
          enum:
            - speech.error
        error:
          $ref: >-
            #/components/schemas/ChannelsAudioStreamInputMessagesSpeechErrorError
        request_id:
          type: string
          description: |-
            Server-side request identifier. Echoes the `Speechify-Request-Id`
            response header.
      required:
        - type
        - error
      description: |-
        Terminal event carrying the standard error envelope, emitted when a
        failure happens after the stream has started and the status code is
        already committed: an upstream fault (`upstream_failure`) or a content
        policy refusal (`content_policy_violation`).
      title: audioStreamInput_speechError
    audioStreamInput_inputText:
      type: object
      properties:
        type:
          type: string
          enum:
            - input.text
        text:
          type: string
          maxLength: 20000
          description: >-
            The next piece of text. Include the spaces and newlines between
            words and sentences exactly as your source produced them: they are
            what tells the server a sentence has ended.
      required:
        - type
        - text
      description: >-
        Client message on GET /v1/audio/stream/ws: text appended to what the
        session will speak. Send it as your application produces it, a few
        tokens at a time is fine; the server speaks each sentence once the next
        one starts, so a reply that arrives in pieces is cut at its sentence
        boundaries, not at the message boundaries. The text is spoken as plain
        text: SSML is read aloud, not interpreted. At most 20000 characters can
        wait to be spoken at once; more is refused with payload_too_large, which
        ends the session.
      title: audioStreamInput_inputText
    audioStreamInput_inputFlush:
      type: object
      properties:
        type:
          type: string
          enum:
            - input.flush
      required:
        - type
      description: >-
        Client message on GET /v1/audio/stream/ws: speak everything sent so far
        now, without waiting for its sentence to end. The server answers with
        speech.flushed once that audio has been sent, at once when nothing was
        waiting. Use it at the end of an LLM turn.
      title: audioStreamInput_inputFlush
    audioStreamInput_inputClose:
      type: object
      properties:
        type:
          type: string
          enum:
            - input.close
      required:
        - type
      description: >-
        Client message on GET /v1/audio/stream/ws: no more text is coming. The
        server speaks everything it was sent, sends speech.done with reason
        closed, and closes the socket. Keep reading until speech.done arrives:
        closing the socket first abandons the audio still to come, and the
        segment it cuts is billed in full.
      title: audioStreamInput_inputClose

```