Skip to navigation

Create Voice

Create a cloned voice for the workspace from a 10-30 second audio sample, with verified consent from the speaker.

Cloning requires proof that the speaker agreed to it. Create a consent challenge with POST /v1/voices/consent-challenges, show the returned phrase to the speaker, record them reading it aloud, and send that recording here as consent_recording together with the challenge’s consent_challenge_id. Speechify transcribes the recording, checks it against the phrase it issued, checks that its speaker is the speaker in your sample, and keeps it as the consent record for the voice. The person consenting therefore has to be the person being cloned. A challenge is single use and short-lived, so record and submit in one sitting.

The clone belongs to the workspace rather than the member who created it, and access follows the caller’s workspace role and API-key scopes exactly as for any other voice: voices scopes to list it, audio scopes to synthesize with it, and the content-management permission plus a write scope on the key to delete it. Cloned voices are usable self-serve on simba-3.0 (and, on a workspace pinned before API version 2026-09-21, on the retired simba-english and simba-multilingual, which from 2026-11-21 are served by our current models and still take cloned voices). simba-3.2 also serves cloned voices.

Callers pinned before Speechify-Version: 2026-09-13 use the previous flow instead: no challenge, and a consent form field carrying the speaker’s name and email as a JSON string. That flow is switched off on 2026-09-23 for every API version: until then each create on it answers with Deprecation and Sunset headers naming the date, and from that date a create that sends consent and no consent_challenge_id returns 400 consent_verification_required.

Authentication

AuthorizationBearer

Enter your API key with the Bearer prefix, e.g. 'Bearer sk_...'.

Headers

Speechify-VersionstringOptional
Idempotency-KeystringOptional<=255 characters

A client-generated key (an opaque string, max 255 chars) that makes a side-effect POST safe to retry: the server runs the operation exactly once and replays the first response (its status and body) for 24 hours. Reusing a key with a different request body, or while the first request is still in flight, returns 409 idempotency_conflict. A replayed response carries the Idempotent-Replayed: true header.

Request

This endpoint expects a multipart form with multiple files.
namestringRequired
Name of the personal voice
localestringOptionalDefaults to en-US

Native language (locale) of the personal voice (e.g. en-US, es-ES, etc.)

genderenumRequired

Gender marker for the personal voice male GenderMale female GenderFemale not_specified GenderNotSpecified

Allowed values:
samplefileRequired

Audio sample of the voice to clone, 10-30 seconds of clean speech.

avatarfileOptional
Avatar image file
consent_challenge_idstringRequired

The id of the consent challenge this create consumes, from POST /v1/voices/consent-challenges. Single use: once a create has consumed it, whether or not that create succeeded, it cannot be used again.

consent_recordingfileRequired

Recording of the speaker reading the challenge's phrase aloud. This is the consent record for the voice, not a second voice sample: it must be the same person as in sample, and it is retained as evidence. 5-30 seconds, at most 25 MB, in any common audio container.

Response headers

Speechify-Request-IdstringOptional

Unique identifier for this request, present on every response (2xx and non-2xx alike). If the caller sends a Speechify-Request-Id request header the server echoes it back (sanitized and length-capped) so one logical request can be traced end-to-end; otherwise the server generates a fresh value. Log it on every response and quote it in support requests

  • it is the stable handle that ties your observation to Speechify's server-side logs, and it matches the request_id field in the error envelope.

The legacy alias X-Request-ID carries the same value and is still accepted on requests, until 2027-07-24. Prefer the un-prefixed name (RFC 6648).

RateLimit-LimitintegerOptional

Request-rate budget: the maximum number of requests in the current window (the bucket capacity). The IETF-draft un-prefixed name; the legacy alias X-RateLimit-Limit carries the same value. Rides every response.

RateLimit-RemainingintegerOptional

Request-rate budget: requests left in the current window. Legacy alias: X-RateLimit-Remaining.

RateLimit-ResetintegerOptional

Request-rate budget: integer delta-seconds until the window fully refills (same unit as Retry-After). Legacy alias: X-RateLimit-Reset.

Response

A created voice
display_namestring
genderenum
Allowed values:
localestring
idstring
modelslist of objects
typeenum
Allowed values:
avatar_imagestring or nullOptional
preview_audiostring or nullOptional
project_idstringOptionalformat: "^proj_[0-9a-hjkmnp-tv-z]{26}$"

The workspace project this cloned voice is filed under, set when a project-pinned key created it. Returned wherever a cloned voice is: the list, a single-voice read, and the create response.

Omitted for a shared-catalog voice and for a cloned voice no project filed, which is shared with the whole workspace and listed for every member of it.

can_managebooleanOptional

Whether this workspace may delete the voice and download its sample through this API. true for a cloned voice the workspace owns. false for a shared-catalog voice, and for a cloned voice that reaches this workspace only through its creator's personal account (a voice cloned before workspace ownership, or under another Speechify product), which is managed where it was made.

tagslist of strings or nullOptional

Errors

400
Bad Request Error
401
Unauthorized Error
402
Payment Required Error
403
Forbidden Error
409
Conflict Error
413
Content Too Large Error
422
Unprocessable Entity Error
429
Too Many Requests Error
500
Internal Server Error
502
Bad Gateway Error
503
Service Unavailable Error