Voice Cloning Quickstart
Clone a voice and synthesize with it in minutes
The fastest path to a working cloned voice, covering both legs of the flow: the consent leg (a challenge phrase, read aloud and recorded), then the cloning leg (one create call with the sample and the recording), and you are synthesizing in the new voice.
Cloning requires a consent recording. The speaker reads a phrase Speechify issues and you send that recording with the create call - so plan for a live speaker at a microphone, not just an audio file.
Record a sample
Capture 10-30 seconds of clean speech, under a minute and under 5MB. Avoid background noise.
Record the speaker reading the phrase
This recording is the consent record. It has to be the same person as in your sample, 5-30 seconds, at most 25MB. The challenge is single use and expires at expires_at, so record and submit in one sitting.
Cloned voices work self-serve on
simba-3.0 (and on simba-english / simba-multilingual for a workspace pinned to an API version before 2026-09-21, until they are switched off on 2026-11-21). simba-3.2 serves them self-serve too, and is English only, so a non-English clone returns 400 there.