Short answer
Voice cloning services want a few minutes of one clean voice: no music, no second speaker, no room echo, no long pauses. Take your best recording, cut out the clean passages, remove the pauses, split the result into 30 to 90 second pieces, listen once, and upload the pieces you would be happy to hear repeated a thousand times. All three steps run in your browser in ChunkAudio, and the audio never leaves your computer.
The clone can only be as good as the samples. Every service, whether ElevenLabs, Resemble AI, PlayHT or an open model, learns from what you give it, and it learns the flaws with the same enthusiasm as the voice: the hum of a fridge, a reverberant room, the co-host who laughs in the background. This guide is about getting those out before you upload.
What the services ask for
The numbers below are typical tiers as of September 2026. Check the current limits of the service you use; they change often.
| Service | Instant clone | Professional clone | Formats |
|---|---|---|---|
| ElevenLabs | 1 to 5 minutes | 30 minutes to 3 hours | MP3, WAV |
| Resemble AI | about 1 minute | 25 sentences and up | WAV, MP3 |
| PlayHT | 30 seconds | 20 minutes and up | MP3, WAV |
| Open models (XTTS, Tortoise) | 6 to 30 seconds | as much clean audio as you have | WAV |
Two rules hold across all of them: the sample should contain one voice, and every second of it should be speech you would accept as the voice's default sound.
What makes a good sample
- One speaker, no music. Intros, jingles and guests must go. The model cannot separate them.
- Dry room. Echo becomes part of the voice. A closet full of clothes beats a kitchen.
- Same microphone and distance throughout. A mix of recordings from different setups produces a clone that sounds like two people.
- Natural delivery with variety. Questions, statements, a laugh, a serious passage. Read-aloud text with a flat tone gives a flat clone.
- No long pauses. Silence is training time wasted, and some services count it against the minimum.
- Technical: 44.1 kHz or higher, 16-bit or better, mono. WAV if you have it; MP3 at 192 kbps or better is fine. Do not "upgrade" a low-bitrate MP3 by re-encoding it, that adds nothing.
The five steps in ChunkAudio
Pick the recording
Choose the recording with the best sound, not the longest one. A solo podcast episode, a narration, a voice memo recorded close to the microphone in a quiet room. Five clean minutes beat thirty mediocre ones.
Cut out the clean passages
Open the file in cut mode, drag the two handles around a stretch where only the target voice speaks, press play to check it, and save the selection. Repeat for each clean passage. The cut is sample-exact on WAV, and for MP3 the frames are copied rather than re-encoded, so the sound is untouched. More on cutting in the audio cutter guide.
Remove the pauses
Switch to remove pauses. It finds every silence longer than the limit you set, one second is a good default for voice samples, and shortens each to a natural gap. You get one file of continuous speech, which is what the models want. The silence remover page explains the settings.
Split into samples
Back in split mode, set 60 seconds as the chunk length and use smart split, which moves each cut to the nearest pause so no word is chopped. Thirty to ninety seconds per sample suits every service in the table; if a service asks for a single file, skip this step and upload the cleaned file from step 3.
Listen, discard, upload
Play each part once. Discard anything with a click, a cough, a change in tone or a second voice. Upload the rest. If a service offers a "remove background noise" switch, leave it off for clean samples; it can soften consonants.
Quality over quantity
The most common mistake is padding the upload with weaker material to reach a duration. The clone averages everything it hears. Three excellent minutes make a better voice than twenty mixed ones.
How much to prepare
Instant clones
Two or three samples of about a minute each, from the same session, covering a question, a statement and something said with feeling. That is enough for ElevenLabs' instant tier and for most open models.
Professional clones
Thirty minutes to a few hours, split into 60 to 90 second samples, all from the same microphone and room. Include different moods and speaking speeds. Review every sample; at this volume one bad recording session can drag the whole model.
Mistakes that ruin a clone
- Processed audio. Heavy compression, EQ or "podcast enhancement" filters get baked into the voice.
- Music under the voice. Even quiet beds leak into the model.
- Mixed microphones. The phone memo and the studio session are two different voices to the model.
- Long silences. Wasted training time, and a clone that pauses oddly.
- Reverb. The room becomes permanent.
Consent
Only clone voices you have explicit permission to clone, including your own. Most services require you to confirm that, and several countries now regulate synthetic voices. Legitimate uses, accessibility, your own content, preserving a relative's voice with their agreement, are exactly what the technology is good for.
Frequently Asked Questions
Split your audio now
Free, private, no signup. Your file never leaves your browser.
Open the splitter