Prepare Audio for AI Voice Cloning: Clean Samples in Five Steps

Short answer

Voice cloning services want a few minutes of one clean voice: no music, no second speaker, no room echo, no long pauses. Take your best recording, cut out the clean passages, remove the pauses, split the result into 30 to 90 second pieces, listen once, and upload the pieces you would be happy to hear repeated a thousand times. All three steps run in your browser in ChunkAudio, and the audio never leaves your computer.

The clone can only be as good as the samples. Every service, whether ElevenLabs, Resemble AI, PlayHT or an open model, learns from what you give it, and it learns the flaws with the same enthusiasm as the voice: the hum of a fridge, a reverberant room, the co-host who laughs in the background. This guide is about getting those out before you upload.

What the services ask for

The numbers below are typical tiers as of September 2026. Check the current limits of the service you use; they change often.

ServiceInstant cloneProfessional cloneFormats
ElevenLabs1 to 5 minutes30 minutes to 3 hoursMP3, WAV
Resemble AIabout 1 minute25 sentences and upWAV, MP3
PlayHT30 seconds20 minutes and upMP3, WAV
Open models (XTTS, Tortoise)6 to 30 secondsas much clean audio as you haveWAV

Two rules hold across all of them: the sample should contain one voice, and every second of it should be speech you would accept as the voice's default sound.

What makes a good sample

  • One speaker, no music. Intros, jingles and guests must go. The model cannot separate them.
  • Dry room. Echo becomes part of the voice. A closet full of clothes beats a kitchen.
  • Same microphone and distance throughout. A mix of recordings from different setups produces a clone that sounds like two people.
  • Natural delivery with variety. Questions, statements, a laugh, a serious passage. Read-aloud text with a flat tone gives a flat clone.
  • No long pauses. Silence is training time wasted, and some services count it against the minimum.
  • Technical: 44.1 kHz or higher, 16-bit or better, mono. WAV if you have it; MP3 at 192 kbps or better is fine. Do not "upgrade" a low-bitrate MP3 by re-encoding it, that adds nothing.

The five steps in ChunkAudio

1

Pick the recording

Choose the recording with the best sound, not the longest one. A solo podcast episode, a narration, a voice memo recorded close to the microphone in a quiet room. Five clean minutes beat thirty mediocre ones.

2

Cut out the clean passages

Open the file in cut mode, drag the two handles around a stretch where only the target voice speaks, press play to check it, and save the selection. Repeat for each clean passage. The cut is sample-exact on WAV, and for MP3 the frames are copied rather than re-encoded, so the sound is untouched. More on cutting in the audio cutter guide.

3

Remove the pauses

Switch to remove pauses. It finds every silence longer than the limit you set, one second is a good default for voice samples, and shortens each to a natural gap. You get one file of continuous speech, which is what the models want. The silence remover page explains the settings.

4

Split into samples

Back in split mode, set 60 seconds as the chunk length and use smart split, which moves each cut to the nearest pause so no word is chopped. Thirty to ninety seconds per sample suits every service in the table; if a service asks for a single file, skip this step and upload the cleaned file from step 3.

5

Listen, discard, upload

Play each part once. Discard anything with a click, a cough, a change in tone or a second voice. Upload the rest. If a service offers a "remove background noise" switch, leave it off for clean samples; it can soften consonants.

Quality over quantity

The most common mistake is padding the upload with weaker material to reach a duration. The clone averages everything it hears. Three excellent minutes make a better voice than twenty mixed ones.

How much to prepare

Instant clones

Two or three samples of about a minute each, from the same session, covering a question, a statement and something said with feeling. That is enough for ElevenLabs' instant tier and for most open models.

Professional clones

Thirty minutes to a few hours, split into 60 to 90 second samples, all from the same microphone and room. Include different moods and speaking speeds. Review every sample; at this volume one bad recording session can drag the whole model.

Mistakes that ruin a clone

  • Processed audio. Heavy compression, EQ or "podcast enhancement" filters get baked into the voice.
  • Music under the voice. Even quiet beds leak into the model.
  • Mixed microphones. The phone memo and the studio session are two different voices to the model.
  • Long silences. Wasted training time, and a clone that pauses oddly.
  • Reverb. The room becomes permanent.

Consent

Only clone voices you have explicit permission to clone, including your own. Most services require you to confirm that, and several countries now regulate synthetic voices. Legitimate uses, accessibility, your own content, preserving a relative's voice with their agreement, are exactly what the technology is good for.

Frequently Asked Questions

How much audio do I need?
One to three clean minutes for an instant clone. Thirty minutes or more of varied, same-session audio for a professional clone. In both cases quality matters more than length.
Can I use podcast audio?
Yes, if you can isolate passages with only the target voice. Cut out the intro music, ads and guest segments first. Solo episodes work best; interviews take more cutting.
Which format should I upload?
WAV at 44.1 kHz or higher if the recording exists as WAV. An MP3 at 192 kbps or better is fine, and cutting it in ChunkAudio does not re-encode it. Converting a low-bitrate MP3 to WAV does not improve it.
Does my recording get uploaded anywhere while I prepare it?
Not by ChunkAudio. Cutting, pause removal and splitting run inside your browser; the file stays on your computer until you upload the finished samples to the cloning service yourself.
Do I need professional equipment?
No. A decent USB microphone in a quiet, soft-furnished room gives excellent samples. Consistency and silence matter more than the microphone's price.

Split your audio now

Free, private, no signup. Your file never leaves your browser.

Open the splitter

Tim Heineccius

Builds ChunkAudio in Bonn, Germany. Writes about splitting, transcribing and cleaning up audio without uploading it anywhere.