Taletune

How to Record a Voice Sample That Actually Sounds Like You

September 8, 2025 · 5 min read

When a cloned voice sounds wrong, people blame the AI. Usually the AI did exactly what it was asked — it faithfully modelled a recording made in a kitchen, at eight in the evening, by someone doing a strange announcer impression because a microphone was on.

The sample is the whole ball game. Here is what actually moves the needle, roughly in order of impact.

1. Kill the background noise

This matters more than everything else combined.

A voice model cannot cleanly separate you from your environment. It learns the room. A dishwasher running two metres away becomes part of what the model thinks your voice is, and every story afterwards carries a faint ghost of it.

What to do:

  • Shut the door. A bedroom with soft furnishings is ideal — beds, curtains and carpet absorb reflections.
  • Turn off anything with a motor or a fan. Dishwasher, extractor, air conditioning.
  • Close the window. Traffic is worse than people expect.
  • Avoid the bathroom and the kitchen. Hard surfaces create echo, and echo is baked in permanently.

You do not need silence. You need the absence of specific competing sounds.

2. Read the way you read to your child

The moment a microphone appears, most people slow down, drop their pitch and start narrating like a documentary. The clone then sounds like that — a slightly pompous stranger who shares your accent.

The fix is a mental one. Picture your child next to you, half asleep, and read to that image. Slightly too casual is much better than slightly too performed.

If you catch yourself doing The Voice, stop and start again. It takes twenty seconds.

3. Hold the phone properly

A phone microphone is good. People use it badly.

  • A hand's width from your mouth. Closer and you get plosive thumps on every "p" and "b". Further and the room creeps in.
  • Slightly off to the side, not straight on. Your breath then misses the microphone.
  • Do not cover the bottom edge. On most phones the main microphone is there, and a thumb over it is the single most common cause of a muffled sample.
  • Take the case off if it is a thick one.

Earbuds are usually a downgrade, not an upgrade. Their microphones are tuned for phone calls — they compress heavily and cut low frequencies, which strips out exactly the warmth that makes a voice recognisable. The phone itself is better.

4. Give it a full minute if you can

Thirty seconds can produce a usable model. A minute produces a noticeably steadier one, particularly on unusual sentences.

More audio means more examples of how you handle different sounds, so the model guesses less. If the app shows you a script, read all of it rather than stopping at the point where it says you have enough.

5. Read the script as written

If there is a script on screen, read those words. Apps often run a speech-to-text check comparing what you said with what you were meant to say, as a proxy for "was this recording clear enough to be usable". Improvising fails that check even when the audio is fine.

If you stumble, most apps would rather you started again than pushed through — a clean sixty seconds beats ninety seconds with a cough in the middle.

6. Be in a normal state

A tired voice at midnight produces a tired-sounding clone. So does a cold, and so does the tail end of a shouting match about teeth-brushing.

You do not need to be fresh. You need to be somewhere near your usual self, because that is the voice your child recognises. If you have a cold, wait two days. The model you make in five minutes is the one you will hear for months.

A sixty-second routine

  1. Bedroom, door shut, everything with a motor turned off.
  2. Phone in hand, a hand's width away, slightly to one side, bottom edge clear.
  3. One practice read of the first line, out loud, to shake off the announcer voice.
  4. Record the whole script at your normal reading pace.
  5. Play it back. Ask one question: does this sound like me talking, or like me presenting?
  6. If presenting, do it once more. It will be better.

What still will not be perfect

Even with a flawless sample, expect the clone to be a slightly calmer version of you. The performance layer — the wolf voice, the dramatic pause, the whisper as they drift off — is where the gap remains.

That is worth knowing so you are not disappointed. What survives is the thing that matters most to a small child: the sound of you. Children consistently recognise it immediately, which is the entire point.

After you record

Two habits worth having:

Listen to a full story before you rely on it. A sample can sound fine in isolation and reveal a problem across three minutes of narration.

Re-record if your circumstances change. Recovered from a long illness, or made the first one in a noisy flat you have since moved out of? A fresh sample takes a minute and you can delete the old profile.

And if you would rather not keep a voice profile permanently, deleting and re-recording later is entirely reasonable — the privacy questions worth asking covers why some parents prefer to work that way.

Ready to try? Taletune is free on iOS and Android, and you will know within about two minutes whether the result sounds like you.

Read next