Voice Cloning in Poko Motion: Narrate Your Product Videos in Your Own Voice
Record a 10-second sample and every video you generate can sound like you, not a generic AI narrator.

Most AI video tools give you one narrator. It's a competent, slightly generic voice. And it starts to sound identical across every product on the internet the moment you've heard it on a few too many demo videos. That's despite voice cloning technology itself having gotten remarkably good over the last couple of years - most tools just don't put a cloning workflow in front of you by default.
Poko Motion lets you skip that voice entirely. Use your own, or a colleague's with their permission. It takes about as long as reading a paragraph out loud.
Here's what voice cloning actually is in Poko, how the recording works, and where the cloned voice ends up once you've saved it.
What voice cloning means here
Voice cloning in Poko is not a professional studio process. You record a short sample of yourself speaking, and Poko turns that sample into a reusable narration voice.
From then on, you can select it for any video project. It works the same way you'd pick a stock voice from the catalog. There's no separate rendering step and no waiting period - once it's saved, it's just another option in the voice picker.
Step 1: Pick a vibe before you record
Rather than asking you to read something neutral and hoping the clone generalizes, Poko has you choose a "vibe" first:
- Neutral - natural, balanced narration
- Energetic - high-energy, punchy delivery
- Calm - soft, soothing, relaxed
- Friendly - warm, conversational, approachable
- Authoritative - confident, precise, expert tone
- Excited - enthusiastic, upbeat, hyped
- Angry - sharp, intense, frustrated
- Whisper - soft, intimate, secretive
- Storyteller - expressive, dramatic, cinematic
- Sales pitch - persuasive, polished, salesy
Each vibe comes with its own short script written for that tone. The energetic script pushes you toward an upbeat, punchy delivery. The calm one reads more like a guided breathing exercise.
This matters because a clone trained on a flat, neutral reading tends to sound flat no matter what text you feed it later. Picking the vibe that matches how you actually want your videos to sound gets you a clone with the right energy from the start.
If you're recording in a language other than English, the vibe scripts step aside. You just get simple instructions: speak naturally for 10 to 12 seconds, use a full sentence or two, and keep the room quiet.
Step 2: Record 10 to 12 seconds
The recorder has one hard requirement: a minimum of 10 seconds of audio before you're allowed to save. A live timer counts up while you talk, and a "Keep going…" prompt lets you know when you're not quite there yet.
Once you cross the threshold, it flips to "Ready," and you can play back the recording before committing to it. There's no multi-take wizard or lengthy studio session - a phone-quality reading of one short paragraph is enough to get started.
Poko's recorder supports more than 40 languages, including:
- English, Spanish, French, German, Italian, Portuguese
- Hindi, Tamil, Telugu, Bengali, Gujarati, Marathi
- Japanese, Korean, Chinese, Vietnamese, Thai, Indonesian
- Arabic, Hebrew, Turkish, Russian, Ukrainian, Polish
Step 3: Name it, confirm consent, save
Before the clone saves, you name it so you can tell it apart from other voices later. You also check a consent box, and the text is specific on purpose. You're confirming that:
- This is your own voice, or you have explicit written permission from whoever's voice you're cloning
- You will not use it to impersonate anyone without their consent
- You understand that violating this can get your account terminated
This isn't a formality Poko buries in the terms of service. It's a checkbox you have to actively tick every time, which is the right amount of friction for a feature that can otherwise be misused.
Where the cloned voice actually shows up
Once saved, your cloned voice appears in the Voice picker inside the Studio sidebar. It sits right next to:
- Poko's built-in voice catalog
- Any ElevenLabs or Cartesia voices you've unlocked with your own API key
From there, selecting it works exactly like selecting any other narrator. It becomes the voice used for the script the AI agent writes when it builds your video. You don't need to re-record for every project, and you can swap back to a stock voice at any time without losing the clone.
Adding your own ElevenLabs or Cartesia API key in settings is optional. It simply unlocks that provider's separate voice catalog in the same picker - it doesn't replace or gate Poko's own cloning pipeline, which works without any external account. If you're already using BYOK for agent usage, it's worth knowing voice narration is billed separately from music and sound effects under that same billing toggle.
Why this is worth the two minutes it takes
A launch video, an onboarding walkthrough, and a changelog update all sound more like they came from the same company when the same voice narrates all three.
That consistency used to mean hiring a voiceover artist for every batch of content, or living with whatever the AI provider's default voice happened to sound like.
Cloning your own voice once, and reusing it across every video Poko generates afterward, removes both problems. It's still automated end to end, but the result sounds like your team, not a stock library.
If you've been putting off recording narration because a real voiceover felt like overkill for a quick product update, this is the middle ground. Ten seconds of audio, a script you can actually read comfortably, and a narrator that's yours for every video after.
FAQs
As little as 10 to 12 seconds of clear speech. Poko's recorder times you and only lets you save once you've hit the minimum, so there's no guessing whether the sample is long enough.
