"AI voice" gets used as a catch-all term for two genuinely different technologies: text-to-speech using a pre-built voice library, and voice cloning that recreates a specific person's voice from a sample recording. They solve different problems, and the distinction matters more than it might seem — especially around what you're allowed to do with the output.
Text-to-speech with pre-built voices
This is what most free and general-purpose tools offer — including this one. A voice library is built by training a model on recordings from voice actors or licensed datasets, producing a fixed set of general-purpose voices you choose from. You don't provide a voice sample; you select from what's already available. The output sounds like "a voice," not "a specific identifiable person you supplied."
Voice cloning
Voice cloning takes a sample recording of a specific person's voice — sometimes just a few seconds, sometimes several minutes depending on the platform — and builds a model that can generate new speech in that same voice, saying things the person never actually said. This is a genuinely different capability, offered by platforms that specialize in it, and it comes with meaningfully higher stakes.
Why the distinction matters
Voice cloning raises consent and identity questions that pre-built-voice text-to-speech mostly doesn't. Cloning someone's voice without their permission — a public figure, a family member, a coworker — can range from a bad idea to outright illegal depending on jurisdiction and use, especially if the output is used to impersonate that person or deceive listeners about who's actually speaking. Most legitimate voice-cloning platforms require you to either use your own voice or provide proof of consent from the person whose voice you're cloning.
Pre-built-voice text-to-speech doesn't carry that specific risk, because the voice was never a real, identifiable individual to begin with — it's a synthesized voice built from training data, not a 1:1 recreation of one named person. That doesn't mean text-to-speech has zero ethical considerations (using AI narration without disclosing it in contexts where listeners would reasonably assume a human is speaking is its own separate question), just a different and generally lower-stakes set of them.
Which one you actually need
If you want narration for a video, podcast, or document, and you don't need it to sound like one specific real person, pre-built-voice text-to-speech is the simpler, more accessible option — free tools like this one exist specifically for that case. If your project genuinely requires recreating a specific individual's voice (with their consent), you need a dedicated voice-cloning platform built for that purpose, not a general text-to-speech tool.