Narration voices
Choosing, previewing, and tuning the voice that carries your video.
Narration typically covers most of a video's runtime, so the voice matters more than any other single setting.
Voice providers
VidNext offers voices from several providers, each with different strengths:
| Provider | Character |
|---|---|
| ElevenLabs | Premium emotional voices |
| Cartesia | Ultra-fast generation |
| Fish Audio | Expressive, natural voices |
| MiniMax HD | Natural voices |
| Gemini TTS (beta) | LLM-driven performance with audio-tag direction |
For non-English videos the provider list narrows to those with strong support for the selected language.

Star any voice to collect it in Favorites for quick reuse across projects:

You can also clone a voice from about 10 seconds of clear speech — useful for keeping the same host across every video on a channel:

Previewing voices
Never commit to a voice unheard. Select one and generate a test clip using one of four preset snippets — Hook, Statistics, Emotional, High energy — or paste your own line. The presets are designed to expose how a voice handles the moments that matter: openings, numbers, emotional beats, and hype.
Voice tests are free or nearly free (Gemini TTS shows a small per-test credit charge after your free tests run out); failed generations are never charged.
Speed and pacing
Narration speed is adjustable — Slowest, Slow, Normal, Fast, Fastest. VidNext also calibrates timing to the actual voice: scene durations follow the real narration audio, not an assumed reading speed, so pacing stays natural at any setting.

Using your own API key
If you have your own ElevenLabs, Fish Audio, or Cartesia account, connect your API key in settings:
- 50% off the TTS portion of every generation
- Access to your own cloned voices
- Usage runs against your own provider quota

A "My key · 50% OFF" chip appears on the voice section when active.
Uploading your own narration
Skip TTS entirely by choosing Upload my narration as the narration source — bring your own voiceover (up to 30 minutes total, multiple files allowed) and VidNext builds the visuals around it. There's also a Music video mode: upload a song and the track becomes the full audio, with visuals cut to it.


Changed your mind after generation? Any narration segment can be given a different voice and regenerated from the editor — no need to redo the whole video.
