VidNextDocs
Creating videos

Narration voices

Choosing, previewing, and tuning the voice that carries your video.

Narration typically covers most of a video's runtime, so the voice matters more than any other single setting.

Voice providers

VidNext offers voices from several providers, each with different strengths:

ProviderCharacter
ElevenLabsPremium emotional voices
CartesiaUltra-fast generation
Fish AudioExpressive, natural voices
MiniMax HDNatural voices
Gemini TTS (beta)LLM-driven performance with audio-tag direction

For non-English videos the provider list narrows to those with strong support for the selected language.

The voice picker: provider tabs, a 700-voice searchable library with previews, favorites and cloned voices

Star any voice to collect it in Favorites for quick reuse across projects:

The Favorites tab with a starred voice

You can also clone a voice from about 10 seconds of clear speech — useful for keeping the same host across every video on a channel:

The My Cloned Voices tab with the clone-a-voice flow

Previewing voices

Never commit to a voice unheard. Select one and generate a test clip using one of four preset snippets — Hook, Statistics, Emotional, High energy — or paste your own line. The presets are designed to expose how a voice handles the moments that matter: openings, numbers, emotional beats, and hype.

Voice tests are free or nearly free (Gemini TTS shows a small per-test credit charge after your free tests run out); failed generations are never charged.

Speed and pacing

Narration speed is adjustable — Slowest, Slow, Normal, Fast, Fastest. VidNext also calibrates timing to the actual voice: scene durations follow the real narration audio, not an assumed reading speed, so pacing stays natural at any setting.

The narration speed selector

Using your own API key

If you have your own ElevenLabs, Fish Audio, or Cartesia account, connect your API key in settings:

  • 50% off the TTS portion of every generation
  • Access to your own cloned voices
  • Usage runs against your own provider quota

The API Keys settings with Fish Audio, ElevenLabs, and Cartesia slots

A "My key · 50% OFF" chip appears on the voice section when active.

Uploading your own narration

Skip TTS entirely by choosing Upload my narration as the narration source — bring your own voiceover (up to 30 minutes total, multiple files allowed) and VidNext builds the visuals around it. There's also a Music video mode: upload a song and the track becomes the full audio, with visuals cut to it.

The three audio modes: AI script & voice, Upload my narration, Music video

Upload my narration mode with the audio file upload zone

Changed your mind after generation? Any narration segment can be given a different voice and regenerated from the editor — no need to redo the whole video.

On this page