How to use ElevenLabs for voiceovers (2026)

Runbo Li
Runbo Li
·
· 5 min read
elevenlabs

Quick answer

To use ElevenLabs for a voiceover, open Text to Speech, paste a short representative script, choose a voice and model, generate a sample, then correct pronunciation and delivery before producing the full recording. Download the approved audio in the format your editor or application needs. Use voice cloning only for a voice you are authorized to reproduce.

ElevenLabs now spans text-to-speech, voice design and cloning, dubbing, conversational agents and other media workflows. This guide stays focused on producing and exporting narration. Product labels, model availability, credits and commercial terms can change; check the live interface, current documentation and pricing page before a large job.

ElevenLabs voiceover workflow

  1. Define the deliverable. Set language, duration, audience, tone, required file format and commercial use.

  2. Prepare a short test script. Include names, numbers, acronyms and emotional range that represent the full project.

  3. Choose a voice. Start with a default, library, designed or authorized cloned voice that matches the language and accent.

  4. Choose a model. Match expressiveness, language coverage, consistency and latency to the job.

  5. Generate a short sample. Do not spend the full script until the difficult words and delivery work.

  6. Correct the input and settings. Rewrite awkward phrasing, spell out ambiguous numbers and adjust only the controls that address a heard problem.

  7. Generate and inspect the final audio. Listen from beginning to end for omissions, artifacts and inconsistent delivery.

  8. Export and document. Save the approved audio, script, voice, model, settings, date and usage terms with the project.

Run a same-script voice test

Generate the same short script in a few current voices. Compare pronunciation, pacing, commercial terms and the finished-video workflow before committing.

Try AI Voice Generator

1. Prepare text for speech

Write in the words the listener should hear. ElevenLabs warns that unusual names, abbreviations, symbols, digits and emoji can destabilize pronunciation, especially in multilingual work. Spell out ambiguous numbers, dates and acronyms, and use punctuation and paragraph breaks to express structure.

Test the hardest 15–30 seconds first. Include the brand name, product terms, a number, a question and the strongest emotional line. A neutral easy sentence can make the wrong voice or model appear acceptable.

2. Choose the voice before fine-tuning settings

ElevenLabs’s TTS product guide says voice selection has the largest effect on output, followed by model selection and then settings. Browse the available voice types and verify that the voice’s permitted use fits the project.

ElevenLabs’s current voice documentation also says its Default voices will expire on December 31, 2026. If a production depends on one of those voices, migrate to a persistent replacement and approve the new sound before the deadline.

3. Choose a current model by the job

Model family

Use when

Check before finalizing

Eleven v3

Expressive or dramatic multilingual delivery and multi-speaker dialogue

Current character limit, tag behavior and consistency on long scripts

Eleven v3 Conversational

Expressive real-time speech

Latency and turn-taking in the actual application

Multilingual v2

Longer-form, stable multilingual narration

Language, accent and pronunciation on the chosen voice

Flash v2.5

Low-latency or cost-sensitive generation

Whether speed preserves the required delivery quality

These are the model families documented on September 13, 2026. Do not reuse an old model recommendation without checking the current model picker and documentation.

4. Adjust settings only after listening

ElevenLabs documents speed values from 0.7 to 1.2 and warns that extreme values can affect quality. Stability and similarity behavior depends on the model and voice. Generate a baseline first, then change one setting at a time so you can hear what fixed or degraded the sample.

  • Wrong pronunciation: rewrite the word, spell out the acronym or use a supported pronunciation control.

  • Flat delivery: improve the script and punctuation before adding stronger style controls.

  • Unstable emotion: test a better-matched voice or a more stable model.

  • Wrong accent: choose a voice recorded or designed for the target language and region.

  • Artifact or omission: regenerate the short segment and inspect the input before rebuilding the whole file.

5. Generate, review and export

Generate in manageable sections when the project is long, but preserve enough neighboring context for natural transitions. Listen across every join, check loudness and room tone, and keep the uncompressed or highest-quality approved master before making delivery copies.

Choose MP3 for common playback workflows or an available PCM/WAV-quality option when the next step is editing. Telephony and streaming applications may require other formats. Export options and higher-quality availability depend on the live plan and endpoint.

How to clone a voice responsibly

ElevenLabs’s voice-cloning documentation distinguishes Instant Voice Cloning from Professional Voice Cloning and explains that poor, short or noisy samples reduce quality. It describes verification as a safeguard, while placing responsibility for authorized use on the creator.

  • Get explicit authorization. State the intended content, channels, duration and whether the voice can be reused or sublicensed.

  • Record clean samples. Use one speaker, consistent microphone placement, low noise and the emotional range needed in output.

  • Verify the final script. A clone can make the speaker appear to say words they never recorded. Approval must cover the generated message, not only the source sample.

  • Protect access. Limit who can use the voice, keep credentials private and remove access when the project ends.

  • Disclose when required. Follow the applicable platform, advertising, contractual and jurisdictional rules.

Using the ElevenLabs API

The current ElevenLabs API quickstart shows the supported SDK flow: store the API key in an environment variable, initialize the client, select a voice ID, model ID and output format, then call text-to-speech conversion. Do not place the key in source code, screenshots, client-side bundles or published examples.

For production, log the model and voice identifiers, version the script, handle rate and quota errors, and validate the returned audio before marking a job complete. A successful API response does not prove the pronunciation, content or rights are correct.

When Magic Hour fits the workflow

Use Magic Hour’s AI voice generator when you want to compare narration in the same broader creation stack. Use voice cloning for an authorized custom voice, AI Talking Photo when the output should animate a still portrait, or Lip Sync when approved audio needs to drive existing video.

Frequently asked questions

Open Text to Speech, paste a short script, choose a voice and model, generate a sample, correct pronunciation and delivery, then generate and export the full recording.

Use Eleven v3 for expressive performance, v3 Conversational for real-time speech, Multilingual v2 for stable longer narration, or Flash v2.5 when low latency or generation cost matters. Verify the current product documentation because model availability changes.

Only when you have the authorization and rights required for that voice, message and use. Platform verification does not replace the speaker’s permission or your legal and contractual obligations.

Commercial use depends on the current plan terms and your rights to the script, voice and other inputs. ElevenLabs’s documentation says commercial usage rights are available with paid plans; recheck the terms for the account and project before publishing.

Use the ElevenLabs alternatives comparison to compare voice cloning, narration, API and finished-video workflows, or the broader AI voice-generator guide for category-level selection. The marketing-video workflow explains how to combine narration with visuals and captions. When the narration is approved and the remaining job is placing it in a finished clip, follow how to add an AI voiceover to video.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

How to Lip Sync a Video With AI
Recommended next
How to lip sync a video with Magic Hour: 5 steps

Lip sync a video in Magic Hour in five steps. See current free limits, accepted files, timing checks, common failures and when to use Talking Photo instead.

Prompting AI Videos Cover
How to prompt AI videos: a practical 10-step guide
Picture of Multiple Faces
How to swap faces: photo and video guide for 2026
veo3
Google Veo 3.1: a beginner's guide to AI video
Runway
Runway Gen-4 guide: 5/10-second image-to-video prompts