

To dub a video with AI, start with a final source video and permission to use the footage and voice. Generate a transcript, translate the meaning, create the replacement speech, align it to the speaker, add matching captions, and have a fluent reviewer check the full export. Magic Hour can combine transcription, translation, voice preservation and lip alignment in one dubbing workflow.
Product capabilities were checked September 13, 2026. Magic Hour currently lists AI Video Dubbing in 53 languages with voice cloning and lip sync included. Language availability and account limits can change, so confirm the choices shown in the product before committing to a delivery.
Upload one authorized source video, choose a target language, then review the translated meaning, names, numbers, voice, timing, lip alignment and captions before publishing.
Try AI Video DubbingTask | What changes | Use it when | Acceptance check |
|---|---|---|---|
AI video dubbing | Spoken language, voice track and visible mouth timing | A visible speaker should deliver the same message in another language | Meaning, voice, timing and mouth movement all pass review |
Video translation | The spoken message is converted to another language | Language conversion is the main job | Names, numbers, claims and intent remain correct |
Voiceover | A new audio track is added or substituted | The speaker is off-camera or mouth matching is unnecessary | Speech is clear and timed to the edit |
Lip sync | Visible mouth movement is aligned to prepared audio | You already have an approved replacement track | The face remains stable and sync looks natural at normal speed |
Captions | Readable on-screen text is added | Viewers need text access or may watch without sound | Text matches the final spoken track and is readable |

Use the highest-quality master available, with clear dialogue and a stable view of the speaker. A single visible speaker and limited background noise are the safest starting point. Finish picture edits first so later cuts do not invalidate the dubbed timing.
Confirm permission for the footage, original performance and any voice cloning. Write down the target audience, language, region, required terminology and words that must stay unchanged.
In AI Video Dubbing, upload the video and choose the target language. The current workflow transcribes the dialogue, translates it, recreates the speaker’s voice and applies lip alignment. If language conversion is the reader’s main concern, the AI Video Translator page describes the same output from a translation-first perspective.
Generate one representative segment before processing a library. Include the hardest material: a brand name, a number, a fast sentence, a visible close-up and a culturally specific phrase.
Compare the source transcript, target script and rendered speech. A fluent reviewer should check meaning, tone, names, numbers, units, product claims, calls to action and pronunciation. A natural-sounding sentence can still be factually wrong.
When a sentence expands in the target language, shorten it without removing a qualification or changing the claim. Keep a glossary for product names and repeated technical terms so every video uses the same wording.
Watch the entire export at normal speed, then recheck close-ups. Reject a version if the voice changes identity unexpectedly, words clip, pauses land on the wrong edit, the mouth drifts from the speech, teeth or face shape distort, or background audio masks the dialogue.
If you already have an approved translated track, use the dedicated Lip Sync workflow to align that audio to a visible face. Compare current alternatives in the AI lip-sync tools guide when the source contains multiple speakers, animation or other requirements outside the direct workflow.
Generate captions after the spoken track is final, then correct them against what viewers actually hear. The Magic Hour Subtitle Generator can create a starting transcript. W3C accessibility guidance explains that captions should include speech and meaningful non-speech audio needed to understand the video.
Check line breaks, reading speed, speaker changes, sound labels, contrast and placement on the intended phone or player. Keep captions clear of faces, product details and interface controls.
Keep the same source edit, offer and destination when comparing language versions. Record the source file, target language, glossary, reviewer, approved export and publication date. Measure qualified viewing, completion, destination actions and conversion for each audience; do not attribute a difference to dubbing when the offer or distribution also changed.
Use separate tools when a human translator must approve the script before audio generation, when you need a specific authorized synthetic voice, or when the target track must be edited in a timeline. Generate speech with an AI voice generator or an authorized voice clone, edit it against the picture, then apply lip sync only after the audio is locked.
Segment by sentence or shot rather than generating one long track. Preserve timecode references and join segments with room tone and consistent loudness. This makes a single correction cheaper and avoids regenerating an otherwise approved video.

Meaning: every sentence preserves the source intent and required qualifications.
Facts: names, numbers, dates, units, prices and product claims match the approved source.
Language: a fluent reviewer accepts grammar, pronunciation, register and regional wording.
Voice: the voice is authorized, intelligible and consistent across the full video.
Timing: speech begins and ends at sensible visual moments without clipped words.
Lip alignment: visible speech looks credible at normal speed and the face remains stable.
Mix: dialogue remains clear against music, effects and room tone.
Captions: text matches the final audio and remains readable in the final player.
Failure | Likely cause | Smallest useful fix |
|---|---|---|
Wrong name, number or claim | Transcript or translation error | Correct the approved script and regenerate only that segment |
Speech overruns the shot | Target sentence is longer | Shorten the wording without changing meaning, or retime the edit |
Mouth movement drifts | Audio timing or face visibility is poor | Align the track first; use a clearer shot or split the segment |
Voice changes between lines | Segments use inconsistent settings or source | Reuse the same approved voice setup and reference |
Dialogue is hard to hear | Music or effects compete with speech | Lower or duck the background mix and review on phone speakers |
Captions disagree with speech | Captions came from an earlier script | Regenerate from the final audio and proofread |
For a pilot, total the platform charge, translation review, audio editing, caption review and failed attempts. Then divide the total by approved delivered minutes. A low generation price can still produce an expensive workflow if linguistic or visual failures require repeated work.
Accepted-minute cost = (generation + human language review + editing + caption review + retry cost) ÷ approved delivered minutes. Use the current account price and actual project records rather than a generic market estimate.
Yes. A one-step dubbing workflow can transcribe, translate, generate replacement speech and align visible mouth movement. A fluent human still needs to approve the meaning, terminology, pronunciation and final audiovisual result.
Use lip sync when a visible speaker’s mouth should match the replacement speech. For narration over slides, screen recordings or off-camera footage, a well-timed voiceover may be enough.
Yes. Use an authorized synthetic voice or a recorded voice actor, then edit and synchronize that track. Decide whether preserving the original speaker or choosing a new narrator better fits the audience and permissions.
Approve one source master and glossary, run the hardest target language as a pilot, then repeat the same review record for each language. Do not assume approval in one language proves the others are correct.
Captions make the final spoken track available as text and can include meaningful non-speech audio. Create them from the approved dubbed audio so the words match what viewers hear.

