Magic Hour
  • Pricing
Video

Start your first video in under 60 seconds

Generate or edit video, image, and audio - free to start.
Start Creating Free
No credit card requiredFree daily creditsNo signup required
Join the Discord
Video
Company
PricingAboutBlogChangelogAPISkillsAll ToolsTemplatesAI ModelsTrust & Data UsePrivacy PolicyTerms of ServiceRefund Policy
Video Products
AI Avatar GeneratorAI UGC Ad GeneratorAI Video DubbingAI Video EditorAI Video ExpanderAI Video ExtenderAI Video TranslatorAI Video UpscalerAnimationAudio-to-VideoCharacter ReplaceColor GraderFace Swap VideoImage-to-VideoLip SyncMusic Video GeneratorSubtitle GeneratorTalking PhotoText-to-VideoVideo ColorizerVideo-to-Video
Image Products
AI Clothes ChangerAI Face EditorAI GIF GeneratorAI Headshot GeneratorAI Image EditorAI Image ExpanderAI Image GeneratorAI Image UpscalerAI Influencer GeneratorAI Meme GeneratorAI Selfie GeneratorAI Storyboard GeneratorBackground RemoverBody SwapFace Swap PhotoGenerative FillHead SwapPhoto ColorizerQR Code GeneratorCharactersMoodboards
Audio Products
AI Audio TranslatorAI Music GeneratorAI Sound Effect GeneratorVideo-to-AudioVoice ChangerVoice ClonerVoice Generator
Support
CommunityFAQHelp CenterContact UsStatus
Social
Instagram
X
TikTok
Facebook
YouTube
LinkedIn
support@magichour.ai
Backed byCombinator

© 2026 Magic Hour AI, Inc.

Back to Blog
  1. Blog
  2. Content Creation

How to make a realistic talking AI avatar

Runbo Li
Runbo Li
·
CEO of Magic Hour
·
Jul 28, 2025· 4 min read
AI Summary:
ChatGPTClaudeGeminiPerplexity
AI Avatar Gen

Contents

Create with Magic Hour
Make videos and images with AI.

Quick answer

To make a realistic talking AI avatar, use one clear portrait you have permission to animate, prepare clean approved speech, generate a short proof in Magic Hour Talking Photo, and inspect the complete result before extending the script. If the source is already a video, use Lip Sync instead. Realism depends on the portrait, audio, selected mode and the review standard; no tool makes every image convincing.

The current Magic Hour workflow and limits below were checked September 13, 2026. The focused make-a-photo-talk guide covers the simpler one-off task; this guide concentrates on building a repeatable presenter and quality-control process.

1. Decide whether you need a talking photo or reusable avatar

  • Talking photo: one portrait performs one supplied audio track. Use it for a proof, greeting, lesson excerpt or short presenter shot.

  • Reusable avatar: the same approved identity, voice, framing, wardrobe and delivery recur across scripts. This requires a documented asset and review system.

  • Lip sync: existing video is synchronized to new approved audio; the source already contains motion.

  • Image to video: animates a still more generally and may not be designed around speech synchronization.

2. Secure the face and voice rights

Record who owns the portrait and audio, who is depicted, what use was approved, where it may be published, whether the voice may be cloned and when permission expires. A public photo or recording is not proof of permission. Do not imply that a person endorsed a product or made a statement they did not approve.

For an original synthetic voice, use AI Voice Generator. Use Voice Cloner only with authorized samples and a permitted purpose. Keep the approved script, source files and consent record with the project.

3. Choose a portrait that can survive animation

  • Use one visible face. Start front-facing or only slightly turned, with the eyes and mouth unobstructed.

  • Give the face enough pixels. Avoid a tiny subject, heavy compression, motion blur or an extreme crop.

  • Use stable lighting. Strong shadows across the mouth, glasses glare and clipped highlights make review harder.

  • Keep the frame plausible. Include enough head and shoulder area for the selected motion mode; hands or props near the face can create occlusion errors.

  • Preserve identity intentionally. Define the details that must not change: age, facial structure, skin detail, hairline, marks, clothing and brand assets.

4. Prepare the speech before generating video

Record or generate the final voice first. Use a quiet source without music or echo, verify names and numbers, and read the script aloud for pace. Split a long message at natural sentence or scene boundaries so a correction does not require regenerating the entire piece.

5. Generate a short proof in the current tool

The current Talking Photo Help Center guide says to add one image and audio, then choose Realistic or Expressive. Realistic is the longer lip-sync-oriented mode and currently accepts 0.5 to 300 seconds; Expressive supports prompt-guided movement up to 45 seconds. Those are input limits, not guarantees that every long clip will pass quality review.

  • Realistic: start here when mouth timing and natural facial detail matter most. It does not currently support 1080p.

  • Expressive: use when broader prompted expression or motion helps and the shorter duration fits.

  • Guest browser test: the public product page currently offers three talking-photo generations per day, up to five seconds, without signup. Plan, resolution and watermark rules differ after signup.

For a first proof, use a sentence containing the hardest name, number, plosive sound, smile or pause in the real script. Generate that before investing in the full narration.

6. Review the avatar like an editor

  • Speech: exact words, pronunciation, emphasis, pace, pause and audio-video timing.

  • Mouth: closure on M/B/P sounds, teeth, lips, corners and transitions between phonemes.

  • Face: identity, eye direction, blinks, skin texture, jaw, hairline, glasses and occlusions.

  • Motion: head and body movement that fits the words without loops, drift or abrupt resets.

  • Frame: background stability, crop, aspect ratio, resolution, captions and room for graphics.

  • Whole clip: watch once at normal speed, then inspect failure timestamps frame by frame.

7. Fix the input before adding more effects

  • Poor mouth sync: use cleaner speech, remove music, shorten the proof or try the more lip-sync-oriented mode.

  • Identity drift: choose a clearer portrait, reduce extreme movement and compare the face at the start, middle and end.

  • Stiff delivery: revise the voice performance first; punctuation and pauses often matter more than extra visual effects.

  • Unnatural motion: simplify the prompt, shorten the shot or use Realistic when expressive movement is not required.

  • Long-form monotony: write scenes and cutaways rather than forcing one face to carry the entire video.

8. Disclose and publish in context

Follow the destination’s current rules. YouTube’s AI-content disclosure guidance explains how “Made with AI” disclosures can appear, while its privacy guidance for synthetic likenesses describes requests involving realistic altered or synthetic depictions. Disclosure does not replace permission, and permission does not make a false claim accurate.

Before publishing, confirm the speaker identity, script approval, disclosure, captions, brand claim, destination policy and the final file. Save the exact source image, audio, mode, settings, output and reviewer decision so later versions can be reproduced.

Make the first proof clip

Start with one clear portrait and a short approved recording. Generate a proof clip, review difficult sounds and facial motion, then decide whether the workflow is ready for a longer script.

Open Talking Photo

Frequently asked questions

A clear permitted portrait, clean speech, plausible movement and a strict review matter more than a generic realism label. Test the hardest sentence and reject identity drift, mouth errors or movements that do not fit the delivery.

The signed-in Help Center currently documents up to 300 seconds for Realistic and 45 seconds for Expressive. Guest use is currently limited to three five-second generations per day. Check the live tool because mode, plan and resolution limits can change.

Commercial permission depends on your rights to the face, voice, script and other assets plus the provider and destination terms. Paid Magic Hour terms may permit commercial output, but they do not grant rights to someone else’s likeness or recording.

Usually, build a sequence of reviewable shots. Cutaways, screen recordings, examples and graphics can carry information while reducing the burden on one continuous generated performance.

Runbo Li
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
See all articles

Continue Reading

How to Lip Sync a Video With AI
Recommended next
Videos
How to lip sync a video with Magic Hour: 5 steps

Lip sync a video in Magic Hour in five steps. See current free limits, accepted files, timing checks, common failures and when to use Talking Photo instead.

Jun 24, 2026
AI Talking Photo Tools (2026)
Images
Best AI talking photo tools (2026): 5 portrait workflows
May 16, 2026
Magic Hour editorial collage comparing AI lip sync with phoneme drawings and a speech waveform
Content Creation
7 best AI lip sync video tools (2026 comparison)
Jul 06, 2025
The Top AI Talking Photo Platforms
Images
7 best AI talking-photo generators in 2026
Aug 25, 2025
Top 6 Best Talking Photo APIs
Images
5 best talking-photo APIs and endpoints for developers
Dec 27, 2025