Magic Hour
  • Pricing
Video

Start your first video in under 60 seconds

Generate or edit video, image, and audio - free to start.
Start Creating Free
No credit card requiredFree daily creditsNo signup required
Join the Discord
Video
Company
PricingAboutBlogAPIAll ToolsTemplatesAI ModelsPrivacy PolicyTerms of ServiceRefund Policy
Video Products
AI Avatar GeneratorAI UGC Ad GeneratorAI Video DubbingAI Video EditorAI Video ExpanderAI Video ExtenderAI Video TranslatorAI Video UpscalerAnimationAudio-to-VideoCharacter ReplaceColor GraderFace Swap VideoImage-to-VideoLip SyncMusic Video GeneratorSubtitle GeneratorTalking PhotoText-to-VideoVideo ColorizerVideo-to-Video
Image Products
AI Clothes ChangerAI Face EditorAI GIF GeneratorAI Headshot GeneratorAI Image EditorAI Image ExpanderAI Image GeneratorAI Image UpscalerAI Influencer GeneratorAI Meme GeneratorAI Selfie GeneratorAI Storyboard GeneratorBackground RemoverBody SwapFace Swap PhotoHead SwapPhoto ColorizerQR Code GeneratorCharactersMoodboards
Audio Products
AI Audio TranslatorAI Music GeneratorVideo-to-AudioVoice ChangerVoice ClonerVoice Generator
Support
CommunityFAQHelp CenterContact UsStatus
Social
Instagram
X
TikTok
Facebook
YouTube
LinkedIn
support@magichour.ai
Backed byCombinator

© 2026 Magic Hour AI, Inc.

Back to Blog
  1. Blog
  2. Videos

6 best lip-sync APIs: inputs, costs and integration

Runbo Li
Runbo Li
·
CEO of Magic Hour
·
Dec 13, 2025· 8 min read
AI Summary:
ChatGPTClaudeGeminiPerplexity
 5 Best Lip Sync APIs for Production-Grade AI Video

Contents

Create with Magic Hour
Make videos and images with AI.

For matching new audio to an existing video, start with Magic Hour, Sync or HeyGen. Magic Hour is useful when lip sync sits alongside other image and video tools; Sync offers several specialist lip-sync models; HeyGen exposes dedicated dialogue-replacement and translation workflows. Choose D-ID or Tavus for a live talking avatar, and Replicate when you want to compare hosted models through one provider.

The first decision is your input: an existing video, a still portrait or a live conversation. These require different endpoints. A service that animates an avatar is not automatically a drop-in API for changing the dialogue in your footage.

Best lip-sync APIs compared

Provider

Best workflow fit

Input and delivery

Billing detail to check

magic hour logoMagic Hour

Recorded-video lip sync within a broader media workflow

Video plus audio; asynchronous output file. Still images use the separate Talking Photo API

Output frames and selected generation mode

Sync logoSync

Choosing among specialist models for existing footage

Video plus audio; asynchronous jobs, SDKs and webhooks

Model-specific frame rate, subscription and usage charges

HeyGen logoHeyGen

Dialogue replacement plus optional translation workflows

Dedicated video-plus-audio lip-sync endpoint; separate avatar and translation endpoints

Current API operation and mode, separate from website subscriptions

D-ID logoD-ID

Photo avatars or live agents that speak supplied audio

Photo-avatar video jobs; separate real-time agent sessions

Avatar/version, session or video allowance and account limits

Tavus logoTavus

A conversational avatar with your own voice or application logic

CVI conversation sessions, including an Echo mode

Conversation plan, usage and integration mode

Replicate logoReplicate

Comparing hosted lip-sync and talking-video models

Model-specific image/video/audio inputs through its prediction API

Exact model endpoint, usage unit and provider price

Test one authorized lip-sync API job

Use one representative video and audio file, inspect the complete mouth movement and timing, then compare retries, correction time and the final project charge.

Try Lip Sync

Reviewed September 10, 2026. This guide compares current documentation and integration fit. It does not claim a new controlled quality or latency benchmark. Magic Hour publishes the guide and appears in the comparison.

1. Magic Hour: existing footage plus a supplied audio track

The Magic Hour Lip Sync API creates a video job from your source footage and audio. You specify the clip’s start and end times, choose a generation mode and retrieve the completed result using its project ID. The API also exposes an output frame-rate cap.

Why shortlist it: A product can use the same platform for lip sync and related operations such as video generation, face swap or captions. For a still portrait, choose the separate AI Talking Photo API, which generates a speaking video from image and audio inputs.

Main limitation: These are asynchronous file-generation workflows. Do not design a live conversation interface around an assumption of instant rendering. Inspect your actual source footage and final output before setting a turnaround promise.

The current model and credit reference lists lip sync at 1 credit per output frame, doubled for Pro generation mode. The create response can contain an estimate; use the completed job’s charge for accounting. Generation mode and subscription tier are separate choices.

2. Sync: specialist models and explicit input constraints

Sync’s model documentation distinguishes lipsync-2, lipsync-2-pro and sync-3. That distinction matters more than a generic recommendation for “Sync Labs.” For example, lipsync-2 and lipsync-2-pro need visible speaking motion in the source, while sync-3 can animate silent lips, with different behavior and pricing.

Why shortlist it: Its developer workflow includes Python and TypeScript SDKs, asynchronous jobs and webhooks. It is a useful specialist option when your application already produces the audio and needs to apply it to footage.

Main limitation: Check the selected model’s handling of static segments, obstructions and multiple speakers. A higher output resolution or a successful request does not establish that every mouth looks correct.

Sync’s billing guide describes subscription plus usage. Rates quoted per second assume 25 fps; billing uses output frames. Higher plans change limits and discounts rather than making generation unlimited. Use its cost estimate for the exact model and files before submitting a large batch.

3. HeyGen: dedicated lip sync and separate translation APIs

HeyGen now documents a standalone Lip Sync API: POST /v3/lipsyncs accepts an existing video and replacement audio. You can choose Speed or Precision mode, then poll the job or receive a callback.

Why shortlist it: Controls include partial clip times, duration adjustment and optional captions. This suits a workflow in which the replacement dialogue is already approved and the remaining job is to align the visible speaker with it.

Main limitation: Lip sync itself does not translate or generate the replacement speech. HeyGen’s translation API adds those operations separately. Its audio-only translation option skips visual lip sync, so selecting “translation” does not by itself guarantee an edited mouth.

Check current API charges in the developer account for the operation and mode you choose. Do not calculate production cost from Creator or Pro website-subscription minutes. Our HeyGen pricing guide explains the distinction between the product plans and API billing.

4. D-ID: photo avatars and live audio-driven sessions

D-ID’s API quickstart separates asynchronous avatar-video generation from real-time agents. Its photo-avatar workflow creates a talking head from an image and speech input; its agent APIs support interactive applications.

Why shortlist it: For an application that already generates speech, D-ID’s Echo Sessions stream supplied audio to a V4 expressive avatar. Your application owns the conversation and voice, while D-ID renders the speaking avatar.

Main limitation: Echo and ordinary conversational sessions have different input behavior. Echo bypasses speech recognition, the language model and text-to-speech; it is not a full conversational assistant by itself. Confirm the avatar type and session interface before integrating.

Evaluate live systems on time to first visible speech, interruptions and turn-taking. Those measurements are different from the time needed to render a finished 30-second video.

5. Tavus: lip-synced faces inside a conversational system

Tavus CVI is designed for real-time video conversations. Its current terminology uses PAL for the configured conversational system, Face for its visual representation and Voice for speech.

Why shortlist it: The pipeline modes include a full conversational pipeline, Echo for directly supplied text or audio, and integrations with LiveKit or Pipecat. That is relevant when a talking face is part of an application that already has conversation logic.

Main limitation: This is a different integration task from uploading an arbitrary video and replacing its soundtrack. Choose it because you need a conversational avatar, then compare the relevant session limits, voice path and usage charges.

Avoid choosing a live-avatar stack solely because a listicle describes it as “a lip-sync API.” The surrounding session architecture may be unnecessary for a recorded-video feature.

6. Replicate: compare models without confusing them with providers

Replicate’s lip-sync collection brings together endpoints from different model publishers. Its catalog includes video-to-video lip-sync models and image-plus-audio models that generate a new talking video.

Why shortlist it: You can explore several models through a familiar prediction interface. For example, sync/lipsync-2 is Sync’s model hosted through Replicate. It is not a separate competing model simply because the API provider changes.

Main limitation: Inputs, output limits, terms and prices belong to the specific endpoint. A collection-level description should not replace the model’s input schema. Some models need a speaking video; others start from a still image and generate additional movement.

Record the model identifier and version, where applicable, with each result. When comparing the same model on two providers, evaluate access, billing, delivery and support separately from model quality.

What does a batch of lip-synced videos cost?

Define the job before comparing prices. For example: 10 clips, each 30 seconds long, with approved audio and one generation attempt per clip. This is 300 output seconds. The estimates below illustrate billing arithmetic, not a quality-equivalent benchmark.

Configuration

Calculation

Generation usage before other charges

magic hour logoMagic Hour at 24 output fps, non-Pro generation mode

10 × 30 × 24 × 1 credit

7,200 credits

magic hour logoMagic Hour at 24 output fps, Pro generation mode

10 × 30 × 24 × 2 credits

14,400 credits

Sync logoSync lipsync-2 at 25 output fps

300 × documented $0.04–$0.05 per second

$12–$15, plus the applicable subscription

Sync logoSync sync-3 at 25 output fps

300 × documented $0.107–$0.133 per second

$32.10–$39.90, plus the applicable subscription

Convert Magic Hour credits using your actual account’s rate. Recalculate if output frame rate changes. Add translation, speech generation, storage or other steps when your workflow needs them. Two attempts for every clip double the generation component.

For a production decision, use total spend divided by accepted output minutes. A cheaper successful render can still be expensive if the result needs to be regenerated or manually repaired.

A practical evaluation before integration

Use footage you have permission to modify and distribute. Compare the same input video and approved audio across candidates. Keep a separate evaluation for still-image animation or live avatars, because those start from different information.

Clip or behavior

What to inspect

Clear frontal speech

Mouth timing at normal playback speed; teeth and jaw detail

Pauses and closed-mouth moments

Whether the mouth keeps moving during silence

Head turn or partial obstruction

Facial blending, flicker and recovery after the obstruction

Your actual target-language audio

The final translated recording, including names, numbers and pauses

Multiple people or a scene cut

Which face changes and what happens between shots

Longer representative clip

Alignment near the beginning, middle and end

This is a suggested evaluation, not a test result. Record technical completion, visual acceptance, correction time and charge separately. Avoid turning one successful portrait into a claim about all languages, face angles or clip lengths.

How to connect a lip-sync API to your application

  1. Validate the inputs. Confirm that the endpoint accepts your source type and that the face and speech are usable.
  2. Prepare the audio. Approve the script or translation before paying to animate it. A lip-sync model cannot repair a wrong product name or an inaccurate translation.
  3. Submit once and save the job ID. Associate the provider job with the user’s request so a page refresh does not create a second paid generation.
  4. Track completion. Use the provider’s documented status endpoint or webhook. Treat queued, running, failed and completed states distinctly.
  5. Retrieve and inspect the output. A completed status establishes that a file was produced, not that it meets your quality requirements.
  6. Deliver the accepted file. Keep the actual charge and input settings with the job, and provide a clear recovery path when generation or review fails.

For a browser walkthrough before writing code, use our step-by-step lip-sync guide. If you also need translation and speech preparation, follow the AI dubbing workflow.

Frequently asked questions

Which lip-sync API is the most accurate?

There is no universal winner established by this guide. Accuracy depends on the model, source face, motion, audio and output settings. Shortlist a workflow fit, then compare representative clips using the same acceptance criteria.

Can I use a still image with a lip-sync API?

Sometimes, but use an endpoint that explicitly accepts it. Magic Hour has a separate Talking Photo API; D-ID offers photo-avatar generation; Replicate hosts image-plus-audio models. Do not send a portrait to an endpoint that requires existing speaking footage.

Is Wav2Lip free for commercial use?

The original Wav2Lip repository restricts its open-source version and pretrained results to personal, research or non-commercial use. A third-party API wrapper does not remove those restrictions. Check the exact model’s terms or a separately licensed commercial service.

Is lip sync the same as dubbing?

No. Dubbing replaces spoken audio. Visual lip sync changes the mouth to match that audio. Translation, voice generation, subtitle preparation and lip sync may be separate paid steps even when a product offers all of them.

Which option should I try first for an existing video?

Try one short clip with the Magic Hour Lip Sync tool, then use its documented API workflow if the output fits your application. Compare Sync or HeyGen on the same files when you need a second option before committing to a larger integration.

Runbo Li
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
See all articles

Continue Reading

How to Lip Sync a Video With AI
Videos
How to lip sync a video with Magic Hour: 5 steps
Jun 24, 2026
6 Best AI Lip Sync Tools
Content Creation
6 best AI lip sync video tools (2026 comparison)
Jul 06, 2025
MH lip sync
Content Creation
Magic Hour lip sync: workflow, pricing & troubleshooting
Jun 28, 2025
best ai image and video apis
VideosImages
9 best AI image and video APIs: costs and integration
Jun 14, 2025
5 Best AI Image Generators & their API
Images
5 best AI image generation APIs in 2026
Oct 13, 2025
6 Best Free AI Lip Sync Tools
VideosTop Choice
Best free AI lip sync tools (2026): limits & watermarks
May 25, 2026