

For matching new audio to an existing video, start with Magic Hour, Sync or HeyGen. Magic Hour is useful when lip sync sits alongside other image and video tools; Sync offers several specialist lip-sync models; HeyGen exposes dedicated dialogue-replacement and translation workflows. Choose D-ID or Tavus for a live talking avatar, and Replicate when you want to compare hosted models through one provider.
The first decision is your input: an existing video, a still portrait or a live conversation. These require different endpoints. A service that animates an avatar is not automatically a drop-in API for changing the dialogue in your footage.
Provider | Best workflow fit | Input and delivery | Billing detail to check |
|---|---|---|---|
Recorded-video lip sync within a broader media workflow | Video plus audio; asynchronous output file. Still images use the separate Talking Photo API | Output frames and selected generation mode | |
Choosing among specialist models for existing footage | Video plus audio; asynchronous jobs, SDKs and webhooks | Model-specific frame rate, subscription and usage charges | |
Dialogue replacement plus optional translation workflows | Dedicated video-plus-audio lip-sync endpoint; separate avatar and translation endpoints | Current API operation and mode, separate from website subscriptions | |
Photo avatars or live agents that speak supplied audio | Photo-avatar video jobs; separate real-time agent sessions | Avatar/version, session or video allowance and account limits | |
A conversational avatar with your own voice or application logic | CVI conversation sessions, including an Echo mode | Conversation plan, usage and integration mode | |
Comparing hosted lip-sync and talking-video models | Model-specific image/video/audio inputs through its prediction API | Exact model endpoint, usage unit and provider price |
Use one representative video and audio file, inspect the complete mouth movement and timing, then compare retries, correction time and the final project charge.
Try Lip SyncReviewed September 10, 2026. This guide compares current documentation and integration fit. It does not claim a new controlled quality or latency benchmark. Magic Hour publishes the guide and appears in the comparison.
The Magic Hour Lip Sync API creates a video job from your source footage and audio. You specify the clip’s start and end times, choose a generation mode and retrieve the completed result using its project ID. The API also exposes an output frame-rate cap.
Why shortlist it: A product can use the same platform for lip sync and related operations such as video generation, face swap or captions. For a still portrait, choose the separate AI Talking Photo API, which generates a speaking video from image and audio inputs.
Main limitation: These are asynchronous file-generation workflows. Do not design a live conversation interface around an assumption of instant rendering. Inspect your actual source footage and final output before setting a turnaround promise.
The current model and credit reference lists lip sync at 1 credit per output frame, doubled for Pro generation mode. The create response can contain an estimate; use the completed job’s charge for accounting. Generation mode and subscription tier are separate choices.
Sync’s model documentation distinguishes lipsync-2, lipsync-2-pro and sync-3. That distinction matters more than a generic recommendation for “Sync Labs.” For example, lipsync-2 and lipsync-2-pro need visible speaking motion in the source, while sync-3 can animate silent lips, with different behavior and pricing.
Why shortlist it: Its developer workflow includes Python and TypeScript SDKs, asynchronous jobs and webhooks. It is a useful specialist option when your application already produces the audio and needs to apply it to footage.
Main limitation: Check the selected model’s handling of static segments, obstructions and multiple speakers. A higher output resolution or a successful request does not establish that every mouth looks correct.
Sync’s billing guide describes subscription plus usage. Rates quoted per second assume 25 fps; billing uses output frames. Higher plans change limits and discounts rather than making generation unlimited. Use its cost estimate for the exact model and files before submitting a large batch.
HeyGen now documents a standalone Lip Sync API: POST /v3/lipsyncs accepts an existing video and replacement audio. You can choose Speed or Precision mode, then poll the job or receive a callback.
Why shortlist it: Controls include partial clip times, duration adjustment and optional captions. This suits a workflow in which the replacement dialogue is already approved and the remaining job is to align the visible speaker with it.
Main limitation: Lip sync itself does not translate or generate the replacement speech. HeyGen’s translation API adds those operations separately. Its audio-only translation option skips visual lip sync, so selecting “translation” does not by itself guarantee an edited mouth.
Check current API charges in the developer account for the operation and mode you choose. Do not calculate production cost from Creator or Pro website-subscription minutes. Our HeyGen pricing guide explains the distinction between the product plans and API billing.
D-ID’s API quickstart separates asynchronous avatar-video generation from real-time agents. Its photo-avatar workflow creates a talking head from an image and speech input; its agent APIs support interactive applications.
Why shortlist it: For an application that already generates speech, D-ID’s Echo Sessions stream supplied audio to a V4 expressive avatar. Your application owns the conversation and voice, while D-ID renders the speaking avatar.
Main limitation: Echo and ordinary conversational sessions have different input behavior. Echo bypasses speech recognition, the language model and text-to-speech; it is not a full conversational assistant by itself. Confirm the avatar type and session interface before integrating.
Evaluate live systems on time to first visible speech, interruptions and turn-taking. Those measurements are different from the time needed to render a finished 30-second video.
Tavus CVI is designed for real-time video conversations. Its current terminology uses PAL for the configured conversational system, Face for its visual representation and Voice for speech.
Why shortlist it: The pipeline modes include a full conversational pipeline, Echo for directly supplied text or audio, and integrations with LiveKit or Pipecat. That is relevant when a talking face is part of an application that already has conversation logic.
Main limitation: This is a different integration task from uploading an arbitrary video and replacing its soundtrack. Choose it because you need a conversational avatar, then compare the relevant session limits, voice path and usage charges.
Avoid choosing a live-avatar stack solely because a listicle describes it as “a lip-sync API.” The surrounding session architecture may be unnecessary for a recorded-video feature.
Replicate’s lip-sync collection brings together endpoints from different model publishers. Its catalog includes video-to-video lip-sync models and image-plus-audio models that generate a new talking video.
Why shortlist it: You can explore several models through a familiar prediction interface. For example, sync/lipsync-2 is Sync’s model hosted through Replicate. It is not a separate competing model simply because the API provider changes.
Main limitation: Inputs, output limits, terms and prices belong to the specific endpoint. A collection-level description should not replace the model’s input schema. Some models need a speaking video; others start from a still image and generate additional movement.
Record the model identifier and version, where applicable, with each result. When comparing the same model on two providers, evaluate access, billing, delivery and support separately from model quality.
Define the job before comparing prices. For example: 10 clips, each 30 seconds long, with approved audio and one generation attempt per clip. This is 300 output seconds. The estimates below illustrate billing arithmetic, not a quality-equivalent benchmark.
Configuration | Calculation | Generation usage before other charges |
|---|---|---|
10 × 30 × 24 × 1 credit | 7,200 credits | |
10 × 30 × 24 × 2 credits | 14,400 credits | |
300 × documented $0.04–$0.05 per second | $12–$15, plus the applicable subscription | |
300 × documented $0.107–$0.133 per second | $32.10–$39.90, plus the applicable subscription |
Convert Magic Hour credits using your actual account’s rate. Recalculate if output frame rate changes. Add translation, speech generation, storage or other steps when your workflow needs them. Two attempts for every clip double the generation component.
For a production decision, use total spend divided by accepted output minutes. A cheaper successful render can still be expensive if the result needs to be regenerated or manually repaired.
Use footage you have permission to modify and distribute. Compare the same input video and approved audio across candidates. Keep a separate evaluation for still-image animation or live avatars, because those start from different information.
Clip or behavior | What to inspect |
|---|---|
Clear frontal speech | Mouth timing at normal playback speed; teeth and jaw detail |
Pauses and closed-mouth moments | Whether the mouth keeps moving during silence |
Head turn or partial obstruction | Facial blending, flicker and recovery after the obstruction |
Your actual target-language audio | The final translated recording, including names, numbers and pauses |
Multiple people or a scene cut | Which face changes and what happens between shots |
Longer representative clip | Alignment near the beginning, middle and end |
This is a suggested evaluation, not a test result. Record technical completion, visual acceptance, correction time and charge separately. Avoid turning one successful portrait into a claim about all languages, face angles or clip lengths.
For a browser walkthrough before writing code, use our step-by-step lip-sync guide. If you also need translation and speech preparation, follow the AI dubbing workflow.
There is no universal winner established by this guide. Accuracy depends on the model, source face, motion, audio and output settings. Shortlist a workflow fit, then compare representative clips using the same acceptance criteria.
Sometimes, but use an endpoint that explicitly accepts it. Magic Hour has a separate Talking Photo API; D-ID offers photo-avatar generation; Replicate hosts image-plus-audio models. Do not send a portrait to an endpoint that requires existing speaking footage.
The original Wav2Lip repository restricts its open-source version and pretrained results to personal, research or non-commercial use. A third-party API wrapper does not remove those restrictions. Check the exact model’s terms or a separately licensed commercial service.
No. Dubbing replaces spoken audio. Visual lip sync changes the mouth to match that audio. Translation, voice generation, subtitle preparation and lip sync may be separate paid steps even when a product offers all of them.
Try one short clip with the Magic Hour Lip Sync tool, then use its documented API workflow if the output fits your application. Compare Sync or HeyGen on the same files when you need a second option before committing to a larger integration.
