7 best talking photo API options: inputs, costs and limits

Runbo Li
Runbo Li
·
· 9 min read
Talking Photo APIs Powering AI Avatars

Quick answer

Start with Magic Hour for photo-and-audio generation in an existing image/video workflow, HeyGen for reusable photo avatars with built-in speech, or D-ID for its documented photo-avatar API. AKOOL and VModel offer additional direct photo-and-audio endpoints. Synthesia fits template-based business videos; SadTalker is a self-hosted model you would need to turn into a service yourself.

Create a talking-photo prototype

Upload one clear portrait and authorized audio, then review lip movement, expressions, timing, and output limits before integrating the API.

Open Talking Photo API Docs

A talking-photo API animates a still face to match speech. It is different from lip-syncing an existing video, and a rendered video endpoint is different from a real-time conversational avatar. Choose the input and output your application actually needs before comparing prices.

This guide is published by Magic Hour. Sources were reviewed on September 10, 2026. The recommendations are based on documented capabilities, billing and integration requirements. They are not a measured ranking of facial realism, latency or failure rates.

Magic Hour publishes this guide and is included where relevant. We selected current options with official product documentation and a distinct fit for the tasks in this guide, then compared documented inputs, controls, limits, exports, pricing mechanics, and workflow fit. This is not a controlled output-quality benchmark unless a retained test is explicitly described below.

Magic Hour publishes this guide and includes its own product in the comparison. Treat the recommendations as editorial guidance from a vendor, and verify the linked first-party product details and your own output requirements before choosing a tool.

Talking-photo API comparison

Option

Starting input

Useful fit

Billing or access distinction

Magic Hour

Image and audio assets

Generate talking photos alongside other image/video tools

Credits vary by mode and funding method; inspect the completed job charge

HeyGen Photo Avatar

Photo avatar, script and voice

Reuse an approved presenter across scripts

Photo Avatar IV at 720p/1080p is $0.05 per output second; creation is separate

D-ID V2 Photo Avatars

Image URL and text or audio script

Direct photo-to-speaking-video integration

API plan, license and 15-second billing increments matter

AKOOL Talking Photo

Image URL and audio URL

Direct animation with optional gesture prompt

Published API table lists 10 credits per five seconds; confirm resolution pricing

VModel Talking Photo Turbo

Image URL and audio URL

Evaluate a clearly priced hosted endpoint

Published rate is $0.002 per second of input audio

Synthesia

An available avatar and script or template

Branded business-video automation

API access on Creator and Enterprise; shared credit allowance

SadTalker

Portrait and audio files

Self-hosted research or custom infrastructure

No hosted service included; compute and operations remain your responsibility

If your input is already a video, use the lip-sync API comparison. If you want to animate a face manually before integrating, try the talking-photo tool.

1. Magic Hour: photo-and-audio generation within a media workflow

Magic Hour exposes POST /v1/ai-talking-photo. The request takes image and audio asset paths, start and end times, and optional style settings. You receive a project ID, then retrieve the completed video. The API reference explains the job response and download workflow.

Useful distinction: The live OpenAPI contract currently defines two modes: realistic, with a maximum clip length of 300 seconds, and prompted, with a maximum of 45 seconds. The optional motion prompt applies to prompted mode. Older aliases are retained for compatibility; new integrations should use the current mode names. These are API limits, not the guest web tool’s allowance.

Choose it when: Your application needs photo animation alongside other Magic Hour image, video or audio operations, and you want to manage those projects in the same platform.

Watch for: The raw endpoint requires an audio asset. If your application starts with text, generate and approve the speech first; do not pass script text into a field documented as an audio path. The SDK and raw REST API also have different conveniences for local files and URLs—follow the upload instructions for the interface you use.

Cost: Read the charge for the selected mode and account funding method. The API’s credits_charged field can be an estimate before completion. Use the final value when reconciling your budget, rather than a fixed “price per video.” Current plans and credit packs describe funding options and commercial-use access.

2. HeyGen Photo Avatar: reusable presenters and scripts

HeyGen’s Photo Avatar workflow creates an avatar from a photo through POST /v3/avatars, then generates a video through POST /v3/videos. The video request includes the avatar ID, script and voice ID. It supports motion prompts, expressiveness settings and a completion callback.

Choose it when: You want to create an approved presenter once and reuse it for different scripts, backgrounds or languages supported by your selected voice.

Watch for: Avatar creation and video rendering are separate operations. Preserve the avatar ID instead of recreating it for every script. The API uses a separate pay-as-you-go balance. An editor subscription is not automatically the same budget as your API integration.

Cost: Self-serve API pricing lists Photo Avatar IV at $0.05 per output second for 720p or 1080p, and $4 per minute for 4K, charged by actual seconds. Photo-avatar creation is $1 per call. For a new 30-second 1080p video with one newly created avatar, that is $1.50 for video + $1 for creation = $2.50, before any separately billed operations. Reusing that avatar avoids another creation operation.

3. D-ID: a direct photo-avatar endpoint

D-ID’s V2 Photo Avatar quickstart uses POST /talks with a source image URL and a script. The example supplies text and a voice provider. D-ID also documents audio-script inputs. Save the returned talk ID, check its status, and download the result when the status is done.

Choose it when: You need a documented photo-to-speaking-video flow and want to evaluate D-ID’s voices or wider avatar services.

Watch for: The quickstart says result URLs are valid for 24 hours. Save the generated file or re-fetch the talk for a fresh URL. A live-avatar integration uses a different workflow: D-ID now directs new streaming integrations toward Agents SDK or Agents Streams, rather than its legacy Talks Streams flow.

Cost: Compare the actual API plan and its license. D-ID’s billing FAQ describes one credit for each started 15 seconds of generated video: a 30-second clip uses two credits, while a 31-second clip uses three. That is why “100 videos” alone is not enough information to quote a bill. Personal-license and commercial-license plans also serve different uses.

4. AKOOL: photo-and-audio animation with gesture prompts

AKOOL’s Talking Photo endpoint accepts talking_photo_url and audio_url, plus an optional prompt and webhook URL. Its current schema lists 720p and 1080p resolution options.

Choose it when: Your integration already produces audio and you want to guide gestures while animating a portrait.

Watch for: An accepted creation response is not a finished video. The overview defines queue, processing, success and failure states; video_status = 3 indicates a completed result. The generated resources are valid for seven days, so copy approved files to your own storage.

Cost: The API pricing table lists ten credits per five seconds for Talking Photo on its displayed self-serve API tiers. A 30-second job at that listed rate is 60 credits. The endpoint documentation notes that higher resolution costs more, so confirm the applicable resolution rate and dollar value of your account’s credits before budgeting a 1080p batch.

5. VModel Talking Photo Turbo: a priced endpoint to evaluate

VModel’s model page documents avatar and speech URL inputs, a model version identifier, a task response and generated video output. It publishes a rate of $0.002 per second of input audio.

Choose it when: You want to include a low listed usage rate in a technical shortlist and can independently evaluate the output and service requirements.

Watch for: The page’s example result is a demonstration, not evidence of your application’s expected speed or quality. Confirm supported duration, resolution, output retention, commercial terms and account limits for the intended deployment. Do not assume those match a higher-priced provider.

Cost: At its published audio-duration rate, a 30-second request is $0.06, and 100 such requests are $6 in generation usage. This calculation does not establish equivalent output quality, retry rates or service guarantees, and it excludes any account funding minimum.

6. Synthesia: avatar videos inside a business template

Synthesia supports reusable personal avatars created from photos, alongside other avatar types. Its API supports programmatic video creation and personalized templates. This is a different setup from sending a new arbitrary image with every animation request.

Choose it when: Your team wants a consistent business-video template, approved presenters and repeatable script variations.

Watch for: Create and make the avatar available in the workspace first. Check supported avatar IDs and template variables rather than assuming a portrait-upload parameter exists on a video-generation endpoint.

Cost: API access is available on Creator and Enterprise. Creator is listed at $89 on monthly billing. Its current credit guide lists 75 credits per minute of API usage, while other features have different rates. Check the shared balance and exact operation; the headline number of video minutes is not a universal allowance for every feature.

7. SadTalker: build and operate your own service

SadTalker is a research implementation that animates a single portrait from audio. Its repository states that the code uses Apache 2.0 and provides installation instructions, model downloads and local inference workflows.

Choose it when: Self-hosting or modifying the pipeline is a real requirement, and your team can maintain the environment.

Watch for: You are selecting software to operate, not a managed API with a service commitment. You must provide authentication, job handling, compute, storage and support. Review the licenses of any additional models or components you deploy.

Cost: There is no fixed hosted per-video rate in the repository. Count GPU runtime, idle capacity, storage, transfer and maintenance. Open-source code can reduce a vendor dependency without making production operation free.

Compare the same job before choosing a provider

For a useful first comparison, define 100 videos of 30 seconds each, using one approved presenter and one output resolution. That is 3,000 seconds, or 50 minutes, of completed video before retries.

Provider and stated assumption

Usage for this job

What remains outside the calculation

HeyGen Photo Avatar IV, 1080p, one new avatar

$150 video + $1 avatar creation = $151

Additional avatars, separately billed operations and wallet funding requirements

D-ID at one credit per started 15 seconds

200 credits

Plan price, commercial-license eligibility and unused allowance

AKOOL at the listed ten credits per five seconds

6,000 credits

Applicable resolution rate, credit purchase terms and plan cost

VModel Turbo at $0.002 per input-audio second

$6 generation usage

Account minimum, required output specifications and retries

Magic Hour

Sum of final credits_charged for the selected mode

Convert credits using your actual funding method; include speech generation if needed

Synthesia at 75 credits per API minute

3,750 credits

Plan capacity and other use of the shared credit balance

SadTalker

Measured compute and operating cost

Hosting, engineering and support; no standard vendor rate

These are billing calculations, not equivalent-quality benchmark results. A provider that charges less per second can still cost more per usable video if it needs more attempts or manual correction. Keep avatar creation separate from rendering, and keep audio generation separate wherever it is billed separately.

A practical first integration

  1. Approve the input. Use a clear portrait and speech you have permission to use. Check the script’s names, product claims and pronunciation before animation.
  2. Create one short job. Start with a 15- to 30-second introduction in the language your users need. Save the provider job ID and input-to-output association.
  3. Wait for completion. Follow the provider’s polling or webhook flow. A submitted or queued status is not a completed video. Handle an explicit failed state instead of waiting indefinitely.
  4. Retrieve the file. Download before a temporary result URL expires. Confirm that the file contains the expected video and audio and can play through to the end.
  5. Review the result. Check identity, speech alignment, unwanted movement, end-of-sentence timing and readability at the final display size.
  6. Expand only after that works. Generate the required script variations within account limits. If a creation request times out ambiguously, reconcile its status before submitting a replacement that could create a second billable job.

For a basic marketing welcome video, a useful script is: “Welcome to [product]. In this short guide, I’ll show you how to [specific task]. Start by opening [actual feature].” Replace every placeholder with approved copy. This is an example script, not a generated testimonial or a claim that a customer endorsed the product.

Questions marketing teams and developers should answer

This guide does not establish a universal winner. Compare the same permitted portrait, audio and output size. A clean front-facing headshot is only one case; also inspect the real framing, glasses, facial hair, speaking pace and language your product must support. Do not infer reliability from a single vendor demo.

Some workflows accept a script and voice selection; others require an audio file. HeyGen’s photo-avatar workflow and D-ID’s text-script example include voice selection. Magic Hour’s raw talking-photo endpoint and the AKOOL endpoint described here require audio assets. Budget the preceding speech-generation step when needed.

A service that renders a downloadable MP4 does not automatically provide an interactive conversation. For a live avatar, evaluate the provider’s separate streaming interface, audio transport, interruption behavior and session billing. D-ID’s current Agents documentation is one example of that distinct integration path.

First confirm that the provider can complete your exact input-to-download workflow and that the result passes review. Then calculate your actual monthly seconds, avatar count and retry usage. If you want to start with Magic Hour, try a talking photo to assess your source image, then use the API reference for integration. For a broader media pipeline, compare the image and video generation APIs.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

lip sync girl
Recommended next
What is AI lip sync? How it works, uses and limits

Learn how AI lip sync turns audio into synchronized facial motion, how it differs from dubbing and face swap, where it fails, and how to review results.

How to Lip Sync a Video With AI
How to lip sync a video with Magic Hour: 5 steps
Magic Hour editorial collage comparing AI lip sync with phoneme drawings and a speech waveform
7 best AI lip sync video tools (2026 comparison)
AI Avatar Gen
How to make a realistic talking AI avatar
AI Talking Photo Tools (2026)
Best AI talking photo tools (2026): 5 portrait workflows