AI video API latency benchmark: 30 days of real jobs

Runbo Li
Runbo Li
·
· 4 min read
8,160 AI Video Jobs Tested editorial benchmark cover

Quick answer

From August 15 through September 13, 2026, median end-to-end time among successful Magic Hour API jobs was 36 seconds for Text to Video, 52 seconds for Auto Subtitle, 71 seconds for AI Video Editor, and 87 seconds for Image to Video. Character Replace had the longest median at 21 minutes 3 seconds. These measurements include queueing and are directional, not an SLA.

The 30-day results

The table covers 8,160 successful jobs across 11 Magic Hour video endpoints. A smaller median means the typical completed job finished sooner inside this specific window; it does not establish output quality, acceptance rate, reliability under every workload, or an advantage over another provider.

Endpoint

Typical (p50)

Successful jobs

What the timing includes

Text to Video (/v1/text-to-video)

36s

1,974

Project creation through completion, including queueing

Auto Subtitle (/v1/auto-subtitle-generator)

52s

60

Project creation through completion, including queueing

AI Video Editor (/v1/ai-video-editor)

1m 11s

25

Project creation through completion, including queueing

Image to Video (/v1/image-to-video)

1m 27s

2,030

Project creation through completion, including queueing

Lip Sync (/v1/lip-sync)

1m 46s

767

Project creation through completion, including queueing

Talking Photo (/v1/ai-talking-photo)

1m 48s

269

Project creation through completion, including queueing

Face Swap Video (/v1/face-swap)

1m 49s

2,431

Project creation through completion, including queueing

Animation (/v1/animation)

2m 43s

18

Project creation through completion, including queueing

Audio to Video (/v1/audio-to-video)

4m 31s

308

Project creation through completion, including queueing

Video to Video (/v1/video-to-video)

6m 20s

79

Project creation through completion, including queueing

Character Replace (/v1/character-replace)

21m 03s

199

Project creation through completion, including queueing

Median end-to-end completion time and successful-job sample size for 11 Magic Hour video API endpoints

Run one traceable API job

Choose the exact endpoint for your workflow, store the returned project ID, and measure one representative job from submission through the downloaded result before setting production timeouts.

Explore Magic Hour API

Methodology

Window: 30 complete UTC days from August 15, 2026 at 00:00 through September 14 at 00:00, covering jobs created through September 13.

Population: Magic Hour VideoProject records marked as API traffic. The latest warehouse row for each project ID was retained before aggregation.

Included: jobs with final status complete and a completion timestamp. Each endpoint's sample is shown in the table.

Excluded: errored, canceled, pending, starting, and other active jobs. The results therefore describe completion time among successful jobs; they are not a technical completion-rate benchmark.

Metric: p50 of completedAt minus createdAt in seconds. This is end-to-end time from project creation to completion, so queueing is included.

Privacy: only endpoint-level aggregates are published. No account, user, project, prompt, input, output, or file-level data appears in the table or chart.

What the results show

Text to Video had the shortest measured median at 36 seconds across 1,974 successful jobs. Image to Video recorded 87 seconds across 2,030 jobs, and Face Swap Video recorded 109 seconds across 2,431. Character Replace was the longest workflow at 21 minutes 3 seconds across 199 successful jobs.

The spread reflects different workloads as well as infrastructure. Character replacement, video transformation, lip synchronization, captioning, and text generation do not perform the same amount of work. Compare an endpoint with its own historical baseline or with another provider performing the same job, duration, resolution, audio setting, and input size.

Google's SRE guidance explains why a median represents typical performance while slower-tail behavior must be evaluated separately. This public table intentionally reports p50 only. Teams choosing timeout, retry, or capacity rules should measure the full distribution for their exact production workload.

How to run a fair video API benchmark

  • Freeze the workload. Use the same input mode, duration, resolution, audio requirement, region, account tier, and concurrency.

  • Name the model and provider. A platform may expose several models and routing paths. Record the exact endpoint and selected model/version.

  • Retain every attempt. Keep successful, failed, canceled, timed-out, and rejected outputs in separate counts instead of selecting the fastest result.

  • Define acceptance before review. Score subject fidelity, motion continuity, text and product accuracy, audio, and any other requirement that determines whether the output can ship.

  • Separate outcomes. Report technical completion, human acceptance, end-to-end latency, retry behavior, and cost per accepted output as different metrics.

Do not turn the median into a timeout

A median says half of the successful jobs in the measured cohort completed within that value. It does not say a slower active job has failed. Use webhooks where practical, preserve the returned project ID, and handle complete, error, and canceled as distinct terminal states.

AWS guidance on retry-safe APIs explains why a client timeout can leave an outcome uncertain. Retrieve the existing project before submitting again; a delayed polling request is not proof that the generation should be duplicated.

Set user-facing timeouts from the workload's tolerance for waiting and test the exact models, settings, input sizes, and arrival pattern that production will use. Reaching the application's timeout should stop or defer the waiting flow while preserving the project ID for later status checks.

Continue learning

Continue learning: best AI image and video APIs, AI media API evaluation framework, and text-to-image API integration guide.

Frequently asked questions

Text to Video had the shortest p50 at 36 seconds among 1,974 successful jobs. This comparison spans different tasks, so it should not be read as a quality or infrastructure ranking.

Yes. The measurement uses project completion time minus project creation time, so it includes queueing and generation until the project reaches complete.

No. The latency table includes only successfully completed jobs. Errored, canceled, and active jobs are excluded, so use a separate reliability analysis when choosing a production provider.

No. They are historical directional medians for one 30-day cohort. Job time varies with endpoint, model, settings, duration, resolution, input size, concurrency, and queue conditions.

Each endpoint name in the table links to its current reference. Use the Magic Hour API documentation for request fields, status retrieval, webhooks, errors, and current availability.

Google’s SRE guidance explains why a median describes typical latency while a high percentile exposes the slower tail that an average can hide. Pick the percentile before collecting results and publish the whole measurement definition.

Queue and generation time can be separated only when the provider exposes trustworthy stage timestamps. Do not infer them from the time between polls.

Publish the model IDs, settings, dates, sample size, prompts, acceptance rubric, measurement definitions, and aggregated results. Keep run-level records available for verification. Until that evidence exists, describe performance expectations as hypotheses rather than measured facts.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

best ai image and video apis
Recommended next
9 best AI image and video APIs: costs and integration

Compare nine AI image and video API providers by models, billing, workflow and deployment, with concrete cost examples and integration guidance.

Median end-to-end completion time and successful-job sample size for 11 Magic Hour video API endpoints
How to evaluate AI media APIs: total time, cost and reliability
Collage of logos from the best text-to-image APIs.
Text-to-image API integration guide: request to accepted image
Multimodal Video APIs
Best multimodal video APIs: inputs, controls & costs
Illustration showing text prompts transforming into AI-generated videos using developer APIs
7 best text-to-video APIs: models, queues & avatars