

Video API latency should be measured from submission to a downloadable result, with retries and output acceptance reported separately. Compare equivalent model settings and publish the sample size, dates, median, and slower-tail performance. The earlier provider tables on this page did not include enough underlying evidence to substantiate their speed, reliability, or cost rankings; those tables have been removed.
Metric | Definition | What to record |
|---|---|---|
Submission time | Time until the provider acknowledges the job | Request start and acknowledgement |
Time to ready | Time until output is available | Submission and completion timestamps |
Time to usable result | Time until an accepted output is obtained | All attempts and review outcome |
Technical completion rate | Completed jobs divided by submitted jobs | Failures and timeouts included |
Acceptance rate | Outputs meeting the predefined brief divided by reviewed outputs | Rubric and reviewer decisions |
Cost per accepted output | Total relevant spend divided by accepted outputs | Charges, refunds, and all attempts |
Queue and generation time can be separated only when the provider exposes trustworthy stage timestamps. Do not infer them from the time between polls.
Choose the same input mode, target duration, resolution, and audio requirement. Record the exact model and provider, test region, dates, account tier, and concurrency. Comparing different models is a workflow comparison, not an isolated test of provider infrastructure.
Use representative scenes and preserve the prompts and input assets. Define acceptance before reviewing results: for example, correct subject, no added objects, continuous motion, and readable product text. Keep policy rejections, technical errors, and visually rejected outputs as separate categories.
Show sample size alongside the median and a slower-tail percentile such as p95. Include unfinished and failed jobs in the accounting; reporting only completed jobs can make an unreliable service look faster. A small sample gives an unstable estimate of tail behavior, so label it accordingly.
Do not rank providers from one best-case clip. Repeat measurements only when the decision warrants the cost and respect each API’s documented limits.
Store the returned job ID, poll that job or use the provider’s documented completion mechanism, and handle transient status failures separately from failed generation. A delayed status request does not prove that the underlying job needs to be submitted again.
For Magic Hour integration details, use the official API documentation. This article does not claim that Magic Hour has a measured latency or reliability advantage over another provider.
Suppose a trial submits ten jobs, nine complete technically, and seven outputs satisfy the brief. Technical completion is 90%; usable acceptance across submitted jobs is 70%. If all relevant charges total $14, cost per accepted output is $2. These are hypothetical numbers illustrating the method, not benchmark results.
Publish the model IDs, settings, dates, sample size, prompts, acceptance rubric, measurement definitions, and aggregated results. Keep run-level records available for verification. Until that evidence exists, describe performance expectations as hypotheses rather than measured facts.
