AI video model benchmark: 60 attempts across 5 models

Runbo Li
Runbo Li
·
· 4 min read
AI Video Model Benchmark

Quick answer

Magic Hour's commercial image-to-video benchmark requested 60 videos across five named models and four source scenarios. Fifty-eight requests completed with downloadable MP4s; two failed. The release includes all attempts, exact prompts, source images, charged credits, and output files. It does not declare a visual-quality winner: completed API jobs have not been scored for human acceptance.

Cloud shadows move slowly across the scene while foreground plants respond to a light breeze. The camera makes one gentle push forward. Preserve the original composition, subjects, and visual style. One continuous shot.

Test an image-to-video brief

Upload a source image to Magic Hour, choose a current model, and review the entire clip against the details your project must preserve.

Try Image-to-Video

What was compared?

The requested models were Kling 3.0, Seedance 2.5, Veo 3.1, Sora 2, and LTX 2.5, all accessed through Magic Hour's image-to-video API. Each received the same source image and motion prompt for a skincare bottle, sneaker, coffee scene, and furnished room, with three attempts per model-scenario pair.

Settings were eight seconds, requested 720p, and audio disabled. The dataset was published September 9, 2026, with request timestamps on September 10 UTC. Magic Hour designed and funded the study and provides the API used for every model. This is one sponsored run through one integration, not an independent comparison of the providers' own applications.

Recorded completion and credits

Requested model

Completed / attempted

Median credits per completed output

Kling 3.0

12 / 12

384

LTX 2.5

12 / 12

384

Veo 3.1

12 / 12

768

Sora 2

12 / 12

960

Seedance 2.5

10 / 12

4,608

The two Seedance failures carried zero final credits in the dataset. These are observed account charges for the recorded requests, not current universal model prices. API completion was 58/60, or 96.7%, in this run. That is not a 96.7% usable-video or customer-success rate.

Use the versioned dataset and DOI for citation and the run-level results to inspect each request. Availability, routing, billing, and model behavior can change.

Inspect the actual commercial examples

The skincare contact sheet below shows the original source and all fifteen requested attempts for that scenario, including the failure. It contains still frames, so it cannot establish motion continuity, label stability across time, or overall clip quality.

Skincare benchmark contact sheet showing the source and all fifteen model attempts, including one failure

Watch the complete output gallery before drawing a visual conclusion. The four scenarios answer different practical questions:

  • Skincare: does a small camera orbit preserve the bottle, label symbol, liquid level, and pedestal?
  • Footwear: do the shoe silhouette, laces, stitching, color blocks, and sole survive a controlled camera move?
  • Coffee: can steam move while the bag, cup, beans, and counter remain consistent?
  • Interior: do furniture identities, straight lines, room layout, and materials survive a dolly move?

These are review questions, not completed scores. Read the frozen study design and prompts to see exactly what each request asked for.

What can this benchmark establish?

It establishes the recorded completion status, charged credits, and downloadable outputs for these attempts. It also documents a shared request schema across the five named model selections through Magic Hour.

It does not establish a best-looking model, a cost per accepted ad, or a general reliability ranking. There were only twelve attempts per model, and no completed human quality scorecard. The collector checked jobs sequentially, so observed terminal time is an upper bound; it should not be presented as exact generation latency or used to rank model speed.

Sora's inclusion is historical access evidence for this run. OpenAI's shutdown notice schedules its API shutdown for September 24, 2026, after closing the consumer experience in April. Check current provider access before planning a new long-term workflow.

Use the evidence for your own product brief

Write the details that must survive before generating: product shape, readable label, material, included items, camera movement, and intended export. Choose one of the eight product-video prompt recipes and adapt it to a source image you have permission to use.

Start in Magic Hour Image-to-Video or use the API documentation for integration. Record the exact model and settings and retain failed and rejected attempts. Review the whole clip against your brief, not just its thumbnail.

Use the video-cost worksheet to calculate total relevant spend divided by accepted outputs. Until the clips have been reviewed, this dataset supports cost per completed output, not cost per usable commercial deliverable.

Citation and reuse

Cite: Magic Hour AI, Inc., Commercial Image-to-Video Model Benchmark via Magic Hour, version 1.0.0, September 9, 2026, DOI 10.5281/zenodo.22683989. Include the sponsor, settings, sample size, and distinction between completion and quality when quoting results.

The repository release includes prompts, source images, outputs, checksums, and reproduction code. Methodology and result metadata use CC BY 4.0; code uses MIT; generated media have separate reuse terms.

For a broader platform decision alongside these controlled results, see the best AI video generators guide.

For a separate retained advertising-output evaluation, read the Fable 5 vs Sol 5.6 taste test.

Frequently asked questions

The requested models were Kling 3.0, Seedance 2.5, Veo 3.1, Sora 2, and LTX 2.5, all accessed through Magic Hour's image-to-video API. Each received the same source image and motion prompt for a skincare bottle, sneaker, coffee scene, and furnished room, with three attempts per model-scenario pair.

Settings were eight seconds, requested 720p, and audio disabled. The dataset was published September 9, 2026, with request timestamps on September 10 UTC. Magic Hour designed and funded the study and provides the API used for every model. This is one sponsored run through one integration, not an independent comparison of the providers' own applications.

It establishes the recorded completion status, charged credits, and downloadable outputs for these attempts. It also documents a shared request schema across the five named model selections through Magic Hour.

It does not establish a best-looking model, a cost per accepted ad, or a general reliability ranking. There were only twelve attempts per model, and no completed human quality scorecard. The collector checked jobs sequentially, so observed terminal time is an upper bound; it should not be presented as exact generation latency or used to rank model speed.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Analog filmmaker workbench comparing AI video generation workflows and finished frames
Recommended next
10 best AI video generators in 2026: models, features, and costs

Compare 10 current AI video models and platforms by generation, editing, audio, references, APIs, commercial use, and cost per accepted clip.

Keep Characters Consistent Without Manual Editing
Best reference image-to-video tools (2026): character and product consistency
Comparison of top AI video generators including Magic Hour, Veo 3, Kling, Hailuo, Runway, and .
Best AI video generation workflow for cinematic commercials
Product-video shot concepts: an amber bottle, a white shoe, and a cream jar
AI product video prompts: 8 shots for product photos
How to Test AI Video Models Without Wasting Credits
AI video pricing index: model rates and cost examples (2026)