AI Video Reliability: How a Deletion Filter Nearly Halved the Measured Error Rate

Runbo Li
Runbo Li
·
· 5 min read
Magic Hour original platform data: What Gets Counted? Dark editorial ledger illustration with one excluded row, lime and violet accents.

Excluding deleted attempts nearly halved the recorded image-to-video error rate for the same 118,008 Magic Hour account IDs: 11.81% with deleted attempts included, versus 6.06% for retained attempts only. The measured rate fell 48.7% relative to the inclusive rate. No model improvement was needed; the filter changed what entered the calculation.

Original Magic Hour platform data. Requests created June 1–August 31, 2026 UTC; extracted September 7, 2026. This is a historical observational snapshot, not current performance, a provider league table or an industry-wide failure estimate.

Same Accounts, Different Denominators

Same 118008 Magic Hour account IDs, June–August 2026 web image-to-video: 53028 ERROR of 448822 requests including deleted attempts, 11.81%; retained only 22314 of 368395, 6.06%. Excluding deleted attempts reduced measured rate 48.7%, not actual reliability. Historical September 7 aggregate snapshot; no prompts or media inspected.

Figure 1. Stored ERROR status divided by all requested attempts in each scope. The same account cohort and creation window are held fixed. Bars use a common zero-to-15% scale; lower is a different measurement, not evidence of better outputs.

Request Scope

Account IDs

Requested Attempts

Recorded ERROR

ERROR / Attempts

Including deleted attempts

118,008

448,822

53,028

11.81%

Retained attempts only

118,008

368,395

22,314

6.06%

The difference is 5.76 percentage points. The 48.7% relative reduction is 100 × [1 − (22,314 / 368,395) / (53,028 / 448,822)]. It is not a 48.7-percentage-point change. No reasons for deletion or characteristics of deleted content were inferred.

Why This Matters for AI Benchmarks

A library containing only projects that still exist is not automatically a complete record of attempted generations. If failed attempts disappear from the observation set, a reported error rate can change before any model, prompt or infrastructure is improved. This dataset supplies one reproducible example of that measurement problem.

When comparing reliability reports, ask what counts as an attempt, which statuses remain in the denominator, whether deleted-attempt metadata is included, and whether the account population is fixed. A completed job also does not establish that a person accepted the output. Completion, policy handling, technical reliability and output quality need different measures.

Cohort and Methodology

The source was Magic Hour’s hosted production-database replica. We selected IMAGE_TO_VIDEO projects with explicit web-mode metadata and creation times from June 1 inclusive to September 1 exclusive, UTC. Staff accounts, deleted accounts and accounts with deletion requests were excluded. Eligibility required at least one retained request in the window. Both rows contain exactly the same 118,008 account IDs; deleted-only accounts do not enter the comparison.

Rates are request-weighted: total records with stored ERROR divided by total requested projects in the relevant scope. Canceled and unresolved states remain in the denominator. A non-ERROR record is not necessarily a completed or usable output. Stored ERROR can include policy-coded rejections and other classifications; this report does not establish technical root causes, actual policy violations or false positives.

A separate direct aggregate query reproduced the retained-cohort totals. Removing the five highest-volume accounts preserved the direction: 11.43% including deleted attempts versus 5.90% retained only. The direction also held in four model strata with at least 1,000 eligible accounts. These checks support the measurement finding, not causal attribution or a model ranking. The public two-row file reproduces headline arithmetic, not those private account-level checks.

Privacy and Responsible Interpretation

All grouping occurred in the warehouse. No customer prompts, inputs, outputs, identities, account IDs, project IDs, emails, IP addresses or individual timestamps appear in the public dataset. Deleted records were used only for aggregate metadata counts; no deleted prompts or media were retrieved. Released cohorts and binary account-outcome cells passed a 100-account minimum. This is aggregation and suppression, not differential privacy.

Accounts mean database account IDs, not deduplicated people or organizations. Magic Hour users are not representative of everyone using AI media. The analysis is exploratory and observational, not randomized or preregistered. It does not show why people deleted projects, how satisfied they were, or which provider is intrinsically more reliable. No naive binomial confidence interval is attached to repeated-request percentages.

The practical lesson is to document the observation set and denominator in privacy-preserving measurement. It does not justify retaining user content after deletion. Any operational attempt ledger must follow applicable deletion, retention and privacy commitments.

Download, Reproduce and Cite

Download the exact two-row CSV and read the fixed data card with formulas. Recompute each rate as errors ÷ jobs × 100. The downloadable values preserve the September 7 extraction; this October 5 publication does not change the observation window.

Suggested citation: Magic Hour Research (2026), AI Video Reliability Measurement: Deletion Sensitivity, June–August 2026 web-mode image-to-video cohort, September 7 snapshot. Include the fixed account cohort, requested-attempt denominator and distinction between measured ERROR and technical failure. Original aggregate CSV, data card and figure are reusable under CC BY 4.0 with Magic Hour attribution.

This is a fixed edition, not a live feed. Future editions should repeat the same eligibility rules, full observation windows, error semantics, concentration checks and privacy thresholds, then release new dated files. Aggregate arithmetic is publicly reproducible; private source-record extraction requires authorized warehouse access.

Related Original Research

The 60-attempt model benchmark exposes downloadable outputs and observed account charges for a controlled set of scenarios. The AI video price index compares calculated listed API charges for matched requested settings. Neither measures cost per human-accepted output.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

AI Video Model Benchmark

AI Video Model Benchmark: 60 Attempts, a 12× Credit Spread

AI Video Pricing Index cover with illustrated video cards and the subtitle How to Test AI Video Models Without Wasting Credits.

AI Video Price Index: What 100 Clips Cost

What Changed This Quarter and What Actually Matters

AI video model release tracker (2026): launch dates and access

Ghibli-style AI animated person wearing large headphones at a desktop computer with a microphone in a dim, RGB-lit room.

Survey: Marketers Say AI Saves Time, Yet 25% Report More Burnout

AI video production workflow from storyboard to finished mobile video

How to make AI videos: a complete workflow for 2026

Analog filmmaker workbench comparing AI video generation workflows and finished frames

10 best AI video generators in 2026: models, features, and costs