AI Video Reliability: How a Deletion Filter Nearly Halved the Measured Error Rate


Excluding deleted attempts nearly halved the recorded image-to-video error rate for the same 118,008 Magic Hour account IDs: 11.81% with deleted attempts included, versus 6.06% for retained attempts only. The measured rate fell 48.7% relative to the inclusive rate. No model improvement was needed; the filter changed what entered the calculation.
Original Magic Hour platform data. Requests created June 1–August 31, 2026 UTC; extracted September 7, 2026. This is a historical observational snapshot, not current performance, a provider league table or an industry-wide failure estimate.
Same Accounts, Different Denominators

Figure 1. Stored ERROR status divided by all requested attempts in each scope. The same account cohort and creation window are held fixed. Bars use a common zero-to-15% scale; lower is a different measurement, not evidence of better outputs.
Request Scope | Account IDs | Requested Attempts | Recorded ERROR | ERROR / Attempts |
|---|---|---|---|---|
Including deleted attempts | 118,008 | 448,822 | 53,028 | 11.81% |
Retained attempts only | 118,008 | 368,395 | 22,314 | 6.06% |
The difference is 5.76 percentage points. The 48.7% relative reduction is 100 × [1 − (22,314 / 368,395) / (53,028 / 448,822)]. It is not a 48.7-percentage-point change. No reasons for deletion or characteristics of deleted content were inferred.
Why This Matters for AI Benchmarks
A library containing only projects that still exist is not automatically a complete record of attempted generations. If failed attempts disappear from the observation set, a reported error rate can change before any model, prompt or infrastructure is improved. This dataset supplies one reproducible example of that measurement problem.
When comparing reliability reports, ask what counts as an attempt, which statuses remain in the denominator, whether deleted-attempt metadata is included, and whether the account population is fixed. A completed job also does not establish that a person accepted the output. Completion, policy handling, technical reliability and output quality need different measures.
Cohort and Methodology
The source was Magic Hour’s hosted production-database replica. We selected IMAGE_TO_VIDEO projects with explicit web-mode metadata and creation times from June 1 inclusive to September 1 exclusive, UTC. Staff accounts, deleted accounts and accounts with deletion requests were excluded. Eligibility required at least one retained request in the window. Both rows contain exactly the same 118,008 account IDs; deleted-only accounts do not enter the comparison.
Rates are request-weighted: total records with stored ERROR divided by total requested projects in the relevant scope. Canceled and unresolved states remain in the denominator. A non-ERROR record is not necessarily a completed or usable output. Stored ERROR can include policy-coded rejections and other classifications; this report does not establish technical root causes, actual policy violations or false positives.
A separate direct aggregate query reproduced the retained-cohort totals. Removing the five highest-volume accounts preserved the direction: 11.43% including deleted attempts versus 5.90% retained only. The direction also held in four model strata with at least 1,000 eligible accounts. These checks support the measurement finding, not causal attribution or a model ranking. The public two-row file reproduces headline arithmetic, not those private account-level checks.
Privacy and Responsible Interpretation
All grouping occurred in the warehouse. No customer prompts, inputs, outputs, identities, account IDs, project IDs, emails, IP addresses or individual timestamps appear in the public dataset. Deleted records were used only for aggregate metadata counts; no deleted prompts or media were retrieved. Released cohorts and binary account-outcome cells passed a 100-account minimum. This is aggregation and suppression, not differential privacy.
Accounts mean database account IDs, not deduplicated people or organizations. Magic Hour users are not representative of everyone using AI media. The analysis is exploratory and observational, not randomized or preregistered. It does not show why people deleted projects, how satisfied they were, or which provider is intrinsically more reliable. No naive binomial confidence interval is attached to repeated-request percentages.
The practical lesson is to document the observation set and denominator in privacy-preserving measurement. It does not justify retaining user content after deletion. Any operational attempt ledger must follow applicable deletion, retention and privacy commitments.
Download, Reproduce and Cite
Download the exact two-row CSV and read the fixed data card with formulas. Recompute each rate as errors ÷ jobs × 100. The downloadable values preserve the September 7 extraction; this October 5 publication does not change the observation window.
Suggested citation: Magic Hour Research (2026), AI Video Reliability Measurement: Deletion Sensitivity, June–August 2026 web-mode image-to-video cohort, September 7 snapshot. Include the fixed account cohort, requested-attempt denominator and distinction between measured ERROR and technical failure. Original aggregate CSV, data card and figure are reusable under CC BY 4.0 with Magic Hour attribution.
This is a fixed edition, not a live feed. Future editions should repeat the same eligibility rules, full observation windows, error semantics, concentration checks and privacy thresholds, then release new dated files. Aggregate arithmetic is publicly reproducible; private source-record extraction requires authorized warehouse access.
Related Original Research
The 60-attempt model benchmark exposes downloadable outputs and observed account charges for a controlled set of scenarios. The AI video price index compares calculated listed API charges for matched requested settings. Neither measures cost per human-accepted output.






