AI video prompt troubleshooting: diagnose motion, identity and camera failures

Runbo Li
Runbo Li
·
· 10 min read
Editorial film contact sheet showing a running subject, a circled motion error, and the corrected frame

Quick answer

When an AI video prompt fails, simplify the shot and change one variable at a time. Diagnose subject, action, camera, environment, style and exclusions separately instead of rewriting the entire prompt after every result.

A prompt is only one input. Before rewriting it, identify whether the miss came from the source image, prompt, model settings, or an unsupported requirement. This guide is for a failed short shot; the realistic-prompt guide covers writing a first prompt, while the character-consistency guide covers multi-shot identity.

For initial prompt patterns, read realistic AI video prompts. For source and identity problems, see character consistency. To clear rights before publishing, use the AI video commercial-use guide.

What a good result looks like

A useful fix changes one observable defect while keeping the rest of the shot intact. Save the rejected clip and record the first timestamp where it fails. Write an acceptance check before regenerating: for example, “the runner remains the same person and the ball stays in frame through the cut.” Do not grade the clip only by its best frame.

The shortest diagnostic record has five fields: source or reference, exact prompt, selected model and settings, first visible defect with timestamp, and the single change made next. If the second result improves one trait but damages another, keep both results; that tradeoff tells you which constraint the current setup handles poorly.

Before you start

Choose the right starting material. Text-to-video invents the opening frame; image-to-video starts from a supplied image. If a product label, character face, or composition must stay exact, a reference image can be more useful than another paragraph of adjectives. A reference guides output; it does not guarantee frame-by-frame preservation. For an already filmed action that must remain recognizable, consider video-to-video instead.

Separate hard requirements from aesthetic preferences. A brand mark, identity, spoken line, or legally required disclosure may need manual review or conventional editing even when the generated motion looks good. Test the hardest few seconds first, with the destination aspect ratio and duration selected in controls. Asking for a longer clip in prose does not override a short duration setting.

Step-by-step workflow

1. Write the visible failure in one sentence

Pause at the first wrong frame. Name what a viewer can see: “face changes at 0:03,” “camera rotates while the subject walks,” or “label letters mutate on the turn.” “Looks bad” does not tell you which input to change. Note whether the error appears from the start or only after a motion, occlusion, cut, or time extension.

2. Reduce the shot to one subject and one action

Rewrite the prompt as a single shot with one main action and an explicit endpoint. Instead of “a runner scores, celebrates, and the camera flies over the crowd,” try “A runner crosses the finish line and slows to a stop; medium side view.” Make the celebration a second clip. Fewer simultaneous events make failures easier to identify and edit.

3. Separate camera movement from subject movement

Subject motion and camera motion are separate instructions. If the athlete runs and the camera also circles, first test the athlete with a fixed camera. If the action becomes readable, restore one modest camera move. If it does not, simplify the subject action or shorten the shot. Do not change framing, motion, model, and style in the same retry.

4. Move fixed traits into the reference or opening description

If a face, outfit, package, or set layout matters, use a clear authorized source or reference when the workflow supports one. Keep the identity description and source fixed while testing motion. For image-to-video, describe what changes after the first frame rather than repeatedly asking the model to redraw what the image already supplies. Inspect turns, occlusion, and the last frame—not just the opening.

5. Add only one measurable constraint

Add a constraint you can verify by looking at the export: “the blue bottle remains in the left hand,” or “one continuous side view.” Avoid a long list of abstract negatives such as “no glitches, no distortion, no mistakes.” If exact text is essential, expect to validate it in every frame or add the text in a conventional editor afterward.

6. Save the prompt, model, settings and output

Keep the rejected output beside its prompt, model, duration, aspect ratio, reference, and other settings. Label the revision with one change and its reason. If two unchanged retries fail in the same way, do not assume a third run will repair the underlying prompt or source. Try a smaller motion, better source, shorter clip, or a different supported workflow.

Test one text-to-video shot

Start with one subject, one action, and one camera instruction. Keep the output and change one variable before your next attempt.

Try Text-to-Video

Practical examples

Example 1: A running subject warps during camera motion

Illustrative repair, not a tested Magic Hour result: “A sprinter accelerates down the track. Handheld camera follows, then orbits around her as she leaps and the crowd erupts.” If the face warps during the orbit, first try “A sprinter accelerates down the track. Fixed medium side view; the camera does not move.” Hold the source, model, duration, and style constant. If that works, add only a gentle tracking move in a later attempt.

Example 2: A product label changes during movement

Illustrative repair, not a measured output: “A bottle spins in midair while the camera zooms through splashing water; the label is perfectly readable.” This asks the model to preserve tiny exact lettering during complex motion. Try an authorized product image with a stationary bottle and a slow push-in; verify the label frame by frame. If exact package copy still changes, composite the real label or real product footage in an editor rather than promising a prompt-only fix.

Example 3: Dialogue and gestures drift together

Illustrative repair, not a measured output: “Two people argue in a crowded café, gesture rapidly, and the camera circles them.” If gestures, lips, and background people drift together, make one short reaction shot per speaker with a static camera and quieter background. Generate the action and edit the dialogue sequence separately. Review any spoken words, lip sync, and identities against the original brief before use.

Common mistakes

  • Writing a multi-scene story as one shot

  • Stacking conflicting camera instructions

  • Using abstract adjectives instead of visible details

  • Changing model and prompt simultaneously

  • Adding long negative lists before fixing the core action

Triage failures without wasting retries

Use a symptom-to-action order. Wrong identity at frame one: replace or clarify the source. Identity changes only on a turn: reduce occlusion or movement and shorten the shot. Wrong action: describe one visible verb and endpoint. Camera chaos: remove competing camera instructions. Text mutates: reduce motion or finish text in an editor. Failure to honor duration or ratio: check selected controls, not just prompt wording.

A fix is useful only if the specific defect improves without breaking the hard constraints. Keep the original and revised clip side by side, and record whether the change improved identity, action, camera, text, or timing. “Better” without a named criterion is too easy to overstate. If the core action remains impossible after a controlled retry, change workflow or scope rather than spending on random adjectives.

A review checklist you can reuse

  • The opening answer matches the actual question and intended user

  • Every factual claim, label, name, number and link is checked

  • Source rights, consent and commercial terms are recorded

  • The complete output passes identity, continuity, text, audio and delivery review where relevant

  • All retries, failures, processing time and finishing work are retained

  • The CTA or next step points to the exact workflow described

How to measure the workflow

For team work, measure accepted outputs, not just generations. Define the acceptance checklist before testing, then record attempts, failed jobs, elapsed time, manual finishing, and exports that pass. Cost per accepted clip is total generation plus finishing cost divided by accepted exports; a cheap first run may be expensive if it takes many retries. Do not treat these examples as benchmark results—we have not run a controlled model comparison here.

To compare two settings or models fairly, hold the brief, source, duration, ratio, and acceptance rule constant where each option supports them. Record version and date because capabilities change. Review full videos, including the first and last frames. A hand-picked vendor sample does not establish comparative performance on your brief.

Accuracy, consent and disclosure

Use only inputs you are allowed to upload or transform. For branded work, check source footage, music, logos, people, and voice separately. A paid generation plan does not grant third-party rights or guarantee that a clip is suitable for an ad. Keep the permission and approval record with the prompt and exported asset.

Disclosure and provenance depend on the destination and context. When realistic synthetic media could mislead viewers about a real event or person, check that platform’s current disclosure controls and applicable rules. A disclosure does not replace consent. The commercial-use rights guide covers the broader clearance checklist.

How to keep the process reproducible

Save a compact attempt log: one row per run with the source, prompt, model/settings, first defect, single change, and pass/fail result. Keep rejected examples. They help another editor avoid repeating the same failure and expose cases where the source or required shot—not wording—is the limiting factor.

When a change helps, freeze the working version before experimenting further. If you also change the source, model, or aspect ratio, start a new comparison. Otherwise the team cannot tell which change caused the improvement. For production, test at least one difficult case, not only a clean showcase.

How to scale after the first pass

Turn an accepted shot into a template only after checking it against a second representative source. Reuse stable instructions and settings, then document where they fail. Different faces, motion speeds, reflective packages, and camera angles may require different source or editing choices.

Review every finished export that carries a real person, product claim, or brand requirement. Track the failure categories over time. If most rejections cluster around one action or source type, fix that input or choose a more suitable workflow before generating more volume.

Decision framework

Choose text-to-video when the scene can be invented, image-to-video when the supplied opening composition matters, and video-to-video when real source motion should guide the result. Use a conventional editor when exact typography, legal copy, product geometry, or timing must be deterministic. The choice of workflow can matter more than a longer prompt.

For a fast first pass, start with one subject, one action, one camera instruction, and a short clip. Add a reference only when the current route supports it and appearance matters. After that, change a single variable per retry. Stop when the clip passes the prewritten criteria or the workflow clearly cannot meet them.

Handoff notes for a team

Hand the next editor both an accepted clip and a rejected clip, with the exact source, prompt, settings, and failure notes. Include the product route and any rights constraints. This makes the workflow reproducible without pretending one prompt works for every model or source.

At sign-off, ask who checked identity, motion, text, rights, and destination requirements. Record unresolved limitations explicitly. A visually impressive preview is not an approved deliverable if a label changes, a face drifts, or a required disclosure is absent.

Official sources and product references

Magic Hour Text-to-Video documentation: tool and API details. For web-app controls and limits, see Text-to-Video step-by-step guide.

Magic Hour Image-to-Video documentation: tool and API details. Runway’s official video prompting guide separately explains camera and subject motion for its model; treat that as vendor guidance, not a universal guarantee.

Frequently asked questions

The prompt may contain too many simultaneous actions or conflicting instructions. Reduce it to one observable shot.

Use concise exclusions after the positive shot is clear; long negative lists can compete with the main instruction.

Save the exact model, version, settings, input asset and accepted output with the prompt.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Cinematographer notebook illustrating realistic AI video prompting with shots, lenses, movement, and continuity
Recommended next
Realistic AI video prompts: 10 patterns with examples

Adapt 10 concise AI video prompt patterns for products, UGC, transitions, process shots and loops, with exact review checks and a controlled test.

Cinematic AI Video Prompt Cookbook
25 cinematic AI video prompts for motion and composition
Image-to-video prompt workflow shown as an editorial motion contact sheet
Image-to-video prompts: 30 examples, templates and fixes
Characters, Styles, and Camera Movement
Kling 3.0 reference guide: Elements, Multi-Shot & prompts
Keep Characters Consistent in AI Video
How to keep characters consistent in AI video