How to prompt AI videos: a practical 10-step guide

Prompting AI Videos Cover

Quick answer

Write an AI video prompt as one filmable shot: subject, action, setting, camera behavior and visual treatment. Add audio only when the selected model supports it. In a Text-to-Video workflow, set duration, aspect ratio, references and other hard controls in the interface; words in the prompt do not override unavailable settings.

A useful starting template is: “[subject] [action] in [setting]. [shot size and camera movement]. [lighting and visual treatment]. [audio direction, if supported].” Keep every instruction compatible with a short clip.

Starting from an approved still image? Use the dedicated image-to-video prompt guide with 30 examples; the image already defines appearance, so the prompt should concentrate on motion, camera behavior and continuity.

Example AI video prompt

A cyclist rides along a quiet coastal road at sunrise. The camera follows steadily from behind in a medium-wide shot. Natural light, gentle motion, no scene change. Ambient wind and distant waves.

Workflow recipe output preview: Kung Fu in Sakura Forest

Workflow recipe

Cinematic wide shot with a gentle upward tilt and slow drifting motion, a highly detailed anime-style young hero with wind-swept hair standing on a hilltop, then transitioning into fluid kung fu movements—graceful strikes, spins, and controlled breathing forms—his stance shifting with precision and power, eyes full of determination, cherry blossom petals swirling dynamically around him, reacting to each motion like flowing energy, set against a dramatic sunset sky in rich orange and pink gradients, distant mountains fading into soft atmospheric perspective, action blending martial arts choreography with elegance and rhythm, style: vibrant colors, soft cel-shading, expressive facial features, clean anime line art, cinematic lighting casting a warm golden glow, dynamic composition with motion trails and wind effects, Studio Ghibli x Makoto Shinkai aesthetic, ultra-detailed environment, emotional, epic, and slightly mystical ambiance.

Model
veo3.1
Format
16:9
Resolution
1080p
Duration
6 seconds

If audio is unavailable in the selected model, remove the final sentence and add sound later. If a supplied image must be the starting frame, use image-to-video instead of trying to describe the image again.

1. Decide what must remain fixed

Before writing, separate hard controls from creative instructions. Duration, resolution, aspect ratio, seed, reference inputs, audio mode and safety controls vary by product. Use the visible controls and current documentation. Magic Hour’s current Text-to-Video guide explicitly notes that requesting a longer movie in the prompt does not override the selected duration.

2. Describe one primary subject

Name who or what the shot is about and include only details that matter to recognition: age range or role, clothing, material, color, shape or product attributes. A long inventory of equally important subjects makes the intended focus unclear.

3. State a visible action

Use an action that can be shown: walks, turns, pours, unfolds, looks up or moves toward the camera. Abstract goals such as “be inspiring” need a visible performance, setting or camera choice.

4. Anchor the setting

Specify the place, time of day and a few scene elements that affect the shot. “A bakery kitchen before sunrise, flour on the steel counter, warm practical lights” gives the model usable visual context without dictating every object.

5. Give the camera one job

Choose a shot size and one camera behavior: static wide shot, handheld close-up, slow dolly in, tracking from behind or overhead lock-off. Multiple incompatible moves in a short generation can create an unclear result. Google’s Veo prompt guide likewise separates subject, action, context, camera and audio elements.

6. Specify the visual treatment

Describe observable choices such as soft overcast light, shallow depth of field, muted documentary color or stop-motion paper texture. Generic tags such as “masterpiece” or “best quality” do not define what the finished shot should look like, and their effect cannot be assumed across models.

7. Treat references as inputs, not adjectives

When a workflow supports a reference image or video, provide the permitted asset through that control and say what it should guide. A reference may guide appearance or structure without locking identity, text or geometry, so inspect the output.

Magic Hour official workflow page screenshot (magic-hour-textvideo), September 25, 2026

Magic Hour input workflow, captured September 25, 2026. This shows the upload or prompt controls, not a completed generation. View official source

8. Prompt audio separately when supported

Describe dialogue, sound effects and ambience in a separate sentence. Quote exact dialogue, name the speaker and keep the line short enough for the selected duration. Do not assume every video model generates synchronized audio.

9. Generate a representative proof

Use the shortest supported setting that still contains the difficult action. Record the complete prompt, model or mode, references and controls. Runway’s current prompting introduction frames prompting as an iterative review process; the useful unit is a documented change and its output, not an unexplained pile of variants.

10. Revise the largest mismatch

  • Wrong subject or object: simplify the subject and move its identifying details earlier.

  • Wrong action: replace an abstract verb with one visible movement.

  • Unstable composition: reduce competing subjects or camera instructions.

  • Wrong look: replace vague quality words with lighting, material, lens or color language.

  • Reference drift: verify that the chosen workflow supports the reference type and test a clearer permitted input.

  • Text or identity errors: treat them as review failures; prompting cannot guarantee exact text or likeness.

A repeatable prompt-testing method

  • Define pass/fail checks. List the required subject, action, composition, protected details and prohibited changes.

  • Create a baseline. Generate once with the simplest complete prompt.

  • Change one variable. Revise the subject, action, camera, setting or treatment, but not all of them together.

  • Review the entire clip. Inspect motion, identity, edges, text, anatomy, continuity and audio rather than one frame.

  • Log accepted output. Keep the prompt, settings, source assets, cost shown and manual fixes for the result you would actually use.

This method works across models because it does not pretend their prompting syntax or controls are identical. For each product, read the current model-specific guide and verify supported inputs before carrying a prompt over.

Test one clear video prompt

Generate one short representative shot, compare it with your acceptance criteria, then revise the single largest mismatch before increasing duration or quality.

Open Text to Video

Frequently asked questions

Long enough to define one shot and its required constraints. There is no reliable universal word count. Remove details that do not change a visible or audible part of the result.

Put the subject and action early because they clarify the shot for readers and models. Do not assume a fixed weighting rule applies to every current video model.

It can be a baseline, but controls, supported references, audio and model behavior differ. Keep the brief constant, map it to each interface and compare the accepted outputs.

Use a permitted reference workflow that explicitly supports character guidance, keep identifying details and wardrobe stable, and review every shot. Text alone does not guarantee identity across separate generations.

Check whether the request is supported by the selected mode, remove conflicts, simplify the shot and revise the largest mismatch. Save the exact settings and output so support can reproduce a persistent issue.

Aastha Kochar - author at MagicHour (SaaS MarTech Content Writer)
Aastha Kochar
Content Manager
Aastha Kochar has spent 5+ years creating content for B2B and B2C SaaS brands in the AI and MarTech space. She is well-versed with AI-powered content tools and offers deep comparisons after trying and testing every tool. Her work has helped companies increase organic traffic, earn AI citations, and most importantly — turn readers into users. With a bachelor's and master's degree in Journalism and Mass Communication, she brings strong research skills, authentic storytelling, and a deep understanding of what makes audiences actually care about what they're reading.
View author →

Continue Reading

Sports zine collage showing selection, pacing, and editing for a highlight reel
How to make a sports highlight reel: recruiting & social
Editorial cover for 10 best face swap apps, with phone screens illustrating face replacement.
10 best face swap apps (2026): photos, videos and mobile
Analog filmmaker workbench comparing AI video generation workflows and finished frames
10 best AI video generators in 2026: models, features, and costs
Picture of Multiple Faces
How to swap faces: photo and video guide for 2026
How to Use Kling 3.0
How to use Kling 3.0: 7 steps for a controlled first clip
Luma Dream Machine
Luma Dream Machine AI review (2026): Ray3.2 & current app