

Write an AI video prompt as a shot description: subject, action, setting, camera and lighting. Keep it to one coherent event, and specify sound only if the selected model supports audio. Generate a short draft, identify the main mismatch, then revise one instruction. Avoid asking a single clip to contain several unrelated scenes or contradictory camera moves.
Text-to-video starting prompt: “A red ceramic mug sits beside a closed book on a wooden desk. Steam rises slowly from the mug. The camera moves gently forward at desk height. Soft window light comes from the left.”
For image-to-video using an approved photo of that scene: “Steam rises slowly. The camera moves gently forward; the mug and book remain in place.”
If the camera moves too much, change only that instruction: “The camera stays fixed at desk height.” If the model invents a hand, simplify the request around the mug and steam before adding another action. These are original practice prompts, not tested outputs or quality guarantees.
Problem in the output | What to change first |
|---|---|
Important subject is missing | Put the subject and action at the start of the prompt |
Camera motion overwhelms the action | Use one explicit, restrained camera direction |
Character or product changes | Check the reference and simplify the movement |
Clip jumps between scenes | Reduce unrelated events or split them into separate shots |
Unwanted or missing sound | Check native-audio support and the actual audio settings |
Result ignores the requested duration | Set duration in the tool rather than relying on the prose |
Images are static: a single frame to describe. But video requires motion, continuity, and narrative. This introduces extra challenges:
A short clip does not need a complete story arc. First decide what must change on screen: a person turns, the camera reveals a room or light moves across a product. Add more beats only when the selected model and duration can support them.

Use the structure below as a starting template. Runway’s text-to-video guide recommends describing the essential visual and motion elements first, then refining details. This article’s examples are suggested prompts, not results from a cross-model benchmark.
The goal is a readable shot brief. Omit a field if it adds no useful direction; a list of every possible cinematic adjective is not required.
Write the subject and action first, then add setting, camera and light as needed. For a multi-shot sequence, make a separate storyboard or shot list before combining the clips. Keep the same character and object descriptions where continuity matters.
For image-to-video, let the source image carry appearance and composition, then describe the intended motion. Runway’s current guide makes this distinction explicit: text-to-video needs visual and motion details; image-to-video should focus on movement. Other models may expose additional reference controls.
Focus on camera language and emotional tone.
Prompt example:
“A young woman stands on a cliff at sunset, wind blowing her hair, wide shot from behind as she faces the horizon, soft golden light and sweeping orchestral feel.”
Separate the visible shot from the soundtrack. The cliff, camera position, light and hair movement describe the picture; an orchestral cue only affects generation in a model that supports audio. Otherwise, add music during editing. A cinematic prompt is not evidence that one model outperforms its rivals.

Focus on motion verbs and camera tracking.
Prompt example:
“A basketball player sprints down a narrow alley, the camera glides alongside him in a fast dolly shot, neon graffiti glowing on the walls, sweat glistening under harsh streetlights.”
For sport or martial-arts scenes, inspect body mechanics, contact with the ground, and continuity throughout the motion. Model labels do not guarantee correct physics.
Describe the visual treatment you want to preserve.
Prompt example:
“Three kids running through a mystical forest, rendered in Studio Ghibli watercolor style, vibrant glowing mushrooms, slow panning shot.”
For stylized video, describe the rendering treatment—such as watercolor backgrounds, clean outlines or soft cel shading—and the motion. If you have an approved reference image, use it where supported. Compare one change at a time, such as a close-up versus a wider tracking shot; do not call a result more consistent without reviewing a sequence.
No prompt is perfect on the first try. Iteration is part of the process:
For an automated workflow, prove the prompt on a small representative sample before sending a batch. Our AI video and image API comparison covers integration and cost differences. A batch runs more attempts; it does not establish that the outputs are usable.
Use the shortest prompt that contains the required visual instructions. There is no universal word count or rule that longer prompts produce more realism. Add detail to correct a specific omission, and check the chosen model’s prompt-length limit.

More models accept not just text but image + text prompts. For instance, starting from a photo and layering text instructions for motion.
This is especially useful for YouTube workflows, as explored in resources on faceless YouTube channels.
For a product photo, begin with a modest camera move that keeps the visible packaging in view. A 360-degree rotation forces the model to invent the unseen back and sides. For a landscape, specify one camera move and inspect whether important landmarks remain unchanged. Neither input guarantees an ad-ready result.
A pro tip for creators is to maintain a prompt library: a personal database of tested structures, styles, and scene instructions.
For example:
Save prompts with their reference files, model versions, generation settings and accepted outputs. A prompt that worked in one model is a starting point for another, not a guarantee. For a recurring character or product, keep a small approved reference set alongside the prompts.

Prompt writing isn’t just a technical trick - it’s part of the broader creator workflow:
A clear prompt reduces ambiguity, but the model, source material and edit also matter. When a result fails, decide whether the problem is the wording, an unsuitable input, an unsupported feature or a shot that would be easier to film or edit conventionally.
Check that the request describes a shot the selected tool can produce:
Choose the smallest draft that can answer your main question, then review the entire clip before increasing spend.
For model and workflow choices, see the best AI video generators guide. For dialogue, Google’s Veo prompt guide covers speech and sound instructions. An attractive final frame is not enough: the shot must work through the whole duration.
