

A realistic AI video prompt should describe one observable shot: subject, action, setting, camera and light. Start simple, generate a short draft, identify the most important error and change one instruction. Example: “A cyclist stops beside a café and places one foot on the pavement. Locked eye-level camera, overcast daylight, six-second shot.” Realism still depends on the model, source and review; prompt wording cannot guarantee it.
The ten embedded videos were present and available when checked September 13, 2026. They are visual references, not retained proof that the prompts below produced them. The named endorsements, “proven” claims and time-saving estimates from the previous version were removed because no reproducible test record supported them.
Runway’s current prompting guide likewise recommends beginning with essential motion and adding subject, camera, scene or style detail one element at a time. Its image-to-video guide says the image already defines composition, subject, lighting and style, so the text should focus on motion and temporal progression.
Subject + action + setting + camera + light + duration and shape + invariants.
Use an invariant only for something that must remain stable, such as identity, product geometry, readable text or chronology. Negative phrases are not consistently honored across models. Prefer a positive state such as “locked camera” or “product remains centered and still,” then verify the result.
Start with one subject, one action and one camera instruction. Review a short output, then change only the instruction tied to the most important error.
Try AI Video GeneratorAn unbranded red stand mixer runs on a kitchen counter. One hand moves the speed lever once. Handheld phone camera moves slightly closer. Soft window light from the right. Seven-second vertical shot.
Review: mixer shape, bowl and lever; hand anatomy; batter motion; contact shadows; any invented text or logo.
Prompt: A person pours coffee into a plain ceramic mug, pauses as steam rises, then sets the pot down. Slow push-in from waist height. Warm morning window light. Eight-second landscape shot.
Review: hands, cup and pot geometry; liquid direction; steam; face and wardrobe continuity; extra actions.
Prompt: Macro close-up of an unbranded steel blade resting on dark wood. Camera slides slowly from the edge toward the guard while one soft rim light moves across the metal. Five-second shot.
Review: edge and guard shape; reflections; material texture; focus behavior; geometry drift.
Prompt: Start on a blurred plant at frame left. One fast whip pan lands on a plain product centered on a table and holds for the final second. Lighting remains unchanged. Four-second vertical shot.
Review: product shape during blur; landing composition; background continuity; exposure change; clean hold frame.
Prompt: Top-down view of two hands placing a blank faceplate over a mounting plate and tightening one screw. Locked camera, even light, continuous eight-second shot.
Review: step order; tool and hardware shape; fingers; cause and effect; whether generated action could be mistaken for real instruction.
Prompt: Slow gimbal move past one textured gym bench toward a shaft of window light. One person crosses the distant background out of focus. Six-second landscape shot.
Review: bench and room geometry; background person; lighting direction; camera path; flicker.
Prompt: A fictional city crosswalk at dusk. Handheld camera follows one person for three steps while other pedestrians cross naturally. No readable signs or brands. Six-second shot.
Review: faces, feet and collisions; traffic direction; signs; crowd continuity; misleading real-location cues.
Prompt: One authorized presenter speaks one short sentence to camera in a plain office. Locked eye-level medium close-up, soft key from the left, unchanged background.
Review: consent; mouth timing; teeth, eyes and face identity; voice rights; disclosure; caption accuracy.
Prompt: A small paper rocket circles a plain blue sphere once. Fixed centered camera, flat illustration, simple background, end composition matches the start. Five-second square loop.
Review: complete orbit; object count; scale; start-to-end jump; small-screen readability.
Prompt: Preserve the supplied bottle, cap, label and colors. Bottle remains centered and still while one soft band of light crosses the background. Slow camera push-in. No new text or hands. Six-second landscape shot.
Review: label spelling; bottle and cap geometry; colors; reflections; contact shadow; newly invented details.

Workflow recipe
Cinematic slow motion of a ripe strawberry sliding down a swirl of whipped cream. Cream folds softly as the berry glides. High-contrast lighting, glossy reflections and a creamy bokeh background.
Subject changes shape: use an approved reference image, reduce rotation and motion, and state the exact feature that must remain stable.
Too much happens: remove secondary actions and keep one subject action plus one camera move.
Camera drifts: describe the positive camera state—locked, fixed point, slow push or left-to-right tracking—and give the subject something visible to do.
Unexpected cut: remove shot-change language, request one continuous shot and fit the action to the selected duration.
Hands or interactions fail: simplify the contact, shorten the action, choose an angle with fewer occlusions or use real footage when the interaction must be accurate.
Product text warps: use a reference image and restrained motion, or composite the real packshot and final typography in an editor.
Loop jumps: reduce independent motion and trim in an editor; asking for a loop does not prove the last and first frames match.
In text-to-video, describe both appearance and motion because the model must create the scene. In image-to-video, the image supplies appearance and composition; concentrate the prompt on what changes over time. If exact product or identity preservation matters, use an authorized reference and inspect every frame.
Choose one difficult shot. Include the identity, interaction, text or camera behavior that matters to the real project.
Lock comparable settings. Keep source, duration, aspect ratio and model fixed while testing wording.
Start with the minimum prompt. Subject action and camera are enough for the first attempt.
Change one variable. Add or revise only the instruction tied to the largest visible error.
Save every attempt. Record prompt, model, seed if exposed, settings, charge, output and why it passed or failed.
Count accepted-output cost. Include rejected generations, correction time, licensing and finishing.
Not automatically. Extra instructions can conflict or obscure the main action. Start with the minimum observable shot, then add one detail after inspecting the result.
Use a lens or camera term only when it communicates a visible result you need, such as macro framing, shallow depth of field or handheld motion. A model may treat numeric camera metadata as style language rather than a physical capture setting, so judge the output instead of assuming the number was obeyed.
No prompt guarantees exact preservation. Use an authorized reference, limit motion and occlusion, inspect every frame and use conventional compositing or editing when exact identity, label or geometry is required.
Reduce the number of simultaneous events, use plausible motion, keep the camera instruction clear and avoid excessive effects. Real source footage remains the strongest choice when the video must prove a product action or factual event.
For more shot patterns, use the AI video examples library. For camera vocabulary, compare the AI video camera angles and prompts.
