

To keep a product consistent in AI videos, begin with an approved product image, define the visual details that cannot change, use image-to-video instead of text-only generation, request one controlled motion per shot, and compare every result with the real SKU before editing. If exact label text, legal copy or packaging geometry must be perfect, preserve or composite the real product rather than asking a model to redraw it.
Product consistency means more than keeping the same color. The bottle, box, garment or device must remain recognizably the same item across frames and shots—even while the camera, setting and surrounding action change.
Create a short acceptance specification before generating. Record the facts a reviewer can verify against the real item:
Silhouette and proportions: overall shape, height-to-width ratio, cap, handle, sole, screen or other defining geometry.
Brand elements: logo placement, label layout, typography zones and required marks.
Color and material: approved color, transparency, gloss, grain, fabric, metal or liquid behavior.
Components: accessories, buttons, seams, closures, package count and included parts.
Relative scale: how the item compares with a hand, face, shelf or another known object.
Do not write “keep the product consistent” and assume every model shares your definition. Name the few visible properties that would make an output unusable if they changed.
Protection level | Use it when | Recommended workflow | Acceptance rule |
|---|---|---|---|
Strict | Real SKU, packaging, regulated claim or catalog asset | Preserve or composite the approved product; generate motion and environment around it | Reject any changed text, geometry, color or included component |
Controlled | Social ad or concept where the real item must remain recognizable | Approved image first, then restrained image-to-video shots | Reject material identity changes; allow harmless environmental variation |
Loose | Mood board, fictional prop or early exploration | Text-to-video or broader transformation | Judge concept and composition rather than SKU accuracy |
Use current, accurate images of the exact SKU. A practical pack includes a clean front view, one side or three-quarter view, a close-up of the label or defining material, and one image that shows scale. Remove obsolete packaging and near-duplicate products from the folder.
Choose one image as the hero reference for each shot. Additional angles help human review and can support models that accept multiple references, but more inputs do not guarantee better identity preservation.
Write two lists. The first contains product invariants that cannot change. The second contains variables the generation may explore: setting, camera move, lighting, particles, hand interaction or background action. This prevents a creative direction from quietly authorizing a product redesign.
When the real product matters, create or approve the first frame before motion. Use the original photograph directly, or make one controlled edit in an AI image editor and compare it with the source. The existing product-photo editing workflow includes prompts and acceptance checks for background, cleanup and composition changes.
Use image-to-video when the source frame already establishes the correct product. Text-to-video is useful for invention, but a written description cannot fully encode an exact label, surface or package shape. The reference reduces what the model must invent; it does not make the product immutable.
Describe the action, environmental motion and one camera behavior. Avoid spending most of the prompt redescribing the object that is already visible. A controlled starting pattern is: “The camera makes a slow left-to-right arc. A narrow highlight travels across the bottle. The product remains rigid and centered. One continuous shot.”
Use these product-video prompt examples or the broader image-to-video prompt guide as starting points. They are templates to adapt, not guarantees across models.
A prompt containing an opening box, pouring liquid, rotating product, moving model and camera transition creates several opportunities for drift. Generate short shots with one visible action. Approve the hardest constraint before producing easier establishing shots.
Small label text, offer copy and legal language are especially fragile. If they must be exact, keep the real packaging visible, composite the approved product render into the generated scene, or add text and graphics in an editor after motion is accepted. Never publish a plausible-looking label without comparing it with the real SKU.
Reuse the same reference pack, model, aspect ratio and product specification. When the workflow supports a final-frame or reference-video input, use an accepted shot to guide continuity. Change the camera or environment separately so a failure has one likely cause. A reused seed or repeated description alone does not guarantee consistency.
Watch the result once for the overall idea, then compare frames with the approved references. Pause at the beginning, middle and end and during any turn, reveal, touch or occlusion. Review the downloaded file rather than relying only on a preview.
Check | Pass | Reject |
|---|---|---|
Shape | Silhouette and proportions match the real item | Cap, handle, sole, screen or package geometry changes |
Branding | Logo and label remain in the correct location | Letters morph, disappear or move |
Color and material | Approved shade and surface behavior remain stable | Color drifts; glass, fabric or metal becomes another material |
Components | Every required part remains present | Buttons, accessories, closures or package count change |
Motion | Product stays structurally stable during camera and environmental movement | Object bends, melts, duplicates or changes scale |
Final export | Correct crop, readable overlays and stable product survive download | Editor preview passes but export introduces a defect |
Symptom | Likely cause | Most useful next step |
|---|---|---|
Label or logo changes | The model is regenerating fine text | Preserve the real product or overlay approved graphics after generation |
Bottle, shoe or device bends | Requested action is too large for the reference | Lock the product; move the camera, light or background instead |
Color shifts between shots | Lighting and style instructions are changing the item | Specify the approved product color separately from scene lighting and compare against a reference |
Product changes during a hand interaction | Occlusion forces the model to reconstruct hidden details | Reduce overlap, shorten the interaction or cut before the product becomes obscured |
Each shot contains a slightly different SKU | Separate generations lack a shared reference system | Reuse the same approved images and specification; change one shot variable at a time |
The output looks attractive but inaccurate | Review focused on aesthetics rather than identity | Add a distinct SKU-accuracy gate before creative approval |
For a simple 12- to 15-second ad, build three independent shots around one approved product image:
Hero shot: locked product, slow camera arc or moving highlight.
Detail shot: restrained push-in on a real material, component or use detail.
End shot: stable product with empty space for exact offer copy added in editing.
Generate and approve each shot separately. Assemble them in an editor, then add the real logo, price, disclaimer and call to action. For more structures, see 12 product-video examples and shot plans.
Condensation moves slowly down the bottle while mist drifts behind it. The camera makes a restrained left-to-right arc. Keep the product, label, shape, colors, and materials unchanged. One continuous shot, no new objects or text.

Workflow recipe
The camera trucks (moves laterally) at the same high speed, perfectly parallel to the car and its driver. The car stays locked in the center of the frame. In the second half of the shot, the car executes a quick, slight S-curve maneuver, kicking up a small plume of white salt dust from its tires. The camera perfectly mirrors this S-curve, moving with the car to keep it locked in the center of the frame, before both straighten out again. Cinematic, car commercial, high-energy.
Upload the exact product photo, describe one controlled movement, and inspect the result before building the full edit.
Try Image-to-VideoThe cheapest generation is not always the cheapest usable result. For every production brief, record total attempts, failed jobs, accepted shots, credits or spend, generation time and repair time. Divide total generation cost by accepted shots, and keep labor visible rather than converting it into a made-up universal rate.
Magic Hour’s AI video pricing index explains why model rate, duration, resolution and audio settings must be normalized. Its 60-attempt commercial benchmark publishes prompts, source images, output files and technical completion, but does not claim that every completed clip passed human SKU review.
This is a production method based on observable product requirements and documented workflow constraints. Magic Hour publishes this guide and links to its own tools. It is not a controlled claim that Magic Hour or any model produces the most consistent product video. Output varies by source image, prompt, model and motion. A current research paper on product identity preservation likewise treats product consistency and text accuracy as distinct evaluation problems; its reported model results should not be generalized to every video workflow.
A generation model may reinterpret pixels across time, especially during large motion, camera changes or occlusion. Separate shots also do not automatically share an exact product identity.
Image-to-video is usually the safer starting workflow when an approved product image already defines the item. It reduces invention but does not guarantee exact labels, geometry or color.
No. A prompt can direct motion and emphasize invariants, but strict SKU accuracy still requires reference images, restrained shot design and human comparison with the real product.
Start with one strong hero image for the shot and keep additional front, side, detail and scale views for review or supported multi-reference workflows. Use only images of the exact current SKU.
Preserve the real labeled product whenever possible. When exact text cannot survive generation, composite the approved product or add the logo and copy in an editor after motion is complete.
Reject changes that alter what the buyer would receive or violate the approved brand asset. Harmless changes in reflections, background motion or framing may be acceptable when the product itself remains accurate.
