How to use reference images for image-to-video

Runbo Li
Runbo Li
·
· 4 min read
Use Reference Images in Image-to-Video

Quick answer

Use a reference image as the starting frame when you want to animate an existing composition. Add an end frame, subject image, style image or motion reference only when the selected model explicitly supports that role. A reference guides generation; it does not guarantee that identity, text, logos, products or unseen angles will remain exact.

Choose the reference pattern before the model

Reference pattern

Use it for

What it controls

Main risk

Starting image

Animating one still

Initial subject, scene and composition

Unseen angles and motion may be invented

Start and end images

A planned transition

Initial and final visual states

The path between frames can distort

Multiple references

Subject, object, style or motion guidance

Only the roles exposed by the selected model

Conflicting references can reduce consistency

Motion reference

Reusing movement or camera behavior

Performance or motion when explicitly supported

Appearance may drift without a separate subject reference

The first pattern is ordinary image-to-video. Magic Hour’s current browser workflow uses an uploaded starting image and optional motion prompt. Other patterns are model-specific: its image-to-video API documents a required starting image and an optional end image for supported models, while ByteDance’s Seedance 2.5 announcement documents multiple image, video and audio references. Check the actual model schema instead of assuming every image field has the same purpose.

Test one controlled image-to-video shot

Upload one clear image, describe one subject action and one camera move, then inspect the whole clip before adding more references or constraints.

Try Image-to-Video

Step 1: define what must stay unchanged

Write a short acceptance list before generating. For a product shot, it may be the label text, cap, color, proportions and background. For a person, it may be recognizable facial features, hair, clothing and accessories. The model may preserve some attributes better than others, so inspect each one explicitly.

Step 2: choose a source that contains the needed information

  • Use visible detail. A model cannot faithfully reveal the back of a product when the source shows only the front.

  • Match the intended angle. A front or three-quarter portrait is a stronger source for similar motion than for a full profile turn.

  • Leave room for the motion. A tight crop gives a pull-back or large subject movement little scene context.

  • Remove accidental ambiguity. Avoid extra people, duplicate products, heavy blur, compression artifacts and unreadable text when those details are not part of the brief.

  • Use the final aspect ratio when possible. This reduces the need for the model or editor to invent missing space later.

Step 3: prompt the change, not the whole image

The image already defines the visible subject, composition, color and lighting. Use the prompt to describe what changes next: one subject action, one camera move and any essential constraint.

The bottle remains stationary as the camera slowly pushes in. Condensation moves down the glass. Preserve the cap, label, proportions and background.

Workflow recipe output preview: 3D figure dancing

Workflow recipe

A small figurine standing on a desk and a digital character inside a computer screen both come to life at the same time. They begin performing synchronized dance moves inspired by the style of Michael Jackson. Their movements are perfectly matched—every step, gesture, and pose happens in exact unison. The dance includes smooth gliding steps, sharp turns, and rhythmic footwork, creating a strong sense of synchronization between the real-world figurine and the digital character. The contrast between physical and virtual space adds a creative and surreal effect. The environment is a simple desk setup with a glowing computer screen. The camera remains completely fixed in one position with no movement, no zoom, and no change in angle throughout the entire scene. Smooth motion, high detail, cinematic lighting, 4K quality.

Model
kling-3.0
Resolution
1080p
Duration
5 seconds
Frame rate
24 fps

Portrait example: The subject looks toward camera and gives a small natural smile. Subtle handheld movement. Preserve facial features, hair, clothing and glasses.

Landscape example: Clouds move slowly from left to right as the camera makes a restrained forward glide. Preserve the mountain silhouette and foreground composition.

Avoid stacking several actions, camera moves and scene changes into a short clip. Use the product-video prompt library for more job-specific examples.

Step 4: use an end frame only for a real destination

An end frame is useful when the final composition matters: a product reaches a target position, a camera move ends on a defined crop, or one approved still must transition into another. The model still generates the path between those states, so inspect acceleration, geometry, continuity and any moment where the source appears to melt or jump.

Step 5: add multiple references by role

When a model supports several references, give each one a job: subject appearance, object design, environment, style, motion or audio. Do not upload redundant or conflicting images and expect the model to infer priority. State the role in the prompt and confirm the host exposes the same reference behavior documented for the underlying model.

  • Subject reference: defines who or what should remain recognizable.

  • Object reference: defines a product, prop, garment or other required asset.

  • Style reference: guides palette, texture, lighting or rendering language rather than identity.

  • Motion reference: guides movement only when the chosen workflow explicitly accepts video or motion input.

Step 6: change one variable per iteration

Keep the source and acceptance list fixed. If identity fails, reduce the action or choose a closer reference angle. If composition drifts, simplify the camera instruction or use a supported end frame. If the clip looks frozen, loosen unnecessary constraints before adding more motion. Save the model, mode, prompt, references, duration, resolution and charge with each output.

Reference-image failure checklist

  • Identity drift: compare eyes, nose, mouth, hairline and accessories throughout the clip.

  • Product drift: inspect labels, letters, logos, geometry, materials and counts frame by frame.

  • Edge failures: watch hands, hair, glasses and objects that cross the subject.

  • Composition drift: compare the subject position, crop, horizon and background to the source.

  • Motion failure: reject jumps, rubbery movement, impossible contact and abrupt camera changes.

  • Audio mismatch: when audio is generated, review exact speech, timing, speaker and unwanted sound separately.

How to compare models fairly

Use the same source, prompt, duration and acceptance list. Keep every attempt and calculate cost and time per approved clip after retries. The commercial image-to-video benchmark shows how to preserve a shared input and rubric; the character consistency guide covers multi-shot identity, and the motion-control guide covers workflows where movement is the primary reference.

Frequently asked questions

Sometimes. In ordinary image-to-video, the uploaded image is the starting frame. A model may also accept separate subject, style, end-frame or motion references, each with a different role.

Only when the selected model and host support multiple references. Assign each input a clear role and remove conflicts before generating.

The motion may expose angles or details missing from the reference, or the brief may ask for too many changes at once. Use a more compatible source, reduce the motion and test a shorter clip.

It provides a concrete starting composition, which usually gives the model more visible information than text alone. It still does not guarantee identity, text or object consistency across generated frames.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Product-video shot concepts: an amber bottle, a white shoe, and a cream jar
Recommended next
AI product video prompts: 8 shots for product photos

Copy eight AI product-video prompts for bottles, shoes, food, and furniture. Choose a source photo, direct one shot, and check product fidelity before export.

Keep Characters Consistent Without Manual Editing
Best reference image-to-video tools (2026): character and product consistency
AI product-photo-to-video tools: a product bottle and animated video frame, with tools, costs and practical prompts subtitle.
7 best AI tools to turn product photos into videos
AI Video Model Benchmark
AI video model benchmark: 60 attempts across 5 models
Edit Product Photos with AI
How to edit product photos with AI: workflow and prompts