How to use reference images for image-to-video


Quick answer
Use a reference image as the starting frame when you want to animate an existing composition. Add an end frame, subject image, style image or motion reference only when the selected model explicitly supports that role. A reference guides generation; it does not guarantee that identity, text, logos, products or unseen angles will remain exact.
Choose the reference pattern before the model
Reference pattern | Use it for | What it controls | Main risk |
|---|---|---|---|
Starting image | Animating one still | Initial subject, scene and composition | Unseen angles and motion may be invented |
Start and end images | A planned transition | Initial and final visual states | The path between frames can distort |
Multiple references | Subject, object, style or motion guidance | Only the roles exposed by the selected model | Conflicting references can reduce consistency |
Motion reference | Reusing movement or camera behavior | Performance or motion when explicitly supported | Appearance may drift without a separate subject reference |
The first pattern is ordinary image-to-video. Magic Hour’s current browser workflow uses an uploaded starting image and optional motion prompt. Other patterns are model-specific: its image-to-video API documents a required starting image and an optional end image for supported models, while ByteDance’s Seedance 2.5 announcement documents multiple image, video and audio references. Check the actual model schema instead of assuming every image field has the same purpose.
Test one controlled image-to-video shot
Upload one clear image, describe one subject action and one camera move, then inspect the whole clip before adding more references or constraints.
Try Image-to-VideoStep 1: define what must stay unchanged
Write a short acceptance list before generating. For a product shot, it may be the label text, cap, color, proportions and background. For a person, it may be recognizable facial features, hair, clothing and accessories. The model may preserve some attributes better than others, so inspect each one explicitly.
Step 2: choose a source that contains the needed information
Use visible detail. A model cannot faithfully reveal the back of a product when the source shows only the front.
Match the intended angle. A front or three-quarter portrait is a stronger source for similar motion than for a full profile turn.
Leave room for the motion. A tight crop gives a pull-back or large subject movement little scene context.
Remove accidental ambiguity. Avoid extra people, duplicate products, heavy blur, compression artifacts and unreadable text when those details are not part of the brief.
Use the final aspect ratio when possible. This reduces the need for the model or editor to invent missing space later.
Step 3: prompt the change, not the whole image
The image already defines the visible subject, composition, color and lighting. Use the prompt to describe what changes next: one subject action, one camera move and any essential constraint.
The bottle remains stationary as the camera slowly pushes in. Condensation moves down the glass. Preserve the cap, label, proportions and background.

Workflow recipe
A small figurine standing on a desk and a digital character inside a computer screen both come to life at the same time. They begin performing synchronized dance moves inspired by the style of Michael Jackson. Their movements are perfectly matched—every step, gesture, and pose happens in exact unison. The dance includes smooth gliding steps, sharp turns, and rhythmic footwork, creating a strong sense of synchronization between the real-world figurine and the digital character. The contrast between physical and virtual space adds a creative and surreal effect. The environment is a simple desk setup with a glowing computer screen. The camera remains completely fixed in one position with no movement, no zoom, and no change in angle throughout the entire scene. Smooth motion, high detail, cinematic lighting, 4K quality.
- Model
- kling-3.0
- Resolution
- 1080p
- Duration
- 5 seconds
- Frame rate
- 24 fps
Portrait example: The subject looks toward camera and gives a small natural smile. Subtle handheld movement. Preserve facial features, hair, clothing and glasses.
Landscape example: Clouds move slowly from left to right as the camera makes a restrained forward glide. Preserve the mountain silhouette and foreground composition.
Avoid stacking several actions, camera moves and scene changes into a short clip. Use the product-video prompt library for more job-specific examples.
Step 4: use an end frame only for a real destination
An end frame is useful when the final composition matters: a product reaches a target position, a camera move ends on a defined crop, or one approved still must transition into another. The model still generates the path between those states, so inspect acceleration, geometry, continuity and any moment where the source appears to melt or jump.
Step 5: add multiple references by role
When a model supports several references, give each one a job: subject appearance, object design, environment, style, motion or audio. Do not upload redundant or conflicting images and expect the model to infer priority. State the role in the prompt and confirm the host exposes the same reference behavior documented for the underlying model.
Subject reference: defines who or what should remain recognizable.
Object reference: defines a product, prop, garment or other required asset.
Style reference: guides palette, texture, lighting or rendering language rather than identity.
Motion reference: guides movement only when the chosen workflow explicitly accepts video or motion input.
Step 6: change one variable per iteration
Keep the source and acceptance list fixed. If identity fails, reduce the action or choose a closer reference angle. If composition drifts, simplify the camera instruction or use a supported end frame. If the clip looks frozen, loosen unnecessary constraints before adding more motion. Save the model, mode, prompt, references, duration, resolution and charge with each output.
Reference-image failure checklist
Identity drift: compare eyes, nose, mouth, hairline and accessories throughout the clip.
Product drift: inspect labels, letters, logos, geometry, materials and counts frame by frame.
Edge failures: watch hands, hair, glasses and objects that cross the subject.
Composition drift: compare the subject position, crop, horizon and background to the source.
Motion failure: reject jumps, rubbery movement, impossible contact and abrupt camera changes.
Audio mismatch: when audio is generated, review exact speech, timing, speaker and unwanted sound separately.
How to compare models fairly
Use the same source, prompt, duration and acceptance list. Keep every attempt and calculate cost and time per approved clip after retries. The commercial image-to-video benchmark shows how to preserve a shared input and rubric; the character consistency guide covers multi-shot identity, and the motion-control guide covers workflows where movement is the primary reference.
Frequently asked questions
Sometimes. In ordinary image-to-video, the uploaded image is the starting frame. A model may also accept separate subject, style, end-frame or motion references, each with a different role.
Only when the selected model and host support multiple references. Assign each input a clear role and remove conflicts before generating.
The motion may expose angles or details missing from the reference, or the brief may ask for too many changes at once. Use a more compatible source, reduce the motion and test a shorter clip.
It provides a concrete starting composition, which usually gives the model more visible information than text alone. It still does not guarantee identity, text or object consistency across generated frames.





