

To use references in Seedance 2.0, upload the actual image, video or audio files, identify them using your provider’s labels, and tell the model what to take from each one. For example: “Use Image 1 for the product’s appearance and Video 1 for the camera movement.” A filename typed into a prompt does not upload the file, and a phrase such as “identity lock” does not guarantee consistency.
Choose image-to-video when an image must be the opening frame. Choose reference-to-video when an asset should guide a new scene. These are different jobs, even when they use the same picture. This guide explains the distinction, the current input limits and six example briefs you can adapt.
The API details below refer to BytePlus ModelArk’s Seedance 2.0 documentation, checked September 10, 2026. Host interfaces may use different labels and restrictions. The example prompts are proposed starting points, not newly generated benchmarks.
Start in Magic Hour AI Video Generator, choose the workflow that fits your source, and review a short draft before producing the final export.
Try AI Video GeneratorWhat you want to preserve or borrow | Input to choose | Example instruction |
|---|---|---|
The exact intended opening composition | First-frame image-to-video | Animate this opening image with a slow camera push |
An opening and ending composition | First-and-last-frame image-to-video | Move from the first supplied frame to the last |
A product or character’s appearance in a new scene | Reference image | Use Image 1 for the subject’s appearance |
The look of a setting | Reference image | Use Image 2 for the room’s layout and lighting |
Camera motion | Reference video | Borrow the camera path from Video 1, not its subject |
A subject’s movement | Reference video | Use Video 1 for the walking action; keep the camera still |
Music, voice or sound direction | Reference audio with an image or video | Use Audio 1 for the soundtrack or specified voice qualities |
A change to existing footage | A supported video-editing workflow | Edit Video 1 and specify what should change |
References condition a generated result. They do not turn it into a pixel-preserving editor or guarantee identical faces, products, timing or choreography. Check the output against the source asset before using it.
Start with the deliverable. For a product reveal, the actual product image usually matters more than a mood board. For a motion study, the useful part of a reference video may be a brief turn or camera move rather than the entire clip.
Give each asset a written job before uploading it:
Use clean assets that show the needed information. Crop away unrelated subjects when that helps identify the intended object, and trim video or audio to the useful segment. Keep your originals so you can compare the result later.
Do not fill every available slot by default. If two images show different versions of the same product, explain which one is authoritative or remove the conflicting image. Multiple references should add information, not require the model to guess which design you sell.
The ModelArk video generation documentation lists these limits for the 2.0 series:
Reference type | Maximum count | Duration | File details |
|---|---|---|---|
Images in omni-reference mode | 9 | Not applicable | Each image under 30 MB; supported formats include JPEG, PNG and WebP |
Videos | 3 | Each 2–15 seconds; no more than 15 seconds combined | MP4 or MOV; each no more than 200 MB; 24–60 fps |
Audio | 3 | Each 2–15 seconds; no more than 15 seconds combined | MP3 or WAV; each no more than 15 MB |
A host can impose additional limits. Audio needs a visual reference: ModelArk does not support audio-only or text-plus-audio input without an image or video. First-frame mode uses one image; first-and-last-frame mode uses two. Those are separate from the nine-image omni-reference allowance.
For real people, check portrait access before building the brief. ModelArk requires its supported portrait-asset routes; an ordinary upload containing a real human face may be rejected. Use a permitted asset and obtain the necessary permission for the person, voice and material involved.
In a browser interface, upload each asset and use the reference labels or selectable mentions that the application creates. Some hosts expose labels such as @Image1. Do not assume another provider uses the same spacing, tag syntax or upload controls.
For ModelArk’s API, the Seedance 2.0 tutorial uses typed content items with reference roles. The prompt identifies assets by media type and position, such as Image 1, Video 1 and Audio 1. Numbering is within each media type, starting at one, rather than one combined sequence for every file.
ModelArk input purpose | Role |
|---|---|
Image used as a reference | reference_image |
Video used as a reference | reference_video |
Audio used as a reference | reference_audio |
Required starting image | first_frame |
Required ending image | last_frame |
These role names belong to the API request structure. They are not bracketed prompt commands. Writing “[reference_image: backpack.png]” does not attach backpack.png. Writing “[identity_lock]”, “[camera_copy]” or “[beat_sync]” does not enable a documented ModelArk switch with that name.
Check the asset order before submitting. If you reorder or replace uploads, make sure “Image 1” still identifies the product rather than the lighting reference.
A useful reference prompt answers four questions: Which asset supplies the subject? What happens? Which reference supplies the movement or look? What must remain unchanged?
Vague: “Use these references to make a cool backpack ad.”
More specific: “The blue backpack from Image 1 stands on a neutral studio floor. Keep its two front pockets and black shoulder straps. Use Image 2 for the soft side lighting only. Follow the slow camera move from Video 1; do not copy the object or background in that video.”
The second prompt makes the source of each decision explicit. It is still a request to a generative model, not proof that every feature will survive. Inspect pocket count, straps, logos and any detail a buyer would rely on.
If the opening composition must match an existing image, switch to first-frame mode. Use reference mode when you want a newly composed scene featuring the same product or character. For resolution, duration and audio choices, see the separate Seedance 2.0 settings guide.
These examples use ModelArk-style labels. Substitute the actual labels in your host. Each prompt assumes the named assets have already been uploaded and accepted.
Assets: Image 1 shows the product. Video 1 shows the desired camera move.
Prompt: “Show the orange desk lamp from Image 1 on a plain gray tabletop. Preserve the curved neck, circular base and single shade. Use Video 1 for the slow left-to-right camera arc only. The lamp stays still and fully visible. Soft studio light; no added writing.”
Check: The lamp’s construction, the direction of the camera move and whether the model borrowed unwanted objects from the video. An orbit can expose a side not shown in your photo; provide another consistent view if that detail matters.
Assets: Image 1 shows an illustrated character; Image 2 shows a location.
Prompt: “The small green robot from Image 1 walks through the workshop shown in Image 2. Keep the robot’s square head, yellow chest panel and two short antennae. Its steps are slow. The camera follows from the side in a single shot. Preserve the illustrated style.”
Check: Antennae, limb count, costume details and visual style throughout the clip. Reusing a reference across shots can guide continuity, but it does not guarantee an identical character every time.
Assets: Image 1 shows the product; Image 2 provides lighting direction.
Prompt: “The ceramic bowl from Image 1 sits on a bare wooden table. Borrow only the warm light from the left and soft shadow from Image 2. Do not include the person or furniture in Image 2. The camera slowly moves closer to the bowl.”
Check: That the product is still the subject and the new scene has not inherited unrelated elements. If the style reference keeps taking over, replace it with a cleaner reference or describe the lighting in text.
Assets: Image 1 shows a permitted character; Video 1 shows a simple movement.
Prompt: “The clay character from Image 1 performs the slow wave shown in Video 1. Borrow the arm movement only. Keep the camera fixed in a medium shot and keep the plain blue background from Image 1.”
Check: Which movement was transferred, whether the camera stayed still and whether hands or limbs deform. For a dedicated image-plus-motion-video task in Magic Hour, compare the available motion-control workflow rather than assuming every image-to-video model accepts motion references.
Assets: Image 1 shows the scene; Audio 1 supplies the intended sound direction.
Prompt: “Animate the paper boats in Image 1 drifting slowly along the stream. Use Audio 1 for the gentle piano soundtrack. Keep the camera above the stream, following the boats. No dialogue.”
Check: That audio is enabled, the soundtrack matches the intended direction and the video ends cleanly. If an exact recording and frame-accurate beat placement are required, retain the approved audio and align the finished video in an editor.
Assets: A prepared opening image and a prepared ending image attached to their dedicated first- and last-frame inputs.
Prompt: “The paper lantern opens gradually from its folded state into the expanded lantern. The camera stays still. Keep the same table, warm lighting and central position throughout the transition.”
Check: Both endpoint compositions and every intermediate frame. A plausible first and last image do not ensure a physically correct transition. Do not replace the actual frame inputs with a made-up “[first_frame_lock]” phrase.
Symptom | Useful next check |
|---|---|
The wrong object becomes the subject | Name the object and its distinguishing features; remove unrelated content from the reference |
A character changes appearance | Check for contradictory reference images and simplify the requested action |
The camera moves instead of the subject | Specify which movement to borrow and whether the camera should remain fixed |
The opening frame differs | Confirm first-frame mode rather than general reference mode; check aspect-ratio cropping |
Upload fails before generation | Check file type, dimensions, duration, size and portrait permissions |
Audio is missing or unrelated | Confirm audio output is enabled and the accepted audio reference is named correctly |
A motion or beat is only approximate | Decide whether a generative approximation is sufficient or an editing workflow is required |
Save the prompt, accepted asset labels, model variant and settings beside the downloaded result. Change one part of the brief at a time when diagnosing a problem so you can compare the effect.
Use Seedance references to guide a new generated clip. If you already have an approved still and only need movement, Magic Hour image-to-video is another starting point. If you already have the video and a final spoken recording, a dedicated lip-sync tool addresses that narrower job.
Compare the full deliverable, including rejected candidates, finishing work and any required commercial plan. The Seedance pricing guide explains why a single attractive generation price is not the same as the cost of usable footage.
For the complete generation workflow, follow how to use Seedance 2.0.
