

For a single Seedance 2.0 shot, start with 5–8 seconds, the final aspect ratio and only the references the shot needs. Use 720p to review the idea if that meets your budget; choose a supported higher resolution when delivery requires it. Enable audio from the first attempt when dialogue or sound timing matters. These are starting points, not settings that guarantee a good result.
The model and provider matter. On BytePlus ModelArk, standard Seedance 2.0 supports resolutions through 4K, while 2.0 Fast and Mini stop at 720p. A host may expose fewer options. This guide covers 2.0, not the newer Seedance 2.5. Model documentation was checked September 10, 2026; the suggested prompts below are worked examples, not newly generated test results.
Use one representative prompt or source image and compare the result before scaling the workflow.
Try AI Video GeneratorSetting | Practical starting point | Change it when |
|---|---|---|
Mode | Text-to-video for a new scene; image-to-video for a required opening image | Choose reference-to-video to borrow an asset’s appearance, movement or sound |
Resolution | 720p for reviewing a concept | The final placement requires 1080p or 4K and your selected model supports it |
Duration | 5–8 seconds for one action | Dialogue or multiple shots need more time; 2.0 supports 4–15 seconds on ModelArk |
Aspect ratio | Match the delivery: 9:16 vertical, 16:9 landscape or 1:1 square | Prepare a separate composition for a different placement |
Audio | On for dialogue, audible actions or music-driven timing | Turn it off for silent footage that will receive its soundtrack in editing |
References | One clear asset per required role | Add another reference only to supply missing information |
Final approval | Review the actual final file at its intended size | Regeneration, upscaling, cropping or editing changes the deliverable |
The current ModelArk Seedance 2.0 tutorial lists these variants:
Model | ModelArk ID | Output resolutions | Duration |
|---|---|---|---|
dreamina-seedance-2-0-260128 | 480p, 720p, 1080p, 4K | 4–15 seconds | |
dreamina-seedance-2-0-fast-260128 | 480p, 720p | 4–15 seconds | |
dreamina-seedance-2-0-mini-260615 | 480p, 720p | 4–15 seconds |
These are ModelArk’s documented options, not a promise that every Seedance app exposes the same menu. Confirm the model name and generation mode before paying. A missing 1080p option on Fast is not a prompt problem.
Use a lower resolution when the decision is broad motion or composition and the provider’s quote makes that worthwhile. Use the intended delivery resolution earlier when the decision depends on fine product detail, faces or small objects. There is little value in approving a low-resolution preview that cannot show the defect you need to catch.
Seedance 2.0 does not have ModelArk’s dedicated draft mode. A new higher-resolution generation is a new candidate; it should not be treated as the approved clip with extra pixels. If you upscale an existing result instead, inspect the processed file for altered detail. Neither workflow guarantees that a label, logo or face stays correct.
Choose the frame for the actual placement. A vertical product ad and a landscape website hero need different space around the subject.
For image-to-video, ModelArk documents center cropping when the requested ratio differs from the source image. A wide product photograph can lose the packaging edges when forced into 9:16. Prepare the reference in the intended ratio first, or use the provider’s adaptive ratio option when preserving the original framing is the priority. See the video generation and cropping rules.
Keep critical text and the product away from edges that may be covered by the destination app’s interface. Add exact prices, subtitles and calls to action in your editor so they remain editable and readable.
A 5-second reveal, an 8-second spoken line and a 15-second sequence solve different jobs. Write the action first, then allow enough time for it to happen at the intended pace.
For example, a 15-second ad could use three separately generated 5-second shots. At two candidates per shot, you would generate 30 seconds to obtain 15 seconds of selected footage, before any additional revisions. That is a planning example, not an observed acceptance rate. Compare the complete job quote in our Seedance 2.0 pricing guide, including reference inputs and the provider’s billing rules.
Seedance 2.0 generates video and audio together, including dialogue, effects and ambient sound. ByteDance describes this in the Seedance 2.0 release.
For a speaking character, test with the intended spoken line from the start. A silent preview cannot establish pronunciation, lip synchronization or whether the sentence fits. Describe the speaker, the exact words and the desired background sound separately.
For a silent product montage with a separately licensed soundtrack, switching generated audio off can simplify the brief. Do not assume this makes the job cheaper or faster: check the selected provider’s current quote. Audio pricing and controls vary by implementation.
When the spoken recording is already approved and the task is to synchronize an existing video, a dedicated lip-sync workflow may fit better than regenerating the scene. Match the tool to the part of the asset you actually need to change.
There is no documented universal “character → face → style → scene” priority stack. Explain what each supplied asset should contribute, and use the correct input mode.
What must carry into the video? | Appropriate input | What to specify |
|---|---|---|
The opening composition | Image-to-video first frame | The movement that should follow this image |
Both opening and closing images | First-and-last-frame mode | The transition between the two frames |
A product’s appearance | Reference image | Which object to preserve and what may change |
A visual style | Reference image | Lighting or color treatment, without replacing the main subject |
Camera or subject motion | Reference video | Whether to borrow the camera move, the action or both |
A sound or vocal performance | Reference audio plus image or video | Which sound to use and how it relates to the scene |
In ModelArk’s omni-reference prompts, assets are identified by type and number, such as Image 1, Video 1 and Audio 1. Other hosts may use selectable tags. Follow your interface’s syntax and identify each asset explicitly. An asset ID or an unexplained list of uploads is not the same as giving it a clear role.
For strict first-frame matching, use the actual first-frame input. Saying “start with Image 1” in a general reference prompt does not substitute for selecting that mode. Our Seedance reference guide covers image, video and audio tagging in more detail.
ModelArk’s current 2.0 documentation permits up to 9 reference images, 3 reference videos and 3 reference audio clips. Reference videos must fit within 15 seconds combined; audio references have their own 15-second combined limit. Audio-only and text-plus-audio inputs are not supported without an image or video. Check your host for additional combined-file or upload limits.
Real-person portrait inputs also have access requirements. ModelArk directs users to its approved portrait-asset workflows; ordinary uploads containing real faces are not a substitute. A rejected portrait upload may be an access or asset-validation issue, not something to fix by adding more prompt detail.
The following examples are proposed starting points. They demonstrate how to connect settings to a deliverable; they do not represent tested performance or guaranteed outputs.
Settings: Image-to-video, 9:16, 5 seconds, 720p for concept review, audio off if adding a soundtrack later. Use a product image already composed vertically as the first frame.
Prompt: “The matte blue travel mug remains upright on the pale stone counter. The camera moves slowly closer as soft window light passes across its curved surface. Keep the handle, lid and proportions unchanged. One continuous shot; the mug stays fully inside the frame.”
Accept only if: The mug retains its shape, the lid stays attached and the final framing leaves room for the ad copy. If the label is a purchase-critical detail, examine it at delivery size. Add the offer and CTA in editing.
To prepare this kind of asset in Magic Hour, start with image-to-video and select from the models and settings available there. The ModelArk-specific limits above do not describe every Magic Hour model.
Settings: An approved character-reference workflow, 16:9, 8 seconds, audio on. Use a character asset accepted by the selected provider.
Prompt: “Medium shot of the illustrated shopkeeper from Image 1 behind a flower stall. The shopkeeper smiles and says, ‘Your order is ready. I added the yellow tulips.’ The camera remains still. Quiet street ambience underneath; no music.”
Accept only if: The sentence is complete, the words are correct, mouth movement follows the speech and the background does not overpower the voice. If the delivery is rushed, shorten the script or allow more time instead of simply raising resolution.
Settings: Text-to-video, 16:9, 10 seconds, audio on, a supported resolution appropriate for review.
Prompt: “Shot 1: A locked wide view of a quiet greenhouse at dawn. Water droplets fall from the leaves onto the tiled floor. Shot 2: Close view of one fern as a drop slides from its tip. Soft dripping and distant birds continue across the cut. No dialogue or music.”
Accept only if: Both requested shots appear, the second image belongs in the same greenhouse and the sound continues naturally. Use your editor for an exact cut time; a timestamp in a prompt is not a frame-accurate edit decision.
Problem | Check first | Next change |
|---|---|---|
Product or face changes | Whether the correct asset and input mode are selected | Clarify which features to preserve; remove references with conflicting identities or designs |
Subject gets cropped | Source image versus output ratio | Recompose the reference before generating |
Action is missing or rushed | Number of events versus available seconds | Remove an event, lengthen the clip or generate separate shots |
Dialogue is wrong or cut off | Exact script, speaker assignment and duration | Simplify the line and retest with audio enabled |
1080p or 4K is unavailable | Model variant and host settings | Select a supporting variant or plan a separate finishing step |
A 4K file completes but will not play | Decoder support | ModelArk’s standard 4K output uses 10-bit H.265/HEVC; try a compatible player or transcode the file |
Higher resolution still looks wrong | Whether the defect is semantic rather than pixel detail | Correct the brief or reference; more pixels do not fix the wrong object |
Change one decision at a time when comparing candidates, and keep the prompt, references and settings with the accepted file. That makes it easier to identify why a result changed and to reproduce the workflow, without assuming deterministic regeneration.
Watch the full final export with sound. Check identity, product details, unexpected objects, spoken words, cropping and any added text. Review it in the actual placement as well as in a desktop player.
Use the provider’s applicable commercial terms and assets you have permission to use. Download accepted files promptly: ModelArk’s generated video URLs expire after 24 hours. Save the final video and the inputs that produced it, not just the preview link.
