

Use text-to-video when the scene does not exist; use image-to-video when an approved image should anchor the opening frame. The source image narrows the starting composition, but it does not lock every later frame. Neither workflow is universally faster, cheaper or more reliable: model, duration, resolution, audio, queue and the number of rejected outputs determine the real result.
This guide was checked September 13, 2026 against Magic Hour’s current text-to-video product, image-to-video product and API references. It explains workflow fit; it does not report an independent retained quality benchmark.
Workflow | Start with | Choose it when | Main uncertainty | Inspect before use |
|---|---|---|---|---|
A written scene brief | The scene, subject or composition does not exist yet | The model must invent the whole frame and motion | Subject, action, camera, objects, text, dialogue and audio | |
An approved source image plus an optional motion prompt | A product, character, artwork or composition should anchor the opening | Generated motion can still change details after the starting frame | Identity, geometry, labels, edges, background, motion and audio |
Upload one approved image, describe only the motion you need, and inspect the complete result for identity, product geometry, labels and edges before scaling the workflow.
Try Image to VideoOnly a concept or script exists. Start with text-to-video to explore subjects, settings, actions and camera ideas.
A product photo, character design or approved key art exists. Start with image-to-video so the opening frame uses that visual source.
Exact labels, interfaces or legal copy must remain unchanged. Generate motion around the asset, then restore consequential text and graphics in a deterministic editor.
A sequence needs the same subject across shots. Use supported references or model controls, retain each accepted output and check continuity across the assembled sequence.
You need to compare the modes fairly. Hold the model, duration, resolution, aspect ratio and audio setting constant where the same route supports both inputs.


Text-to-video begins with a written prompt. The selected model generates the subject, setting, composition and motion instead of receiving a source frame. Magic Hour’s current web flow asks for the scene, action and camera movement; supported aspect ratios, duration, resolution and audio vary by model.
This mode fits a blank page: a new establishing shot, visual metaphor, storyboard option or creative direction. Its freedom also creates more variables to inspect. A detailed prompt can guide generation, but it cannot make a probabilistic model deterministic.
[Subject] + [one visible action] + [setting] + [shot size] + [one camera move] + [lighting or audio constraint]
Example: “A small red delivery robot crosses a rain-soaked alley at night, medium tracking shot from the side, slow dolly movement, blue storefront reflections, wheel sounds and light rain, no dialogue.” Generate candidates, then check every subject, action, object, word and sound in the complete clip.
A small red delivery robot crosses a rain-soaked alley at night, medium tracking shot from the side, slow dolly movement, blue storefront reflections, wheel sounds and light rain, no dialogue.
Image-to-video begins with a source image and, where supported, a motion prompt. Magic Hour’s current flow accepts an image and an optional description of subject movement, camera motion or style. The image supplies the opening visual context; the model generates the following frames.
Use it when the starting appearance matters: a product photo, portrait, illustration, environment or storyboard frame. The source is an anchor, not a guarantee. Motion can alter faces, hands, lettering, logos, product proportions, edges, reflections and background objects, so review the entire output.
[What should move] + [how it moves] + [camera behavior] + [what should stay visually stable] + [audio constraint]
Example: “The camera makes a slow five-degree push toward the bottle while condensation moves naturally; keep the bottle centered and avoid added objects; quiet room tone, no speech.” If exact packaging or copy matters, compare against the approved source and restore it during finishing.
The camera makes a slow five-degree push toward the bottle while condensation moves naturally; keep the bottle centered and avoid added objects; quiet room tone, no speech.
Do not infer control from the input type alone. Some models expose start and end frames, references, camera controls or native audio; others expose a smaller set. A web app, direct API and third-party host can expose different options for the same underlying model. Record the exact product, model, mode and date.
Do not use a fixed render-time range. Generation time varies with model, duration, resolution, source input, audio, queue and demand. Magic Hour provides an ETA after a render starts. Google’s Veo documentation likewise describes long-running generation jobs rather than an instant synchronous response.
Compare accepted-output cost, not subscription price or one successful request. Magic Hour’s current text-to-video API reference and image-to-video API reference return estimated charges after completion and state that failed jobs are refunded; a completed output you reject is still different from a failed job. Google’s Veo documentation similarly treats generation as a long-running operation.
Freeze three briefs. Use one scene that can begin from an approved still, one product shot and one character shot.
Create the source image deliberately. For image-to-video, use the image that would actually be approved for production. Save its dimensions and provenance.
Hold shared settings constant. Use the same model, duration, resolution, aspect ratio and audio setting when the route supports both modes.
Give both modes equal attempts. Save prompts, inputs, settings, outputs, errors and rejected candidates. Do not rescue one workflow with extra iterations.
Score complete clips. Check brief adherence, identity, object geometry, hands, text, motion, camera, cuts, dialogue, lip timing and audio.
Calculate accepted-output cost. Include credits, completed rejections, review time, source-image creation and deterministic finishing.
A production can use both without treating one as a universal first step. Explore a new visual direction with text-to-video, approve or create a key image, animate that image, then assemble only retained shots. For another brief, an existing product image may make text exploration unnecessary.
Keep the evidence chain with each deliverable: source rights, exact input, prompt, model, settings, retained output and manual corrections. This makes revisions and factual review possible even when a platform changes its available models.
It is the better starting point when an approved image should anchor the opening. Text-to-video is the better starting point when the scene must be invented from a written brief. Output quality still depends on the selected model, inputs and acceptance criteria.
No. The source constrains the opening frame, but generated motion can change identity, geometry, labels, edges or background details. Inspect the full video and restore consequential elements in a deterministic editor.
There is no universal winner. Compare the same model and settings, then include every completed rejection, source-asset step, review minute and correction in the cost of an accepted clip.
Yes. Magic Hour provides separate text-to-video and image-to-video web and API workflows. Available models and their input, duration, resolution and audio options can differ.
Use the best AI video generators guide for platform selection, the best image-to-video generators comparison for image-led options, the AI video workflow guide for assembly and finishing, or the image-to-video API guide for developer routing.
