Hailuo 02 guide: current status vs MiniMax H3


Quick answer
Hailuo 02 is a 2025 MiniMax video model, not MiniMax’s current flagship. It introduced text-to-video, image-to-video, native 1080p at six seconds, and later start/end-frame controls. For a new workflow in September 2026, begin by evaluating MiniMax H3, the newer multimodal model with native stereo audio and text, image, video and audio context. Keep Hailuo 02 only when a provider still exposes a control or price that your existing workflow needs.
This guide uses MiniMax’s official release material and Magic Hour’s current model implementation, checked September 13, 2026. Provider claims are identified as such. We did not run a retained Hailuo 02 versus H3 quality benchmark for this update.
Hailuo 02 and MiniMax H3 compared
Model | Current role | Documented inputs and output | Use it when |
|---|---|---|---|
2025 MiniMax release; still useful where a platform exposes it | Text-to-video and image-to-video; 6–10 seconds depending on resolution | An existing workflow depends on its start, end or physics-oriented controls | |
Current general-purpose multimodal MiniMax video model | Text, image, video and audio context; native stereo video generation | A new workflow needs multimodal references, audio or current model support |
Try the current MiniMax model
Use one prompt and one approved reference, generate a short clip, and inspect identity, product geometry, motion and audio before increasing duration or resolution.
Open MiniMax H3What Hailuo 02 actually supports
MiniMax’s Hailuo 02 launch documents text-to-video and image-to-video, with 768p clips at six or ten seconds and 1080p clips at six seconds. MiniMax later added start-and-end-frame and end-frame-only controls in its web and mobile products. Those are useful historical capabilities, but they do not make Hailuo 02 the default choice for a new 2026 integration.
Do not treat Hailuo 02 as a talking-avatar or automatic dialogue-editing product. The previous version of this article claimed custom speech, multilingual dialogue, multi-character scene logic and measured lip-sync superiority without retained tests or first-party support. Those claims have been removed.
What changed with MiniMax H3
MiniMax’s H3 release describes H3 as a general-purpose multimodal generation model that uses text, images, video and audio as context and generates video with native stereo sound. MiniMax documents output up to 15 seconds at 2K in its direct model announcement.
A platform can expose a different supported subset or add workflow-level controls. Magic Hour’s current MiniMax H3 page supports text-to-video, image-to-video and saved references with its own listed resolutions, durations and credit rates. Verify the settings in the interface or API you will actually use rather than copying the provider maximum into every platform comparison.
When should you still use Hailuo 02?
Existing reproducibility: an approved campaign or evaluation depends on Hailuo 02 and changing the model would invalidate the comparison.
Start/end-frame workflow: the specific provider exposes the frame controls you need and H3 access there does not.
Known cost envelope: retained jobs show Hailuo 02 meets the acceptance threshold at a lower full cost for your exact shot.
Before keeping it, confirm the endpoint is still supported, pin the exact model ID when possible, and archive prompts, input assets, settings and accepted outputs. A provider label such as “Hailuo” is not specific enough for reproducible work.
A three-shot evaluation for H3
Reference consistency: move one approved character from bright exterior light into a dim interior; inspect face, clothing, proportions and color.
Product handling: rotate a labeled product; inspect hands, logo, label text, shape and continuity.
Audio timing: prompt one visible impact and reaction; inspect whether sound, action and response occur in the intended order.
Run the same shot more than once because generative output varies. Record model ID, platform, resolution, duration, references, prompt, credits, processing time, rejection reason and correction time. Choose by cost per accepted clip, not the cheapest listed generation.
How to use MiniMax H3 in Magic Hour
Open MiniMax H3 in Magic Hour, choose text-to-video or image-to-video, add only the references needed for the shot, and write one observable action plus camera behavior. Generate a short draft first. Review the entire clip before increasing duration or resolution.
Developers can use the Magic Hour MiniMax H3 API page for current SDK, endpoint and billing guidance. Store the Magic Hour project ID and selected model with the request so retries and later model changes remain traceable.
Frequently asked questions
No. MiniMax released H3 in July 2026 as a newer general-purpose multimodal video model. Hailuo 02 remains a historical model that may still be available through some providers.
MiniMax’s Hailuo 02 launch documents text-to-video and image-to-video, not a custom-audio lip-sync workflow. Do not infer dialogue controls from promotional example videos. Use a documented audio-capable model or a separate lip-sync tool when exact supplied speech is required.
Yes. MiniMax’s official H3 announcement describes native stereo sound. Check the platform’s current H3 implementation for the exact input modes, output settings and billing.
Start with H3 for a new evaluation. Keep Hailuo 02 only when a specific retained workflow, provider control or measured cost advantage requires it.
Use the AI video model release tracker to verify release dates and current Magic Hour access before relying on an older model guide.







