

Vidu AI is a short-form video generator for text-to-video, image-to-video, start/end-frame animation, and reference-guided video. Its clearest differentiator is reference-to-video: you can reuse character, object, scene, style, camera, and effect references across shots. Vidu is free to try with signup and daily-login credits, but the public web pricing page does not expose a stable plan table without loading the app, so confirm the live credit allowance and renewal price before paying.
Vidu is most relevant when you need recurring characters or products across several short shots, or want Vidu Q3's documented native audio and clips up to 16 seconds. It is less suitable when the job depends on a full editing timeline, long-form assembly, or a predictable budget based only on a subscription headline. This review separates Vidu's web app from its API and shows how to compare alternatives with the same input.
We checked Vidu's product pages, web pricing FAQ, API documentation, API pricing, and current Q3 documentation on September 12, 2026. We did not run a controlled visual-quality or speed benchmark. Capability claims below are attributed to current first-party documentation; use the included test protocol to evaluate output for your own brief.
Vidu is a generative image and video platform from ShengShu Technology. The consumer product runs in a browser and mobile app. The separate Vidu API platform exposes model endpoints for developers.
The product name can refer to three different things:
Keep those layers separate. A feature shown on the consumer site may use different settings, credits, or availability from an API endpoint. A model price in the API documentation is not the cost of a web subscription.
This table compares product structure and workflow fit. It is not a visual-quality ranking.
Option | Best fit | Reference workflow | Native audio | Developer route | Main tradeoff |
|---|---|---|---|---|---|
Recurring characters, products, and short multi-shot concepts | Dedicated reference-to-video with multiple reusable references | Q3 documents dialogue, voiceover, sound effects, and music | Separate Vidu API with per-credit pricing | Web and API allowances differ; test continuity across several shots | |
Trying multiple image and video workflows in one browser-based suite | Image, video, and model-specific inputs vary by tool | Available on supported video workflows and models | Magic Hour API | Model and feature availability vary by workflow; verify the selected route | |
Cinematic prompt and image-driven generation | Character, element, and model-specific reference controls | Available on supported Kling models | Separate Kling API | Credits, model versions, and app/API prices require separate checks | |
Generation combined with editing and production controls | Reference and input controls vary by model and tool | Available in supported generative workflows | Separate Runway API | Deeper workflow can add complexity and app/API billing is separate | |
Fast social clips, effects, templates, and mobile creation | Image and character controls vary by current model | Available in supported features and models | PixVerse API | Template effects and general generation solve different jobs; check the exact mode |
Try a short representative brief, inspect the full result, and compare the cost of outputs you would actually publish.
Try AI Video GeneratorVidu is free to try, but “free” means a limited credit allowance rather than unlimited production use. The current Vidu pricing FAQ says new accounts receive signup bonus credits and can earn credits through daily login, events, competitions, or the Artist Program.
Vidu describes three credit types:
The plan cards load dynamically and were not reliably exposed in the public page we checked. That means a fixed free-credit count or monthly plan price copied from another review may already be wrong. Open the live pricing screen in your own account and record the plan price, renewal cadence, included credits, expiration, watermark or download rules, concurrency, and commercial terms before subscribing. Vidu's FAQ also says Stripe subscriptions renew automatically and that purchases are nonrefundable.
The Vidu API pricing page currently lists one credit at $0.005. It prices generation by model, mode, duration, and resolution. Current Q3 Pro examples include:
The same table lists lower off-peak rates for supported routes. Earlier Q2 modes use starting charges plus per-second credits, so do not reuse a Q3 calculation for a Q2, reference, start/end, or multi-frame request.
For example, ten eight-second Q3 Pro generations at 720p contain 80 generated seconds. At the listed $0.10 per second, the request bill is $8.00 before tax and any separate image, audio, editing, storage, retry, or delivery cost. If only two clips are acceptable, the generation cost per accepted clip is $4.00, not $0.80.
Use this budget formula:
cost per accepted clip = total generation charges ÷ clips that pass the delivery checklist
Record failed requests and rejected outputs. Counting only the final export hides the cost of iteration.
Vidu Q3 is the current flagship described on Vidu's site. Its product page says one generation can produce up to 16 seconds at a time and can create visuals with dialogue, voiceover, sound effects, and music together. The page lists English, Japanese, and Chinese output, multi-speaker conversations, and camera and pacing controls.
Those are provider claims, not independent measurements. Before choosing Q3 for production, test whether the requested words are spoken correctly, the intended speaker owns each line, sound effects arrive at the right moment, music does not bury speech, and the final container works in your editing stack.
Longer clips can reduce stitching, but duration alone does not guarantee narrative continuity. A 16-second result can still drift in wardrobe, product shape, background, lighting, or shot logic. Evaluate every required beat rather than treating maximum duration as usable duration.
Vidu Reference to Video is for reusable subjects and visual attributes across a prompted scene. Vidu says the web workflow accepts one to seven reference images and can derive character, object, style, composition, camera, scene, and effect references.
Use reference-to-video when the identity of a character or product must survive across several shots. Use image-to-video when you have one exact starting frame and mainly need to add motion. Use start/end-to-video when the opening and closing compositions matter. Use text-to-video when you are exploring ideas and have no required source visual.
These modes are not interchangeable. A single image can anchor a first frame while leaving later identity underconstrained. Multiple references provide more information, but conflicting angles, lighting, outfits, or product versions can confuse the model.
List the identity features that matter: face shape, hair, wardrobe, product silhouette, label placement, material, color, scene geometry, or art direction. If everything is “important,” the review becomes subjective.
Choose sharp images with a readable subject and minimal occlusion. For a character, include useful angles without switching wardrobe or age. For a product, keep the exact SKU, packaging, cap, label, and proportions consistent. Avoid mixing concept art, photography, and screenshots unless the style change is intentional.
Name each subject in the prompt and say what the reference controls. Describe the action, setting, camera, timing, and elements that must remain fixed. Do not make the model infer whether one image defines a person, outfit, location, or color grade.
Run the shortest useful shot before spending credits on a full sequence. Confirm identity and composition, then change one variable at a time. A new prompt, reference set, model, duration, and resolution in the same retry cannot tell you what fixed or broke the result.
Pause on the first, middle, and final frames for identity and object drift. Watch at normal speed for motion and camera problems. Listen once without watching for speech, timing, noise, and music balance.
Use one eight-second brief that resembles the work you plan to publish. Keep the same source assets, prompt, aspect ratio, duration, and resolution when comparing Vidu with another tool.
Reference set: one front three-quarter character or product image, one alternate angle, and one environment or style reference.
Prompt structure: “[Subject from reference] performs [one action] in [setting]. Preserve [three identity details]. Camera: [one movement]. Lighting: [specific condition]. End on [clear final composition]. Include [dialogue or sound cue] only if native audio is being evaluated.”
Generate three attempts in the same mode. Record the model, settings, queue time, credits charged, failures, and whether each clip passes:
Compare cost per accepted clip, not the prettiest single frame.
Magic Hour's AI Video Generator is the practical alternative when you want video generation alongside image creation, talking photos, face swap, lip sync, animation, and other media workflows in one account. Start here when the decision is broader than one Vidu model or when the next production step matters as much as the initial clip.
Kling AI is relevant when you want to compare current Kling models, motion, camera behavior, native audio availability, and reference controls against Vidu. Match model version and settings; “Kling vs Vidu” is too broad if one side uses a fast model and the other uses a flagship route.
Runway is the better comparison when the workflow needs generation, transformation, timeline work, or production controls in one environment. Separate its web subscription from API billing and calculate the cost of the actual model and edit path.
PixVerse is relevant for effect-led clips, templates, mobile creation, and social formats. Compare its exact generation mode with Vidu reference-to-video rather than treating a template effect as the same product.
Vidu offers signup bonus credits and ways to earn bonus credits, including daily login and events. The amount and what those credits can generate may change. Check the live account for the current allowance, watermark, download, and commercial-use terms.
Vidu is best matched to short text-, image-, start/end-, and reference-guided clips. Its most distinctive documented workflow is reference-to-video for recurring characters, objects, scenes, and styles.
Vidu AI is the overall product and developer platform. Vidu Q3 is a model family used inside supported video workflows. A web plan, model, and API endpoint can have different settings and prices.
Vidu's current Q3 page says it can generate dialogue, voiceover, sound effects, and music with video. Test pronunciation, speaker assignment, timing, and mix quality for your own brief; the product-page claim is not an independent quality result.
Vidu currently documents up to 16 seconds in one Q3 generation. Available duration depends on the selected model, route, resolution, and current product settings.
Yes. Vidu provides separate API endpoints and pay-as-you-go credit pricing for text-to-video, image-to-video, start/end-to-video, reference-to-video, and other supported routes. API access and consumer subscriptions are separate.
Commercial use depends on the current plan, terms, model, and rights to the inputs. Check Vidu's live terms before publication and retain proof that you can use each likeness, voice, image, logo, product, and audio asset.
