

Veo 3’s defining change from Veo 2 was native audio generation. Veo 3 also moved the stable Vertex model to 720p or 1080p and 4, 6 or 8-second outputs, while the stable Veo 2 model is documented at 720p and 5–8 seconds. For a new integration, evaluate Veo 3.1: Google now marks Veo 3 API models as deprecated in the Gemini API.
This is a generation comparison, not a claim that every historical preview feature appeared in every interface. Google exposed different capabilities across Vertex AI, Gemini API, Flow and preview model IDs; always identify the exact endpoint and date.
Model generation | Native audio | Documented inputs | Current API status and output |
|---|---|---|---|
Text and image; some Veo 2 variants added references, extension or frame controls | Stable Vertex model; 720p, 24 fps, 5–8 seconds | ||
Stable model: text; preview routes also supported image input | Deprecated in Gemini API; stable Vertex model documents 720p/1080p and 4, 6 or 8 seconds | ||
Text, image and supported video extension; references and first/last frames on applicable variants | Current Gemini API family; 720p, 1080p or 4K by variant and mode |
Use one prompt and one starting image, generate a short clip, then inspect audio, prompt adherence, motion and cost per accepted output.
Open AI Video GeneratorAudio: Veo 3 introduced native sound effects, ambience and dialogue; Veo 2 generated silent video.
Resolution: the current stable Vertex Veo 2 page documents 720p; the stable Veo 3 page documents 720p and 1080p.
Duration: stable Veo 2 documents 5–8 seconds; stable Veo 3 documents 4, 6 or 8 seconds. The old claim that Veo 3 created “10+ seconds” per generation was incorrect.
Inputs: both generations had text and image workflows in at least some routes. “Veo 2 was text only” was incorrect.
Current choice: Google’s Gemini API documentation lists Veo 3 models as deprecated and documents Veo 3.1, 3.1 Fast and 3.1 Lite as the active family.
Google’s Veo 3 model page lists sound generation for Veo 3, while the Veo 2 model page does not. Audio prompts can specify dialogue, ambience and effects, but generated speech and synchronization still need a full listen.
Judge audio separately from visuals: transcription accuracy, speaker identity, timing, clipping, unintended music and rights can fail even when the video looks acceptable.
The stable Veo 2 Vertex documentation lists text-to-video, image-to-video, prompt rewriting and reference-image generation. Some experimental or preview Veo 2 routes also exposed extension, first/last frames and object editing. Availability depended on the exact model ID.
Veo 3’s stable Vertex route documents text-to-video and audio generation, while image-to-video appeared through preview routes. A broad “Veo 3 supports more input types” statement loses these endpoint-specific differences.
For the stable Vertex model pages, Veo 2 is documented at 720p, 24 fps and 5–8 seconds. Veo 3 is documented at 720p or 1080p, 24 fps and 4, 6 or 8 seconds. These are generation lengths, not a guarantee about extension workflows or every consumer interface.
Current Veo 3.1 Gemini API documentation adds 720p, 1080p and 4K options by variant and mode. It requires 8 seconds for 1080p, 4K, extension or reference-image cases, and limits extension to 720p.
Google DeepMind’s current Veo page presents Veo 3.1 as the leading generation and documents reference images, style matching, character consistency, scene extension, first and last frames, outpainting, object edits and control features. The exact API subset still depends on the selected 3.1 variant.
For an API build, save the full model ID with every accepted generation and follow deprecation notices. For a creator workflow, check which controls the current interface actually exposes rather than assuming every research or API capability appears in the app.
Use one prompt and starting image. Include motion, camera, one physical interaction and one short audio cue.
Record the complete route. Save product, model ID, variant, resolution, duration, aspect ratio and date.
Score separate dimensions. Prompt adherence, subject identity, motion, physics, composition, dialogue, ambience and audio synchronization.
Count rejected generations. Compare cost and time per accepted clip, not only list price or the strongest sample.
Retain outputs. A quality claim cannot be audited without the prompts, settings, source assets and full unedited videos.
Medium tracking shot of a bicycle mechanic rolling a repaired red bicycle from a workshop into light rain. The camera moves backward at walking speed. Tires cross one shallow puddle with physically coherent splash. Natural workshop ambience and rain; the mechanic says, ‘Ready for the road.’ No music, no on-screen text.
Native audio. Veo 3 added generated dialogue, effects and ambience, while Veo 2 outputs were silent. Resolution, duration and input differences must be compared by exact model route.
Google’s current Gemini API documentation marks the Veo 3 models as deprecated and lists Veo 3.1 variants as the active family. Existing Vertex deployments and consumer interfaces can have different availability, so check the exact surface.
Yes. Google’s stable Veo 2 Vertex model page lists image-to-video and reference-image generation. The earlier claim that Veo 2 accepted text only was false.
The stable Veo 3 Vertex model page documents 4, 6 or 8-second generations. Longer sequences require a supported extension or editing workflow; do not describe that as a single native 10+ second generation.
Google provides Veo through Gemini, Flow, Google AI Studio and APIs. Magic Hour’s AI Video Generator also lists Veo 3.1 among its current model options. Use the Veo 3.1 guide for the current workflow or the Sora 2 versus Veo 3.1 comparison for a current model decision.
