How to prompt Veo 3.1 for the best results

veo3

Quick answer

Prompt Veo 3.1 as one clearly directed shot. Describe the subject, one action, setting, framing, camera movement, lighting and sound; then state what must stay unchanged. Choose duration, resolution and aspect ratio in the tool’s settings. For longer sequences, generate separate shots and edit them together. The examples below are original prompt templates, not claims of tested performance.

This guide focuses on writing and revising Veo prompts. For access and a first generation, use our Google Veo beginner’s guide. For advice that applies across models, start with the practical AI video prompting guide.

What Veo 3.1 can do—and what belongs in settings

Google’s current Gemini API documentation lists native audio, landscape 16:9 and portrait 9:16, and 4-, 6- or 8-second generation. Veo 3.1 and Fast support 720p, 1080p and 4K, with eight-second requirements for higher resolutions and reference-image workflows; extension uses 720p. Lite has different limits. Check the exact model and access surface: Gemini API, Vertex AI, Flow and a third-party tool need not expose identical controls.

A prompt does not override a model’s settings. Writing “4K” or “30 seconds” in a description is not a substitute for selecting a supported output setting. Likewise, a Google subscription name is not a model identifier. This guide does not assume a particular subscription, price or provider implementation.

Build a prompt around an acceptance check

Before writing, decide what would make the clip usable. For a product reveal, the bottle shape and printed label may matter more than dramatic lighting. For a dialogue scene, intelligible speech and the intended speaker may matter more than background detail. Put the important requirement early and keep the rest subordinate.

Use this working order: shot and framing; subject; one main action; setting; camera path; light and texture; audio; invariants. It is a planning aid, not a special syntax. Google’s official Veo 3.1 prompting guide covers cinematography, subject, action, context and ambiance. DeepMind’s prompt guide also discusses character descriptions, location, dialogue and sound.

Prompt ingredients at a glance

Element

Why It Matters

Example

Subject

Defines who/what the video is about

A woman in her 40s with curly hair wearing a red jacket

Context

Anchors the subject in space

Standing in a crowded subway station at night

Action

Brings the subject to life

She turns, looks into the camera, and waves

Style

Directs the look & feel

Cinematic, anime, claymation, 8-bit retro

Camera

Adds immersion

A dolly zoom, handheld shot, or aerial pan

Composition

Shapes framing and perspective

Wide shot vs close-up

Ambiance

Sets lighting and mood

Warm golden hour light, neon glow, rainy night

Audio

Veo’s edge: define dialogue, music, effects

A man says: Hello, with faint jazz music playing

Start with one action and one camera move

A useful first draft should be easy to storyboard. “A cyclist stops beside a café” names a visible action. “A cinematic journey through ambition, nostalgia and freedom” leaves the physical scene undefined. Add emotion through an observable choice: a hesitant hand on the brake, a lingering look through the window, or warm light on an empty chair.

Keep the first attempt simpler than the finished sequence. A pan, orbit and zoom in one short clip can conflict with a chase, wardrobe change and dialogue. Choose the movement that serves the shot. If a locked camera already proves the action works, add a slow push-in on a later attempt instead of changing everything at once.

Six original Veo prompt examples

These templates are editorial examples written for this guide. They are not recorded test results, guaranteed outputs or endorsements of a specific model’s quality. Replace the details with your own authorized subjects and assets, and evaluate every generated clip.

1. Product reveal with a stable shape

Prompt: A close-up of an unbranded brushed-steel wristwatch resting on dark stone. The camera makes one slow quarter-circle move around the watch while the watch remains still. Soft white side light reveals the case and strap texture. The watch face, proportions and strap remain unchanged throughout. Quiet studio room tone, no dialogue, no music. One continuous shot.

Frame from one eight-second Magic Hour Veo 3.1 generation of the article’s wristwatch product-reveal prompt, September 28, 2026

Frame from one eight-second Veo 3.1 generation using the exact prompt above (September 28, 2026). The orbit and side lighting are visible; small dial markings are not reliable product copy. Review the complete clip before using it.

Review: Check the case geometry, hands, strap and reflections frame by frame. If the watch changes shape, remove the orbit and try a locked shot. Add a logo or exact small print in an editor when it must be accurate; a plausible-looking label is not an acceptable substitute for the real asset.

2. A short line of dialogue

Prompt: A medium close-up of one adult florist standing behind a wooden counter in a quiet flower shop. She looks toward the camera and says, “These are ready for your first dinner party.” Her delivery is relaxed and conversational. The camera stays still. Soft daylight, faint shop ambience, no other voices, no music, no on-screen text.

Frame from one eight-second Magic Hour Veo 3.1 generation of the article’s florist dialogue prompt, September 28, 2026

Frame from one eight-second Veo 3.1 generation using the exact prompt above (September 28, 2026). The florist and shop match the visual brief. This still does not verify the exact spoken words or lip-sync; review the full video and audio before publishing.

Review: Check the exact words, voice, lip movement and background audio. If the line is rushed or cut off, shorten it before changing the visual scene. Keep a single speaker when diagnosing speech problems. An instruction to exclude captions is a request, not a guarantee; inspect the final frames for unwanted text.

3. Animate an existing image

Prompt: Use the uploaded image as the starting composition. The person gently turns toward the window and smiles once. Keep the face, clothing, hairstyle and room layout consistent with the image. Locked camera, soft window light. A quiet indoor soundscape; no dialogue. No new people or objects enter the frame.

Review: When the input already establishes appearance and composition, focus the prompt on motion and what must remain stable. Look for identity drift, extra hands and changes to the background. If the task needs a specific reference-image feature, confirm that the selected model and tool expose it; a normal starting frame and multiple reference images are different controls.

4. A vertical creator clip

Prompt: Portrait composition. One adult travel presenter stands beside a market stall, with her face and hands comfortably inside the frame. She points once toward a bowl of fruit and says, “Here is the one thing I would try first.” Gentle handheld movement, daylight, low market ambience beneath the speech. No captions or music.

Frame from one eight-second Magic Hour Veo 3.1 generation of the article’s vertical market-presenter prompt, September 28, 2026

Frame from one eight-second Veo 3.1 generation using the exact prompt above (September 28, 2026; 9:16). The portrait framing keeps the presenter visible. This still does not verify the spoken sentence or sound; inspect the full clip before use.

Review: Select 9:16 in the actual output settings. Leave room around the subject for interface overlays and later captions. Check whether the hand gesture stays visible and whether background sound obscures speech. Do not rely on a landscape crop to rescue a shot whose important action extends outside the portrait frame.

5. A thriller insert rather than an entire scene

Prompt: Extreme close-up of an adult’s hand placing a small brass key on a scratched wooden table. The hand pauses, then withdraws. Locked camera, shallow depth of field, cold side light. A distant ventilation hum and one quiet metallic tap as the key touches the table. No music, no cut, no additional action.

Frame from one eight-second Magic Hour Veo 3.1 generation of the article’s brass-key suspense prompt, September 28, 2026

Frame from one eight-second Veo 3.1 generation using the exact prompt above (September 28, 2026). The key, hand and cold light appear in one shot; inspect contact and sound in the complete clip before use.

Review: This is one insert shot, not a complete suspense sequence. Generate the face reaction and room-wide shot separately, then edit them together. Check contact between fingers, key and table. If the object floats or multiplies, simplify the action rather than adding more adjectives about suspense.

6. A controlled sports moment

Prompt: A side-on medium shot of one adult table-tennis player preparing to serve on an empty practice court. The player tosses the ball once and strikes it forward. The camera remains fixed. Even overhead lighting, realistic practice-room ambience, one racket impact, no music or spectators. Keep the player and table in the same positions.

Frame from one eight-second Magic Hour Veo 3.1 generation of the article’s table-tennis serve prompt, September 28, 2026

Frame from one eight-second Veo 3.1 generation using the exact prompt above (September 28, 2026). The player and table stay in frame. A still cannot establish that ball toss, contact, path or racket sound are correct; review the whole clip.

Review: Inspect the ball path, hand contact and racket motion. Treat fast interactions as acceptance checks, not as evidence that any model always succeeds or fails at sports. For a finished highlight, combine short accepted actions with editing instead of asking one generation to handle a rally, crowd, slow motion and a branded end card.

Try one prompt

Start with one shot, then check the full result against your acceptance criteria.

Troubleshooting: change the instruction tied to the failure

The shot looks generic. Replace abstract praise with physical detail. “Premium” can become brushed metal, a matte surface, controlled rim light and a clean composition. “Cinematic” can become a specific camera height, lens feel or movement. Keep only details that change what a viewer can see or hear.

The subject changes. Reduce simultaneous actions and appearance changes. Use an authorized starting image or reference control when available. Keep descriptions stable between shots, but do not promise that repeated wording will guarantee identical characters.

The camera ignores the plan. Choose one camera behavior, state its direction and pace, and remove conflicting movements. Compare a locked shot with a slow push-in before trying a more complex move. Evaluate the actual clip rather than assuming film terminology was followed.

Audio is distracting. Separate the requested speech, ambient sound, effects and music in the description. Name the speaker and keep the spoken line brief. If audio is not usable, repair or replace it in an editor rather than claiming the video is finished.

Text or subtitles appear. State the desired clean-frame result and inspect the output. Generate a simpler scene if signage is creating unwanted lettering. Add accurate titles and captions during editing. Do not claim that a particular negative phrase reliably removes all text.

Measure usable clips, not impressive screenshots

Save the model identifier, provider, settings, prompt, input assets and every attempt included in the comparison. Define acceptance before generating: correct subject, action, stable geometry, intended camera, intelligible audio and no unwanted text. A selected success without the rejected attempts cannot establish a success rate.

Record accepted clips divided by total completed attempts, total generation cost, repair time and the final export requirements. Include failed or rejected attempts in the cost calculation when they incurred cost. Keep generation time separate from editing and review time. This is a proposed evaluation method, not a benchmark we have already run.

Compare tools on the same brief and inputs, with a fixed retry allowance. State whether the objective is dialogue, product consistency or a particular visual style. Our AI video generator comparison and Kling versus Veo guide provide broader selection context; neither replaces your task-specific acceptance check.

Frequently asked questions

Yes. Google documents 9:16 as well as 16:9. Select the aspect ratio in the tool or API settings and verify that your chosen provider exposes the control.

Use JSON as an organizational aid if it helps you write a clear brief. Do not treat a JSON-shaped text prompt as an API configuration object. Actual duration, model and resolution settings belong in the documented request or tool controls.

Use enough detail to specify the shot and its important constraints. There is no useful universal word-count target. Remove conflicting actions and decorative language before adding length.

No prompt wording guarantees consistency. Use supported image or reference controls, simplify changes between shots, and review every output before accepting it.

No. Choose by the job and measure accepted output, cost and repair effort. Native audio is useful when the task needs it; it does not establish overall superiority for every workflow.

Related creative workflows

For inspiration rather than proof of model performance, explore top AI influencers on Instagram and product video examples. Keep any creative reference distinct from a measured claim about Veo.

Magic Hour official workflow page screenshot (magic-hour-textvideo), September 25, 2026

Magic Hour input workflow, captured September 25, 2026. This shows the upload or prompt controls, not a completed generation. View official source

If the workflow needs a portrait asset first, see AI headshot generators. For existing footage that needs dialogue synchronization, see the lip-sync workflow guide. These are separate production steps, not required add-ons for every Veo generation.

For broader context, our creative-economy statistics and Studio Ghibli AI trend article cover industry and style discussions. If you are choosing a creation workflow, start with Magic Hour’s AI video generator overview.

Model facts checked September 25, 2026 against the Google sources linked above. Prompt templates and evaluation advice are editorial guidance; this article makes no claim of a completed first-party Veo benchmark.

David's Portrait
David Hu
Co-founder & CTO of Magic Hour
David Hu is the Co-founder and CTO of Magic Hour, where he leads engineering for AI video, image, audio, and developer products. Previously, he was a full-stack engineer at Skillz and led product development teams as the company grew from Series B through its IPO. His work spans user interfaces, APIs, media systems, and infrastructure. He writes about AI media engineering, model integrations, APIs, and production workflows.
View author →

Continue Reading

veo3
Recommended next
Google Veo 3.1: a beginner's guide to AI video

Learn what Veo 3.1 is, where to access it, how to write a first prompt, when to use images or references, and how to review the generated video.

Prompting AI Videos Cover
How to prompt AI videos: a practical 10-step guide
Analog filmmaker workbench comparing AI video generation workflows and finished frames
10 best AI video generators in 2026: models, features, and costs
Kling 3.0 vs Veo 3.1 AI video model comparison showing cinematic video generation and motion quality differences
Kling 3.0 vs Veo 3.1 (2026): controls, audio & API cost
VEO3
Veo 3.1 dialogue prompts: a practical speaking guide