Magic Hour
  • Pricing
Video

Start your first video in under 60 seconds

Generate or edit video, image, and audio - free to start.
Start Creating Free
No credit card requiredFree daily creditsNo signup required
Join the Discord
Video
Company
PricingAboutBlogChangelogAPISkillsAll ToolsTemplatesAI ModelsTrust & Data UsePrivacy PolicyTerms of ServiceRefund Policy
Video Products
AI Avatar GeneratorAI UGC Ad GeneratorAI Video DubbingAI Video EditorAI Video ExpanderAI Video ExtenderAI Video TranslatorAI Video UpscalerAnimationAudio-to-VideoCharacter ReplaceColor GraderFace Swap VideoImage-to-VideoLip SyncMusic Video GeneratorSubtitle GeneratorTalking PhotoText-to-VideoVideo ColorizerVideo-to-Video
Image Products
AI Clothes ChangerAI Face EditorAI GIF GeneratorAI Headshot GeneratorAI Image EditorAI Image ExpanderAI Image GeneratorAI Image UpscalerAI Influencer GeneratorAI Meme GeneratorAI Selfie GeneratorAI Storyboard GeneratorBackground RemoverBody SwapFace Swap PhotoGenerative FillHead SwapPhoto ColorizerQR Code GeneratorCharactersMoodboards
Audio Products
AI Audio TranslatorAI Music GeneratorAI Sound Effect GeneratorVideo-to-AudioVoice ChangerVoice ClonerVoice Generator
Support
CommunityFAQHelp CenterContact UsStatus
Social
Instagram
X
TikTok
Facebook
YouTube
LinkedIn
support@magichour.ai
Backed byCombinator

© 2026 Magic Hour AI, Inc.

Back to Blog
  1. Blog
  2. Trends

Veo 3.1 dialogue prompts: a practical speaking guide

Runbo Li
Runbo Li
·
CEO of Magic Hour
·
Aug 27, 2025· 5 min read
AI Summary:
ChatGPTClaudeGeminiPerplexity
VEO3

Contents

Create with Magic Hour
Make videos and images with AI.

Quick answer

For reliable dialogue in Veo 3.1, identify one visible speaker, put the exact line in quotation marks, describe delivery separately, and keep the action simple enough for the clip. Example: “Medium close-up of one baker facing the camera. She says, ‘The bread is ready,’ in a calm, warm voice. Quiet kitchen ambience; no music.” Generate with audio enabled, then inspect the actual words and lip movement before publishing.

Veo can generate native audio, dialogue, sound effects, and ambience, but a prompt is direction rather than a guarantee. Shorter lines, clear speaker attribution, and one controlled shot make failures easier to diagnose.

A reusable Veo dialogue prompt formula

[Shot and camera] + [speaker] + [action] + [exact dialogue] + [delivery] + [ambience and sound] + [visual style].

Google's general Veo framework is cinematography, subject, action, context, style, and ambience. For speaking scenes, add explicit speaker attribution and the exact quoted line. Do not bury the dialogue inside several unrelated actions.

prompt method

Three prompt structures that work for different jobs

1. One speaker to camera

Locked medium close-up of a woman in a quiet pottery studio, looking into the camera. She says, ‘Every piece starts with a patient hand,’ in a natural, thoughtful voice. Soft room tone and a faint pottery wheel; no music. Warm documentary lighting.

Try in Text-to-Video

Use this for a product line, testimonial-style scene, presenter moment, or social hook. State whether the person addresses the camera or another character.

2. Two-person exchange

Static two-shot in a small train compartment at night. The conductor looks at the traveler and quietly asks, ‘Are you sure this is your stop?’ The traveler glances toward the dark window and replies, ‘It has to be.’ Low train rumble, no music, restrained suspense.

Try in Text-to-Video

Give each person one short line and an observable action. If speaker attribution, timing, or lip sync fails, generate each line as a separate shot and assemble the exchange in an editor.

3. Voice plus environmental sound

Wide shot of a mechanic closing the hood of a vintage car in an open garage. He says, ‘Now listen to that engine,’ with quiet pride. The engine starts and settles into a smooth idle; distant birds and light workshop ambience; no background music.

Try in Text-to-Video

Separate dialogue, sound effects, and ambience so each has a clear role. Avoid asking for several loud effects under a quiet line.

Test a Veo 3.1 dialogue prompt

Generate a short Veo 3.1 dialogue shot, then review the spoken words, speaker attribution, lip movement, ambience and disclosure before using it.

Create with Veo 3.1

How to make the speech easier to generate

  • Use one primary speaker per shot. Multi-person speech increases attribution and timing complexity.
  • Keep the line short. A line that does not fit naturally in the chosen duration is likely to rush, truncate, or drift.
  • Put exact speech in quotation marks. Google explicitly recommends quotation marks for specific dialogue.
  • Name the delivery. Describe tone, pace, volume, accent only when necessary, and emotional state in plain language.
  • Describe visible mouth orientation. A speaker facing away, covered, tiny in frame, or moving rapidly gives the model less visual information for lip movement.
  • Control the soundstage. State the important ambience and effects; say “no music” only when silence behind speech matters.
  • Limit simultaneous action. One camera move and one character action are easier to evaluate than a crowded mini-film.

Structured, narrative, or reference-led prompting?

  • Structured prompt: best when exact speech, staging, and sound cues matter. Label the shot, speaker, dialogue, delivery, SFX, and ambience.
  • Narrative prompt: useful for a fast first exploration, but harder to debug when the spoken result is wrong.
  • Image-to-video: useful when the first frame should establish the speaker, wardrobe, product, or composition. The image guides the visual start; it does not guarantee identity or speech accuracy.
  • Reference images: Veo 3.1's Gemini API supports up to three references for a person, character, or product. Use only images you have permission to process.
  • Separate shots: usually safer for an exchange or longer script. Generate one short line per shot, then edit the scene and sound continuity.

Common dialogue failures and what to change

  • Wrong words or extra speech: shorten the line, remove competing narration, and keep one quoted utterance.
  • The wrong person speaks: use one speaker, a closer shot, explicit names or descriptions, and separate the exchange into shots.
  • Weak lip sync: face the speaker toward camera, reduce motion, shorten speech, and inspect at normal speed and frame by frame.
  • Rushed delivery: reduce word count or use a longer supported duration in the provider you are using.
  • Speech masked by audio: simplify effects, lower the described ambience, and remove music from the generation or mix it later.
  • Character drift: use a permitted reference or first frame, hold the description stable, and generate separate shots. Repetition can help direction but cannot ensure identity.
  • Blocked generation: review the prompt and reference media against the provider's safety rules. Google notes that audio processing can also cause a Veo generation to be blocked.

A controlled way to improve a prompt

  • Start with one speaker, one short line, a locked or simple shot, and quiet ambience.
  • Generate several candidates and save the prompt, model, provider settings, duration, aspect ratio, resolution, audio setting, and output.
  • Score the exact words, speaker attribution, lip movement, vocal delivery, visual continuity, and sound balance.
  • Change one variable at a time: line length, framing, delivery, motion, or ambience.
  • Once speech works, add one visual or sound detail at a time.
  • Choose the lowest-cost candidate that clears the intended quality bar, then finish captions and audio in an editor when needed.

Current Veo 3.1 limits depend on where you use it

In the Gemini API, Google documents Veo 3.1 clips of 4, 6, or 8 seconds, with 8 seconds required for some higher-resolution or reference-image modes. It supports native audio and 16:9 or 9:16 output. Vertex AI and consumer Gemini access, quotas, prices, and controls are separate surfaces and can change.

Inside Magic Hour's Veo 3.1 workflow, current platform settings can differ from Google's direct Gemini API. Check the selected interface for duration, resolution, aspect ratio, audio, price, plan, and commercial-use terms instead of carrying over an old Gemini Ultra price or temporary free-access promotion.

Dialogue QA before publishing

  • Transcribe the generated speech and compare it word for word with the intended line.
  • Confirm the intended person speaks each line and no extra voice appears.
  • Watch lip movement, teeth, tongue, jaw, eyes, hands, and identity at full size.
  • Listen on headphones and a phone speaker for clipping, noise, phase issues, abrupt ambience, or masked speech.
  • Add accurate captions and identify synthetic media when context could mislead viewers.
  • Confirm consent for any recognizable person or voice and rights for references, brands, music, and other assets.
  • Reject a plausible-looking result when the words, identity, safety, or disclosure are wrong.

Frequently asked questions

Yes. Google's Veo guidance explicitly recommends quotation marks for specific speech. Also identify who says the line and how they deliver it.

Yes, Google's examples include multi-person dialogue. For production reliability, keep turns short and attribution explicit; generate separate shots if the speakers or timing drift.

No. Stable descriptions and reference inputs can guide continuity, but they do not guarantee identity. Review every shot and use permitted references where the chosen Veo surface supports them.

No. Google's current Gemini API documentation lists 4-, 6-, and 8-second choices, with mode-specific restrictions. Other providers can expose different controls, so verify the selected interface.

It should not be treated as an exact recording system. Verify the words, speaker, voice, lip movement, rights, and disclosure before publishing or using it in customer-facing work.

Official sources checked

Current Gemini API capabilities, durations, audio cues, safety notes, and reference behavior come from Google's Veo 3.1 developer guide. Prompt structure and soundstage guidance come from the Google Cloud Veo 3.1 prompting guide. Current Magic Hour access comes from the Magic Hour Veo 3.1 model page. Sources were checked September 13, 2026.

Runbo Li
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
See all articles

Continue Reading

veo3
Recommended next
Videos
Google Veo 3.1: a beginner's guide to AI video

Learn what Veo 3.1 is, where to access it, how to write a first prompt, when to use images or references, and how to review the generated video.

Jul 18, 2025
Kling 3.0 vs Veo 3.1 AI video model comparison showing cinematic video generation and motion quality differences
Videos
Kling 3.0 vs Veo 3.1 (2026): controls, audio & API cost
Mar 05, 2026
Editorial triptych comparing Veo 3.1, Kling 3.0 and PixVerse V6 video workflows
Videos
Veo 3.1 vs Kling 3.0 vs PixVerse V6: which video model?
Aug 14, 2025
Veo
Videos
Veo 2 vs Veo 3: differences and current status (2026)
Jul 21, 2025
AI Skills
Trends
9 AI skills for building useful, paid services
Aug 02, 2025