How to make an AI avatar video in Synthesia (2026)

Runbo Li
Runbo Li
·
· 4 min read
synthesia

Quick answer

To make a Synthesia video, choose a starting point, set the aspect ratio, break the script into scenes, assign an avatar and voice, add supporting media, preview, then generate. Synthesia is strongest for presenter-led training, onboarding and explainers. It also includes generative media, but an avatar presentation and a cinematic text-to-video clip remain different production jobs.

The fastest reliable workflow starts from an approved script and a short representative scene. Build and review that scene before converting an entire course or deck.

What you need before opening Synthesia

  • Audience and outcome: one sentence describing what the viewer should know or do.

  • Approved script: facts, claims, names and pronunciations checked by the owner.

  • Brand assets: logo, colors, fonts, product captures and any licensed media.

  • Presenter choice: a stock, personal or custom avatar with the required consent and usage rights.

  • Distribution format: landscape, portrait or square, plus captions and localization needs.

Step 1: choose the right starting point

Synthesia’s current video-creation documentation lists prompt, template, script, file, PowerPoint and blank-canvas starting points. Use a script when wording is already approved; use a template for repeatable layout; import a deck or document only when its structure deserves to survive.

A prompt-generated draft is useful for options, not factual approval. Compare it with the source and remove invented claims before assigning a presenter.

Step 2: set the aspect ratio and scene system

Choose the output format before laying out scenes. Synthesia’s editor guide recommends setting the aspect ratio before design and describes a scene list, canvas, inspector and script box.

Use one scene for one idea. Keep a stable grid for presenter position, headline, evidence and supporting visual so updates do not require redesigning every slide.

Step 3: choose an avatar and document consent

Synthesia’s avatar page describes stock, customizable and personal avatars. A personal avatar workflow requires a consent recording; enterprise controls and allowances vary by plan.

Match the presenter to the audience and context. Avoid a synthetic spokesperson for testimony, news, sensitive advice or any situation where the viewer could reasonably infer a real endorsement.

Step 4: write for spoken delivery

  • Lead with the answer. State the action or conclusion in the first sentence.

  • Use short clauses. Split dense policy or product language into separate spoken beats.

  • Mark pronunciation. Test names, acronyms, URLs, numbers and domain terms.

  • Show evidence. Put the relevant interface, chart or source on screen while the presenter explains it.

  • End with one action. Avoid several competing calls to action.

Step 5: assign voice and timing

Paste the script scene by scene and select the voice and language in the script box. Preview difficult lines before styling the whole video. Adjust the copy before forcing unnatural speed or pauses.

For localization, translate meaning and on-screen layout together. Names, dates, currency, captions and text expansion require a human language review even when the voice sounds fluent.

Step 6: add evidence and B-roll

Use product captures, diagrams and licensed footage to prove what the narration says. Synthesia also documents avatar and space B-roll through supported generative models in the editor.

Generated B-roll should have a written job: demonstrate a concept, bridge a cut or maintain attention. If the script needs an original cinematic insert, compare it with a dedicated AI video generator and keep the shot that survives factual, continuity and rights review.

A red ceramic mug sits on a wooden café table beside a window. Steam rises slowly while the camera makes a gentle push-in. Soft morning light, realistic materials, one continuous shot, no people or readable text.

Create one supporting AI shot

Write one missing supporting shot from your avatar script, generate it, and judge continuity, factual accuracy and edit time inside the finished scene.

Try AI Video Generator

Step 7: preview the complete experience

  • Meaning: script, image and presenter communicate the same fact.

  • Voice: pronunciation, pace, emphasis and pauses sound intentional.

  • Visuals: no cropped UI, tiny text, accidental brands or unsupported claims.

  • Transitions: cuts and scene changes do not interrupt a sentence.

  • Accessibility: captions match the final audio; color contrast and reading order work.

  • Disclosure: label synthetic presenters where context, policy or law requires it.

Step 8: generate, publish and measure

Generate only after the representative scene passes review. Synthesia’s current AI video feature page describes avatar-led videos plus generated assets, and its pricing page is the source of truth for current credits, downloads, collaboration and plan access.

Measure completion rate, the first major drop, comprehension or task completion, edit time, correction count and cost per approved minute. For training, also measure whether viewers can perform the task after watching.

When Synthesia is a good fit

  • Strong fit: repeatable training, onboarding, internal communication, product explanation and localized presenter videos.

  • Weak fit: documentary proof, authentic customer testimony, live events, unscripted emotion or a story that depends on physical performance.

  • Separate job: cinematic generative clips and complex product footage may need a dedicated model or conventional production.

Frequently asked questions

Synthesia currently offers a Basic plan for testing avatars, voices and the editor. The allowance and export features can change, so check the official pricing page before planning a full project.

Synthesia supports personal avatars and voice options with consent requirements and plan-specific allowances. Follow the live capture instructions and do not create an avatar for another person without permission.

Use one idea per scene and let the spoken sentence determine the length. Preview the scene and split it when the viewer must read, compare or follow several steps at once.

No. Its core workflow builds a structured presenter video from scripts, scenes, avatars, voices and media. Its editor can also generate assets, but a short cinematic model output does not replace the complete presentation workflow.

Use the AI avatar generator guide for broad selection, the Synthesia pricing guide for current plan detail, and the Synthesia versus HeyGen guide for a focused platform comparison.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

Sports zine collage showing selection, pacing, and editing for a highlight reel
How to make a sports highlight reel: recruiting & social
Prompting AI Videos Cover
How to prompt AI videos: a practical 10-step guide
AI Voice Cloning Laws & Ethics (2026): Consent, Licensing, and a Risk Checklist
Is AI voice cloning legal? Consent, licensing & disclosure
Turn Audio Into a Video With AI
How to turn audio into video: 5 reliable workflows (2026)
Editorial audio studio illustration comparing ElevenLabs plans, voice credits, and cost
ElevenLabs pricing (2026): plans, credits & voice cloning
ComfyUI Beginner’s Guide
ComfyUI beginner guide (2026): cloud, desktop & first workflow