7 best subtitle APIs: caption files, video output and costs

Runbo Li
Runbo Li
·
· 8 min read
Best subtitle APIs for creators and developers in 2026

Quick answer

For a finished video with styled captions, shortlist Magic Hour, VEED or Submagic. For editable SRT/VTT subtitle files, start with AssemblyAI, Deepgram, OpenAI Whisper or Google Cloud Speech-to-Text. The best subtitle API depends first on the output you need: a captioned video, a separate subtitle track or timestamped transcript data.

Test subtitle output in the browser

Upload a clip with real names, numbers, accents, and background noise. Correct the transcript and timing before committing an automated pipeline.

Open Auto Subtitle API Docs

Those outputs are different products. A speech-to-text API can recognize every word correctly and still leave your application responsible for caption timing, line breaks, styling and video rendering.

Magic Hour publishes this guide and is included where relevant. We selected current options with official product documentation and a distinct fit for the tasks in this guide, then compared documented inputs, controls, limits, exports, pricing mechanics, and workflow fit. This is not a controlled output-quality benchmark unless a retained test is explicitly described below.

Magic Hour publishes this guide and includes its own product in the comparison. Treat the recommendations as editorial guidance from a vendor, and verify the linked first-party product details and your own output requirements before choosing a tool.

Best subtitle APIs compared

API

Best workflow fit

Output to expect

Main integration decision

Magic Hour

Adding styled captions within an image/video production workflow

Rendered video from the Auto Subtitle Generator

Choose the clip segment and caption style; retrieve the completed video job

VEED through fal

Automated social-video caption rendering

Video with burned-in subtitles

Basic versus dynamic styling, resolution and optional translation affect cost

Submagic

Reviewing words before rendering short-form captions

Editable project transcript, then exported video

Disable automatic rendering when a person must approve the captions

AssemblyAI

Downloadable subtitle files and custom caption interfaces

SRT/VTT exports plus timestamped transcript data

Choose the recognition model and caption length

Deepgram

A speech pipeline with your own caption formatter

Timestamped JSON; convert it to SRT/VTT

Separate speech recognition from caption layout and rendering

OpenAI Whisper

Subtitle files in an existing OpenAI integration

SRT/VTT or detailed timestamps with whisper-1

Match the output format to the specific transcription model

Google Cloud Speech-to-Text

Batch caption files in a Google Cloud workflow

V2 BatchRecognize SRT/VTT outputs

Configure batch output formats, storage and the applicable recognition rate

Reviewed September 10, 2026. This is a documentation-based comparison of workflow fit, outputs and billing, not a new accuracy or speed benchmark. Magic Hour publishes this guide and is included in the comparison.

Choose the output before the provider

Burned-in captions become part of the video image. Use them when the text needs to appear consistently in a social clip, including when playback is muted. Changing a word afterward requires another render.

SRT or WebVTT files are separate timed text tracks. Use them when a video platform, player, editor or localization workflow accepts subtitle files. Viewers may be able to toggle a supported track on or off; the player controls how it appears.

Timestamped JSON gives your application words or segments with timing information. It is useful for a custom editor, search, speaker views and word highlighting. Your application must turn that data into readable captions or a rendered video.

For example, a podcast publisher may need an editable VTT track for its website and a vertical video with highlighted words for social media. One transcription can feed both workflows, but the final deliverables require different formatting and rendering steps.

1. Magic Hour: captioned videos alongside other media tools

The Magic Hour Auto Subtitle Generator API takes an uploaded video asset and a start/end segment. Its style settings include a karaoke template and controls for font, size, text and highlight colors, stroke and placement.

Why shortlist it: The output is a video job, which fits an application that also uses Magic Hour for generation, editing or other media operations. Save the returned project ID and retrieve the finished video rather than treating the create response as a download.

Main limitation: This endpoint documents a captioned-video workflow. Do not assume it also returns a separate SRT file, translates the dialogue or generates a dubbed audio track. Those are separate requirements to verify.

The model and credit reference lists Auto Subtitle Generator at approximately 4.8 credits per second for 24 fps video. Use the completed job's charge for accounting and your actual account rate to convert credits into dollars.

2. VEED: styled rendering with optional supplied subtitles

VEED exposes its Subtitles API through fal. It produces a rendered video with burned-in captions. Basic and dynamic presets offer different styling; input resolution and optional translation also change the charge.

Why shortlist it: The API schema accepts either a video for transcription or supplied SRT text/a file URL. Supplied SRT skips transcription, so an application can correct a transcript before sending it to the renderer. A vocabulary field helps specify brand names when using automatic transcription.

Main limitation: A rendered MP4 is the documented output. If your player also needs an editable subtitle track, retain that track separately. Check the actual preset and resolution in your request; an attractive dynamic preset is not billed at the basic styling rate.

3. Submagic: transcript review before the final export

The Submagic API creates projects from video uploads or URLs and supports caption templates, language selection and completion webhooks.

Its useful distinction is the documented edit-before-render workflow. Create a project with autoRender set to false, wait for transcription, edit the words and their timings, then explicitly export the video.

Why shortlist it: This provides an approval step for product names, technical vocabulary and other text that must be correct before publishing.

Main limitation: In this workflow, a completed transcription status does not mean the final video exists. The download URL appears after export. Keep transcript completion and rendered-video completion separate in your application. Confirm API access, usage allowances and export limits for your account instead of treating a consumer subscription price as a universal API rate.

4. AssemblyAI: SRT/VTT exports and caption data

AssemblyAI's transcript export documentation supports SRT and VTT from completed transcripts, with a chars_per_caption parameter for caption length. The transcript also exposes word-level timing, which can support a custom editor.

Why shortlist it: It gives a direct path from audio recognition to a subtitle file. Use its word or utterance data when you need more control than the standard export provides.

Main limitation: The plain text field is not a timed subtitle track. Speaker labels and readable line breaks need deliberate handling; exporting text alone will not preserve everything your caption UI needs.

Current pay-as-you-go pricing lists pre-recorded Universal-2 at $0.15 per hour and Universal-3.5 Pro at $0.21 per hour, with separately priced add-ons. Select the model for your language and recording conditions rather than assuming every feature is included in the base rate.

5. Deepgram: recognition plus a separate caption formatter

Deepgram's caption guide explains how to turn its timestamped transcription response into SRT or WebVTT using its caption utilities.

Why shortlist it: This separation is useful when your application already owns the editor or player. You can preserve recognition data and change how words are grouped without pretending a plain transcript is the finished viewing experience.

Main limitation: The speech API and caption formatting layer do not by themselves burn animated words into a video. You still need a renderer for that output.

On the current pricing page, pre-recorded Nova-3 pay-as-you-go rates are $0.0043 per minute for the monolingual option and $0.0052 for multilingual. Streaming prices and optional add-ons are different; use the pre-recorded rate for an uploaded-video calculation.

6. OpenAI: choose Whisper when the output needs subtitle formats

OpenAI's file-transcription guide recommends gpt-transcribe for general recorded-speech transcription, while directing specialized requirements such as word timestamps and subtitle formats to the appropriate model. These models are not interchangeable output-format choices.

For a direct subtitle-file workflow, whisper-1 supports SRT/VTT. For word-level timing, request verbose JSON and the required timestamp granularity. Check the transcription API reference when selecting a response format.

Why shortlist it: It can fit an existing OpenAI application without adding another transcription provider.

Main limitation: The documented file-transcription upload limit is 25 MB. Compress or split longer recordings appropriately, and account for timing offsets if you combine chunks. The API produces text/data, not a styled video.

The pricing page lists Whisper at $0.006 per minute. A newer transcription model's lower rate does not imply that it accepts the same subtitle response formats.

7. Google Cloud: caption files from batch transcription

Google Cloud documents SRT and VTT caption outputs through Speech-to-Text V2's BatchRecognize method. Configure output_format_config and retrieve the asynchronous result. Outputs can go to Cloud Storage or be returned inline, with supported multi-format configurations.

Why shortlist it: It is a practical fit when recordings and processing already live in Google Cloud and the deliverable is a subtitle file.

Main limitation: This documented caption-output feature belongs to V2 batch recognition. Do not assume the same request works with every real-time or older API endpoint.

Speech-to-Text pricing distinguishes standard recognition, dynamic batch, volume tiers and discounts. At the first standard V2 tier, the listed rate is $0.016 per minute. Storage and other cloud services can add costs.

What does 100 minutes of subtitles cost?

This example uses 100 source minutes, one processing attempt and no optional add-ons unless specified. It compares clearly named operations, not equivalent end products: recognition and styled video rendering include different work.

Operation

Calculation

Example usage cost

Magic Hour Auto Subtitle Generator, 24 fps reference

100 × 60 × approximately 4.8 credits

Approximately 28,800 credits; convert at your account rate

VEED through fal, basic styling, at or below 1080p, no translation

100 × $0.10

$10 for rendered captions

VEED through fal, dynamic styling, at or below 1080p, no translation

100 × $0.10 × 2

$20 for rendered captions

AssemblyAI Universal-2, pre-recorded base recognition

100 ÷ 60 × $0.15

$0.25 before add-ons

Deepgram Nova-3 monolingual, pre-recorded

100 × $0.0043

$0.43 before add-ons/rendering

OpenAI Whisper

100 × $0.006

$0.60 before rendering

Google Cloud V2 standard, first volume tier

100 × $0.016

$1.60 before other cloud services

VEED's listed quote applies a two-times multiplier above 1080p and a two-times multiplier for dynamic styling; they compound. Translation adds $0.20 per minute. Its one-minute minimum charge also matters: 100 separate 30-second clips can cost more than one 50-minute input.

Treat these as reproducible calculations from the linked rate cards, not invoices or measured production costs. Add human correction, retries, rendering, storage and delivery where applicable. Compare total cost per accepted captioned minute when selecting your production workflow.

A practical workflow for readable, accurate captions

  1. Choose the deliverable. Decide whether you need an MP4 with visible text, an SRT/VTT track or both.
  2. Transcribe representative audio. Include your actual accents, background sound, speaker changes and product vocabulary.
  3. Correct meaning before styling. Review names, numbers, negations and technical terms. A wrong price in attractive captions is still a wrong price.
  4. Inspect timing and layout. Watch the opening, middle and end. Check drift, rapid line changes, awkward phrase breaks and words hidden behind interface controls.
  5. Export the intended formats. Keep the corrected text and timestamps when you also create a rendered social clip.
  6. Check the delivered result. Open the caption file in the intended player or watch the final video. A successful API response alone does not prove readable captions.

Run this as a small evaluation before integrating at volume. Record correction time, failed jobs and actual charges. This guide has not run a new cross-provider accuracy test, so it does not assign unsupported accuracy percentages or fastest-provider rankings.

Frequently asked questions

No. It manages caption tracks associated with YouTube videos. The caption download endpoint requires authorization with permission to edit the video. It is not a general-purpose transcription service or a universal downloader for public videos. For footage you are authorized to process, use a transcription API when you need to create a new track.

Only if the selected endpoint explicitly includes that operation. Transcription preserves spoken content as text; translation changes its language; dubbing produces replacement speech. Lip sync is another step. See our AI dubbing workflow if you need the whole sequence.

Choose a renderer with a documented highlighting style, such as Magic Hour's karaoke template, or build from word timestamps and a separate video renderer. An ordinary SRT export is not evidence that a service produces animated, word-by-word video captions.

Start with the Magic Hour Auto Subtitle Generator on a short representative clip. For an editor comparison, see our AI subtitle tools guide. Use the API comparison above when you are ready to automate the workflow that fits your output requirements.

Runbo Li
Runbo Li
CEO of Magic Hour
Runbo Li is the Co-founder and CEO of Magic Hour, where he builds AI video and image tools for content creation. He is a Y Combinator W24 founder and former Data Scientist at Meta, where he worked on 0-1 consumer social products in New Product Experimentation. He writes about AI video generation, AI image creation, creative workflows, and creator tools.
View author →

Continue Reading

AI Subtitle Generators (2026)
Recommended next
7 best AI subtitle generators by workflow

Compare Magic Hour, Descript, VEED, CapCut, Canva, Happy Scribe and Subtitle Edit for transcription, styling, translation, export and subtitle QA.

AI Subtitle Gen
5 best free AI subtitle tools: limits and exports (2026)
bestaitools
Best AI tools by task: a practical shortlist for 2026
Collage of logos from the beat creative automation platforms
8 best creative automation platforms in 2026: choose by workflow
5 Best AI Image Generators & their API
5 best AI image generation APIs in 2026