New 250GB Plans LIVE now. See plans →
All posts
April 26, 2026 · AI

Create SRT File From Video: The Fast Way to Client Sign-Off

The workflow I use to create an SRT file from video, catch misspelled names before export, and get client approval on caption wording at the exact frame instead of over voice notes.

SM
Saumyajit Maity
Co-founder, PlayPause

The SRT file for Onboarding_Module_MV3.mp4 is due to the client's training team in the morning, and at 00:01:14 the founder says her own product name out loud, and the machine transcript under the player has spelled it three different ways inside one minute, right. So before anybody can create an SRT file from this video and ship it, somebody has to agree on how the words are spelled.

And that somebody is never me, right, because I genuinely do not know whether the product name carries a capital in the middle, or whether the founder's surname has a second h, or whether the figure she quotes at 00:02:31 is the one legal signed off on. In my agency this is the part that quietly eats a whole day, because the transcript sits in one tool, the cut in another, and the corrections arrive as a voice note at seven in the morning.

So let me walk you through the workflow that gave me that day back, basically, where the transcript is generated in the same place the client watches the cut, the client flags each wrong word at the exact frame, and the caption file only leaves once the wording has a yes against it.

What an SRT file actually is, and why the client keeps asking for one

An SRT is about as simple as a file format gets, right, it is plain text made of numbered blocks, and each block carries a start time, an end time and the line or two of words that sit on screen between them. You can read one in any text editor like a shopping list, which is why nearly every video platform accepts it. If you want the formal background on how subtitles are defined and how that differs from closed captioning for accessibility, those two pages are ten minutes well spent.

What one SRT block looks like

The formatting is where hand-made files break, so here is one block in plain words. The first line is the block number, say 14. The second line holds the start time, 00:01:14,200, then an arrow made of two hyphens and a greater than sign, then the end time, 00:01:17,900. The third line, and optionally a fourth, is the caption text, and then one empty line before block 15 begins. The thing that trips people up is the comma before the milliseconds, right, it has to be a comma and not a full stop, and plenty of players will quietly reject the whole file over that one character.

My house rule for the text itself, borrowed from the streaming style guides, is two lines at most, roughly 42 characters a line, and no block on screen for more than about seven seconds, because anything longer becomes a paragraph the viewer reads instead of watching the picture.

The reason a client wants the file rather than words burned into the picture is that the file keeps working after handover. YouTube reads it, LinkedIn reads it, most learning portals read it, and a translation vendor can turn it into every language they need, while burned-in captions are just pixels. I covered the platform side in making an SRT for YouTube, and the reach argument lives in the captions and accessibility guide.

The catch here is that because the format is so simple, everybody assumes it is a five minute job, and it is, right up until somebody says a proper noun, a product name, a price or a date, and then you are running an approval job, which moves at a very very different speed.

Approve the words before the file leaves

A client asking for the SRT is usually feeding several systems with it, from video platforms to translation vendors, so the wording has to be agreed before export.

How to generate SRT from video with AI transcription

Here is the sequence I run, and it is deliberately boring. First I wait for picture lock, because captioning a rough cut of the edit means every later trim shifts every timecode after it. Then the locked cut goes up as a version on the project card, so the third pass lands as MV3 on the same card as MV1 and MV2 instead of as a fresh upload with a slightly different name, right.

AI transcription then runs on that version, which is on the Agency plan and up, and it gives you a clickable transcript beside the player, so you click any line and the playhead jumps to that moment. That sounds small and it is honestly the whole trick, because I can audit a fifteen minute video in roughly the time it takes to read it.

The reason I transcribe inside the review tool, to be very honest, is about where the corrections end up living. In my agency, the second a transcript gets mailed around as a document it forks, and one of those two versions is going to end up in the SRT, so I want every correction sitting as a comment on the same link as the video. Does that make sense, right.

1Lock picture and upload the cut as a version
2Run AI transcription and read the clickable transcript
3Send the same share link to the client for wording notes
4Export the SRT and apply every agreed correction
5Import the file into your editor and style it there

If you have never run transcription inside a review workflow, getting started with AI transcriptions covers the setup, and the team habits are on the pages for AI transcription for corporate training teams and searchable transcripts for podcast producers, because a compliance module and a two hour interview need different handling.

Check the names, the jargon and the numbers before you export

Before the client sees anything I do one pass of my own, and on a fifteen minute corporate video it takes me about six minutes, because I read the transcript instead of watching the video and I only hunt for five kinds of word. I start with people, companies and places, then product names and internal jargon, because a machine has never heard of somebody's project codename and will cheerfully turn it into two ordinary English words.

After that come acronyms, which get spelled out when they should be letters, then numbers, units and currencies, and last the homophones, which are real words so nothing flags them, and that is how their and there slide through every automated check, right.

  • People, company and place names
  • Product names and internal jargon
  • Acronyms that should be letters not words
  • Numbers, prices, dates and units
  • Homophones the machine will never flag

When the brand name shows up three different ways, I know I have a question rather than an error, and I never guess, because guessing is how the CEO's name ends up spelled wrong on a public LinkedIn post. I leave each one as a comment at its frame, phrased as a question, so the client answers something specific instead of proofreading fifteen minutes of text.

For instance, on corporate pieces I regularly see a compliance acronym turned into an ordinary word, and usually it is harmless, but now and then that one word flips the meaning of the sentence, and your ear fills in what it expects on a casual watch, which is why I read the text cold every time.

Let the client correct the captions at the exact frame

This is the part that actually changes the timeline of the job, right. The client opens one link in a browser, no account and no install, and leaves a comment that sticks to the exact frame where the wrong word is spoken. For a phrase that runs a few seconds they can leave a range comment, and they can draw on the frame if pointing is easier. Replies thread underneath, @mentions pull in whoever knows how the surname is spelled, and a custom status like wording approved gives everybody a visible line in the sand. That is what an SRT file for client review should look like, the words checked against the picture rather than in a spreadsheet.

Compare that with the usual route, a document of tracked changes, a separate video link, and an email saying around 1:14, the name. Both get you to correct captions eventually, but the document route costs extra round trips, and in my agency round trips are the only thing that has ever made me miss a deadline. I unpacked why the frame matters in what frame-accurate commenting means, and the short-form version of this caption approval workflow is in getting sign-off on auto-generated captions.

Old way

Transcript in a document, video on a separate link, corrections described as somewhere around the one minute mark

With PlayPause

Transcript beside the player and every correction pinned to the frame where the word is spoken

When the client goes quiet instead of approving, the fix is contractual rather than technical, and I wrote that up as a deemed acceptance clause. Captions stall more than anything else, because approving wording feels like somebody else's job to everybody on the client side, and if it is part of a bigger pattern, dealing with difficult video clients covers the people side.

Review_Cut_v4.mp4In Review
212160p · ProRes
00:34 / 02:18
SR
Sarah 0:34

Frame-accurate note, everyone sees the exact same thing.

In PlayPause, every comment is pinned to the exact frame, no more “which part?” email threads.

Export the SRT and bring it into your editor

Once the wording is agreed, the export is the easy bit, right. SRT export sits with AI transcription on the Agency plan and up, so you take the transcript out as a file, open it in a plain text editor, and work down the client's comments one by one. Every comment carries the timecode of its frame, so you find the block whose times wrap that moment and fix the word, and for a product name that appears forty times, one find and replace does the lot. Save it as plain text in UTF-8 with the .srt extension, and on a Mac switch TextEdit to plain text first, because stray formatting or odd encoding is the most common reason an import fails.

I want to be straight about the boundary here. PlayPause gives you the transcript, a place for the client's corrections and the SRT export, and it does not style your captions. Font, position, safe area, how lines pop on and off, the karaoke look, all of that happens in your editor, because the look of captions is a design decision while the words are a text decision.

In Premiere Pro you import the .srt like any media and drop it on the timeline, where it becomes a caption track you can style, and Premiere is also where the PlayPause panel lives, so comments show inside Premiere and clicking one jumps the playhead, which is lovely when a wording fix needs a small re-time. In DaVinci Resolve the file comes in through File, Import, Subtitle, and in Final Cut Pro through File, Import, Captions. For those two, and for After Effects, the workflow is export, upload, review, then work from the notes on a second screen. If you cut in Resolve, my notes on Resolve export settings for review links will save you a slow upload.

The caption file is a text decision and the caption look is a design decision, and mixing the two is why approvals stall.

And if you still put burned-in timecode on review copies out of habit, you probably do not need it once every comment carries its own frame, which I argued in when you can drop burned-in timecode.

What I do when the project is not on Agency yet

To be very honest, plenty of my smaller projects sit on Creator at nine dollars a month, and Creator does not include AI transcription or SRT export, so there I write the transcript myself or fix whatever the platform auto-generates, then build the file by hand using the block format above. What Creator does give you is what matters most for approvals, so frame-accurate comments, drawing on the frame, version stacks for MV1 through MV4, and who-watched analytics that tell you whether the approver has actually opened the link.

The moment a client wants a caption file as a proper deliverable, I move that workspace to Agency at nineteen dollars a month, which covers up to 50 members on one flat price and adds transcription, SRT export, side-by-side version compare and guest version upload. Pricing is per workspace and not per person, which I'm pretty sure surprises people most when they come from a per user tool. The breakdown is on the PlayPause pricing page, and the education version of this lives on AI captions for lecture capture teams.

Frequently asked questions

What is the difference between an SRT file and burned-in captions?

An SRT is a separate text file with timecodes, so the platform draws the words on top, the viewer can switch them off, and search engines and translation tools can read them. Burned-in captions are baked into the picture as pixels, so they cannot be turned off, translated or indexed. I deliver both quite often, right, the burned-in version for social feeds where people scroll on mute and the SRT for the client's own systems.

Can I create an SRT file from video without a transcription tool?

You can, you just type it in a plain text editor using the numbered block format, and on thirty second cuts that is honestly faster than uploading anything. Past about two minutes it stops being sensible, because you are hand entering start and end times for every line and one slip shifts everything after it, so machine transcription plus a human correction pass wins on time and on accuracy.

How accurate is AI transcription on accents or industry jargon?

Ordinary conversational speech comes back clean enough that my own correction pass stays short. Accents, crosstalk and internal jargon are where it drops, and it drops predictably, mangling names, codenames and acronyms far more often than ordinary sentences. That is why I audit proper nouns first, leave the general wording alone unless something reads wrong, and send the client specific questions rather than a whole transcript to proofread.

Do clients need an account to fix the caption wording?

No, and this is the bit I would never give up, right. They open one link in a browser with no install and no sign up, and they leave comments that stick to the exact frame. You can password protect that link and revoke it instantly if a project ends badly, and in my agency every extra login screen between a client and an approval tends to cost a day.

Where I would actually start with this

If you are handing a caption file to a client this week, run the next one this way, basically. Lock picture, upload the cut as a version, transcribe it there, do your own proper noun pass, send the one link, and export the SRT only once the wording has a yes against it. Every plan comes with a 7-day free trial, so you can run a real job through it, and PlayPause pricing shows which plan carries transcription and SRT export. Trust me on any level, the day you stop reconciling a document against a video is the day captions stop holding up delivery.

So yeah. That's my way of saying it.

SM
Saumyajit Maity
Co-founder, PlayPause

Saumyajit co-founded PlayPause after years watching review and approval quietly eat creative teams' deadlines. He writes about the workflow side of video, feedback, versioning, and getting to a clean sign-off.

Related resources

Keep reading

Bring your team into one review space

Centralize feedback, lock approvals, and deliver faster, start free today.

Sign Up for Free