How to Make SRT File for YouTube From Your Edit's Transcript
How I make an SRT file for YouTube from the transcript of the approved cut, proofread the names and jargon, export it, and upload clean captions in YouTube Studio.
The SRT file for ep_112_MV3.mp4 is the last thing standing between us and publish, so I have the approved cut open in one tab and YouTube Studio in the other. At 00:18:42 the guest says the name of her company, right, and the automatic captions have turned it into two random words, and with the thumbnail done and the sponsor read approved, the only thing left is a caption track that would make it look like we never listened to her.
So I do what we always do in my agency now, which is basically ignoring the auto captions, pulling the transcript from the approved cut, reading it once against the picture and exporting a proper file. This post is exactly that, how to make an SRT file for YouTube from your edit's transcript without typing a single timestamp, walked through the way we actually do it, mistakes included.
To be very honest, captions are the last thing anyone thinks about on a YouTube project, right, they come after the color, the mix and the thumbnail argument, so they get rushed. The catch here is that a viewer watching on mute reads every caption word for word, so a wrong name at 00:18:42 is the one thing they remember.
What an SRT file for YouTube captions actually is
An SRT file is basically a plain text file, the kind you open in Notepad or TextEdit, and it holds your captions as numbered blocks. Each block is a number, a timing line and one or two lines of whatever was said in that window, with an empty line before the next block starts. Block 214 from that episode, read top to bottom, is the number 214 on its own line, then 00:18:42,300, an arrow made of two hyphens and a greater-than sign, and 00:18:45,900, then the guest's sentence split across two short lines, then one empty line. That really is the whole format, right, which is why almost every tool reads and writes it.
The Wikipedia page on subtitles covers the history if you're curious, but for a channel all that matters is that YouTube accepts SRT directly as a timed caption file, so it never has to guess where a line goes.
Auto captions get the easy words right and the important words wrong, meaning the guest's name, the product, the city, the inside joke from episode 40 and so on. An uploaded SRT replaces that guesswork with a track somebody has actually read. I've written more about why this matters for reach in our captions and accessibility guide.
Numbered blocks, a start and end time, and one or two lines of dialogue. YouTube reads it directly, so every line lands where it was spoken.
How to make an SRT file for YouTube from a transcript
You can type the file by hand, which I did exactly once years ago for a four-minute video, and it ate most of an afternoon because every timestamp had to be found by scrubbing. If it's a sixty-second clip and you want to try it anyway, open Notepad, or TextEdit switched to plain text from the Format menu, type the blocks in the pattern above and save it as UTF-8 with a .srt extension, and it will work, it's just slow. YouTube Studio also lets you paste a plain transcript and auto-sync it, which in my experience is fine for a solo talking head and shaky with crosstalk or a music bed underneath.
What we do now is generate the file from the transcript of the cut we're already reviewing, because that transcript already knows where every word sits in time. In PlayPause this lives inside AI transcription, which comes with the Agency plan and up. Once the cut is uploaded, from the browser, the desktop app or straight from the timeline through the Premiere Pro panel, you run transcription and get a clickable transcript where clicking a line jumps the player to that moment. From there you export an SRT, and the timings come from the audio of that exact file, so nobody types or guesses a timestamp. If you've never set up transcription for a team, getting started with AI transcriptions covers the setup side.
For instance, on a 42-minute podcast episode my editor uploads MV3, the transcript comes back, and I read it while the video plays at normal speed, and on that one I caught the guest's company name and two product names that came out wrong before the SRT went anywhere. The typing part basically disappears, and what's left is the part that actually needs a human, which is reading. If your starting point is a raw video file rather than an edit already in review, I covered that in creating an SRT file from video.
Transcription is only as good as the audio, right, and accents, crosstalk and very fast talkers all make it work harder, which I've gone into in AI transcription accuracy with accents, so always plan for a proofreading pass.
Proofreading captions against the approved cut
This is the section I really really want you to take seriously, because the biggest caption mistake in my agency is captioning the wrong version, right. Somebody transcribes MV2 on Tuesday because it looked nearly done, the client asks for a four-second trim in the cold open on Wednesday, MV3 gets approved, and the SRT from MV2 goes to YouTube anyway. Every caption after that trim is now four seconds early, which on a talking-head video is a full sentence, so the viewer reads the next line before the person says it.
So our rule is simple, captions come from the approved MV and nothing else. It's basically version control for a caption file, the SRT is married to one specific export, and does that make sense, right, the moment the picture changes the file stops matching. In PlayPause new versions stack on the same card as MV1, MV2, MV3 and so on, so it's obvious which one is current, and side-by-side version compare on Agency shows you exactly where a trim happened. After an editor once uploaded the wrong cut on us, I wrote a post on fixing a wrong cut upload, and captions teach the same lesson, the file you caption has to be the file you publish.
Then I play the video at normal speed with the transcript beside it and read along, hunting for names of people, companies and products, numbers, and industry words the model has probably never heard. A finance guest says a fund name, a fitness guest says a supplement brand, and when I find one like that I leave a frame-accurate comment at that moment so the editor sees it too, because the same misspelling usually shows up in the lower thirds or the description. The sponsor read gets the same treatment, since it has to match the approved script, which I covered in reviewing sponsor segments before upload.
For recurring shows we keep a spelling list of guest names, products and the host's catchphrases inside a Playbook, which is basically a per-client document of creative style and direction, so the proofreading pass becomes checking the transcript against that list. For a 40-minute episode it's one sitting with a coffee, and trust me on any level, that sitting saves you from a comment section pointing out that you spelled your own guest's name wrong.
- Transcript made from the approved MV
- Guest and host names spelled right
- Brand and product names checked against the Playbook
- Numbers and prices match what was said
- Sponsor read wording matches the approved script
Frame-accurate note, everyone sees the exact same thing.
Exporting the SRT and uploading it in YouTube Studio
Once the read is done, I export the SRT and open it in a plain text editor to fix the names I noted, and since SRT is just text, a find and replace on the guest's name fixes every instance without touching a single timing. Save it with a .srt extension and UTF-8 encoding, and check that your computer didn't quietly add .txt to the filename, because that one has caught at least two of my editors.
In YouTube Studio, open Subtitles from the left menu, pick the video, add the language if it isn't there, click Add under subtitles, choose Upload file, pick With timing, select your SRT and publish. The button wording shifts every so often, so keep YouTube's help page on adding subtitles open the first time you upload subtitles to YouTube, right, it's the page I send every new editor to.
After it's published, I watch the first minute, one random spot in the middle and the last minute with captions on. If the start is right and the middle is off, that's the version problem from the last section, and if every line is off by the same amount from the very first caption, somebody added or removed an intro after the transcript was made. I'm pretty sure this three-spot check has saved us more awkward client emails than any other habit we have.
an afternoon of scrubbing for timestamps and a track that drifts after every trim
timings straight from the audio, one proofreading pass, and one clean file for YouTube
Common SRT mistakes that break YouTube captions
Version drift is the big one, and after that it's formatting, right, where someone opens the file in a word processor and on save it adds smart quotes, hidden formatting or a different encoding. YouTube then either refuses the file or shows strange characters wherever there was an accent or apostrophe.
Then there's the decimal point, because SRT uses a comma before the milliseconds and a similar format called WebVTT uses a period, so a hand conversion or timings copied from a spreadsheet give you a file that looks fine to a human and fails to load. A missing empty line between two blocks makes YouTube glue them together or skip one. Overlap, where one block's end time runs past the next block's start, usually comes from hand-edited timings and stacks two captions on top of each other.
Line length is more of a quality problem than a technical one, because on a phone a thirty-word caption fills half the screen and covers the speaker's face, so I keep blocks to one or two short lines and split any monster block in the text editor. At the end of the day, if you never touch the timings and only fix words, you avoid almost all of this, which is the whole point of exporting from the transcript.
Caption the version that actually goes live, because a track made from an older cut starts drifting the moment one sentence moves.
Captions for Shorts and podcast clips
Shorts are where the version problem comes back in a sneakier form, right, because a Short cut from the long episode has its own timeline starting at zero, so the long video's SRT is useless for it. We treat every clip as its own deliverable with its own approved MV, transcript and SRT, and the proofreading is quick because the long episode already taught us every tricky word.
For a lot of Shorts we also burn in big word-by-word captions for style, and I still upload an SRT, because the burned-in text is for the look while the SRT gives viewers a track they can switch on and off. Streams work the same way, and if you're cutting highlights from a Twitch VOD, the stream highlights editor workflow post shows where caption files fit.
Podcast teams also have speakers to think about, so with two or three people talking over each other I add a short speaker label when the speaker changes and it isn't obvious on screen. You see what I mean here, the transcript ends up doing more than captions, and our pages on video review for podcast teams and searchable transcripts for podcast producers go deeper on that.
Frequently asked questions
Can I just rely on YouTube's automatic captions instead of an SRT?
You can, and for a quick vlog with simple vocabulary they're often fine, right. The problem is names, brands and jargon, which are exactly the words auto captions miss and exactly the words your audience cares about. An uploaded SRT that you've proofread against the approved cut replaces the guesswork, and it gives you a reusable file for clips, translations and show notes. For a channel that's building a brand, I'd always upload my own track.
Which PlayPause plan do I need to export an SRT?
AI transcription, with the clickable transcript and the SRT export, sits on Agency and Enterprise, so if captions are part of what you deliver, Agency is where I'd start, because it also gives you side-by-side version compare, which is exactly how you catch the trim that would have thrown your captions off. Creator still covers frame-accurate comments and version stacks, it just doesn't transcribe, and every plan starts with a 7-day free trial, so you can test the export on a real episode first.
Should my SRT include speaker names?
For solo videos I leave them out, because they just clutter the screen. For podcasts and interviews with two or more people, a short label when the speaker changes helps a lot, especially when the camera isn't on the person talking. Keep it consistent, use the same spelling from your Playbook every time, and only label the handovers rather than every single block, so the captions stay readable on a phone.
Can I reuse the long video's SRT for a Short cut from it?
Not directly, because a Short has its own timeline starting at zero and the timings in the long video's SRT belong to a different file. Treat the Short as its own deliverable, upload it, get it approved as its own version, and make a fresh transcript and SRT from that exact file. The proofreading goes much faster the second time around because you already know every tricky name in the episode.
If you want to try this, run the approved cut of your latest episode through AI transcription, proofread it once and export the SRT. Every plan on the PlayPause pricing page comes with a 7-day free trial, Agency is the one with transcription and SRT export, and the video review page for YouTube creators shows how the rest fits together.
So yeah. That's my way of saying it.
Saumyajit co-founded PlayPause after years watching review and approval quietly eat creative teams' deadlines. He writes about the workflow side of video, feedback, versioning, and getting to a clean sign-off.
Related resources
Keep reading
Bring your team into one review space
Centralize feedback, lock approvals, and deliver faster, start free today.
Sign Up for Free