How to extract subtitles from a YouTube video

Four free ways to get subtitles out of a YouTube video, which one fits your job, how accurate the text is, and what to do when there is no track.

Faisal AshfaqMaintains the extractor
Published 12 June 2026Updated 2 October 202621 min read
Free YouTube subtitle extractor: caption strips, ink pot and quill on paper
A subtitle extractor reads the caption track a video already carries. It does not listen to the audio.
What you need to know
  • 01A YouTube subtitle extractor copies the caption track a video already has. It does not listen to the audio, so every honest tool returns the same words.
  • 02Four routes work today: a paste-a-link tool, YouTube's own transcript panel, the file from YouTube Studio or the API, and running speech recognition yourself.
  • 03If a video has no caption track, no extractor can produce one. An empty result is the honest answer.

To extract subtitles from a YouTube video, paste the video link into a free subtitle extractor, choose SRT or VTT, and download the file. For most public videos that takes seconds and needs no account. It only works when the video already has a caption track, because an extractor copies that track. It does not listen to the audio.

The confusing part is that four different things all get called a free YouTube subtitle extractor, and they behave nothing alike. One copies a file that already exists. One shows you text on screen and makes you drag your mouse over it. One is a developer endpoint that mostly refuses strangers. One transcribes the audio from scratch and charges you in compute time. Pick the wrong one for your job and you either lose an hour or you get a file with no timestamps in it.

This page covers all four routes as they behave in October 2026, how to judge one tool against another, how accurate the text really is, how to read and repair the file once you hold it, and the law that put so many captions on YouTube in the first place. One disclosure first: YouTubeScribe is our product, so treat the sections about it as the vendor talking. The limits we list are the real ones, including the videos we cannot touch.

What a subtitle extractor actually does

A YouTube video and its captions are two separate things. The picture and sound live in one stream. The captions live in a timed text file that the player fetches on the side and paints over the frame when you press CC. Nothing about the words is burned into the pixels. So getting subtitles out of a YouTube video is not image processing and it is not audio processing. It is asking for a file that is already there, and keeping a copy.

That single fact explains almost every behaviour people find odd about caption downloaders. Extraction is fast because it is a file transfer, not a computation. It can be free because no model runs. No tool can beat another on accuracy for the same video, because everyone is copying the identical source. And some videos return nothing at all, whichever of the twenty tools you try.

Extractor, downloader, transcriber

These three words get used as if they mean the same thing. Underneath, they describe two jobs.

  • An extractor reads a timed text track a video already has and copies it out as a file.
  • A downloader is the same job under another name. People searching for a YouTube subtitle downloader and a YouTube subtitle extractor want the identical result.
  • A transcriber runs speech recognition on the audio and writes new text. It works when no track exists, but it costs compute time, and it can invent plausible words for audio it misheard.

Naming the job before you search saves more time than any tool choice. If your video has no track, ten extractors will all fail you in the same way and only a transcriber will help. If it does have a track, a slow paid transcription service is the wrong tool.

Two kinds of caption track

YouTube holds two kinds of track and an extractor will take either one.

  • Creator uploaded tracks. Somebody wrote or corrected these and added them to the video. They carry real punctuation, speaker labels where the author bothered, correct spellings of names and products, and sensible line breaks. This is the good stuff.
  • Automatic captions. YouTube's speech recognition made these after the video was uploaded. They often lack punctuation and speaker labels, proper nouns are often mangled, and the line breaks fall at roughly even intervals rather than at the end of a clause.

You can usually tell them apart in a second. Open the caption menu in the player. A track labelled with a language name and the words auto-generated after it is machine output. A bare language name is human work. That difference decides how much editing you are about to do.

Automatic captions do not cover every language. YouTube's help page on automatic captioning lists 67 languages, running alphabetically from Afrikaans to Zulu. That number is smaller than most people assume, and it is the reason a perfectly clear video in a less common language sometimes has no track at all.

Extraction is not transcription

ExtractionTranscription
InputThe caption file YouTube already storesThe audio track
Speed for 20 minutesSecondsMinutes, sometimes longer than real time on weak hardware
CostUsually freeCompute time or an API bill per minute
Accuracy ceilingExactly as good as the existing trackDepends on the model, can beat a bad auto track
Works with no caption trackNoYes
TimestampsInherited from the original trackGenerated, quality varies by model

YouTubeScribe sits firmly on the extraction side. It reads the caption track a public or unlisted video already has. It never invents speech and it does not run speech recognition on the audio. If YouTube never wrote a track, the answer you get back is nothing, and that is the correct answer rather than a bug.

A miss is a valid result

Some tools respond to a track-less video by producing plausible sentences anyway. That output has no tie to what was said. If you are quoting a source, citing a lecture or building a dataset, a blank result you can verify beats a paragraph you cannot.

What a caption cue looks like inside

Open any SRT file in a text editor and the structure is obvious within five seconds. Each cue has four parts: a sequence number on its own line, a timing line with a start time, an arrow and an end time, one or two lines of text, then a blank line before the next cue. Times run hours, minutes, seconds, then a comma and three digits of milliseconds.

WebVTT holds the same information with two differences that trip people up. The file must begin with the word WEBVTT on the first line, and the milliseconds separator is a full stop rather than a comma. Sequence numbers are optional. For ordinary subtitle work that is most of the difference, which is why converting between the two is a find and replace rather than a rebuild.

Once you can read a cue, you can audit any extractor's output in about thirty seconds. Check that the first cue starts near the first spoken word. Check that the last cue lands near the end of the video. Scan the middle for a gap of several minutes, which usually means the track itself is partial rather than the download failing.

The four ways to get subtitles out of a video

Here is the honest comparison. Times are for a single 20-minute talk with an existing English caption track, measured start to finish, including the fiddly bits everybody forgets to count, like opening a new tab and cleaning up a paste.

MethodTime for a 20-minute videoTimestampsBest forCost
Paste a link into YouTubeScribeAbout 10 secondsYes, SRT and VTT keep themAnyone who wants a file, one video at a timeFree, no account
YouTube's own transcript panel3 to 8 minutes with the copying and tidyingInline text only, not a real cue fileA quick read or one quote, no third partyFree
Download the caption file from Studio or the API5 to 20 minutes, longer the first timeYes, the original fileVideos you own or can editFree, but needs a signed-in account with edit rights
Run speech recognition yourself5 to 40 minutes depending on hardwareYes, generated freshVideos with no caption track at allCompute time, or per-minute API charges

Method 1: paste the link into a free extractor

The fastest route by a wide margin. You copy the video URL, paste it, choose a format, and a file lands in your downloads folder. There is nothing to install, nothing to authorise, and no account to create.

What makes one tool better than another here is not accuracy, since everybody reads the same source track. It is the boring operational stuff: which track it picks when there are several, whether SRT costs money, whether your links are logged, and whether the tool admits it found nothing instead of inventing a transcript.

On YouTubeScribe the job is split across free pages. Our free YouTube subtitle downloader returns the caption file as SRT or VTT, with the timings left as YouTube has them. The Transcript Generator and the home page return readable text you can copy or save as TXT. We read 125 languages where a track exists. That is not the same claim as YouTube's 67 automatic caption languages: those are the languages the machine will listen in, while ours counts every language a track can be in once one exists, including human translations.

Method 2: YouTube's own transcript panel

YouTube ships a reading view for captions and most people have never opened it. On a computer, open the description under the video and click Show transcript. YouTube Help gives the same step for the Android and iPhone apps. A panel opens with the caption text, each line beside its start time, and clicking any line jumps the video to that moment.

It is genuinely useful for reading. For finding the one sentence you half remember, it beats everything else because you can scan and jump in the same motion.

The problem is getting the text out. There is no download button. You select the text with the mouse, which on a long video means a scroll and drag inside a small panel, and what lands on your clipboard is a block of text with the timestamps mixed into it. It is not a cue file. Nothing that expects SRT will accept it. For one quote it is fine. For a 90-minute podcast it is a long, dull job.

  • Good for: one quote, a fast skim, checking whether a track exists before you bother with anything else.
  • Bad for: any file you need to feed to an editor, a translator, a search index or a script.
  • Watch out for: the panel showing an automatic track when a better human track exists. Check the caption menu in the player for other tracks.

Method 3: download the caption file from Studio or the API

There are two versions of this and people mix them up constantly.

The first is YouTube Studio. If the video is on your own channel, open Studio, go to Subtitles, pick the video, open the track and use the download option. You get the real file, exactly as stored. This is the cleanest source there is. It is also useless for anybody else's video, which is the case most people are actually in.

The second is the YouTube Data API. It has a captions resource with a list method that names the tracks on a video, and a download method that returns the track itself. The download method is the one that stops people. Google's reference says it needs OAuth authorisation and that the signed-in user must have permission to edit the video. Each call also costs 200 units of your daily API quota. It was built for a creator or their caption vendor to pull their own files, not for anybody to collect anybody's.

Budget an afternoon for the API route the first time: a Google Cloud project, the API switched on, an OAuth consent screen, a client secret, a token exchange, then the two calls. If you own the videos and need this to run nightly on a hundred uploads, it is the right tool. For a handful of videos you do not own, it is the wrong shape entirely.

Method 4: run speech recognition yourself

The last resort, and the only route that works when no caption track exists. You pull the audio and run it through a speech model, locally or through a paid API, and it writes fresh cues with fresh timings.

It is the only method here that can beat the original: a good modern model on clean audio will often out-punctuate and out-spell a weak automatic track. It is also the only method that can be badly wrong in ways you will not notice, because a confident model produces fluent text for mumbled audio and you have no source file to check it against.

  • Use it when: the video has no track at all, the automatic track is unusable, or you need a language YouTube's automatic captioning does not cover.
  • Skip it when: a human-uploaded track already exists, because you will spend forty minutes producing something worse than a free ten-second download.
  • Always check: the first and last minute, and any section with music or crosstalk. That is where invented text hides.
Extracting YouTube subtitles from a video link into SRT and TXT caption files with timestamps
Extraction copies the caption track YouTube already stores, then writes it out as SRT, VTT or plain text.

Step by step with a free extractor

This is the paste-a-link route, written out in full so you can see where the seconds go. The tool is the Subtitle Downloader on YouTubeScribe, but the shape of the process is the same on any honest extractor.

  1. Copy the video URL. The share button gives you a short youtu.be link and the address bar gives you the long watch link. Both work, and so does a Shorts link. A timestamp on the end is ignored.
  2. Open the Subtitle Downloader and paste the link into the field.
  3. Run it. The page does not ask for a language. It takes the creator's English track first, then the automatic English one, then the first track the video lists. If you need another language, the home page has a language picker.
  4. Choose SRT or VTT. SRT is selected by default and the preview below shows the file as it will be saved.
  5. Click Download. The file is named after the video.
  6. Check the file before you build on it. Confirm the first cue lines up with the first words spoken, and scroll to the end to confirm it did not stop early.

Step six is the one everybody skips, and it is the one that saves you. A partial track looks exactly like a complete one until you reach the end of it.

Choosing the output format

Match the format to what happens next, not to what sounds most technical.

  • TXT when a person is going to read it, a language model is going to summarise it, or you want to paste it into a document. Get it from the Transcript Generator or the home page, both free.
  • SRT when the text is going back onto video, into a subtitle editor, into a translation workflow, or into anything that needs to know when each line appears. It is the most widely accepted subtitle format. Free on the Subtitle Downloader.
  • VTT when the destination is a web player using the HTML track element, or a platform that asks for WebVTT. Free on the Subtitle Downloader.

One habit worth forming: keep the SRT even when you only wanted the text. Going from SRT to TXT is easy at any point later. Going from TXT back to SRT means re-timing every line by hand, and nobody has ever enjoyed that.

More than one video at a time

One video is easy. Forty is where the work changes. If you are pulling a whole course, a conference series or a channel's back catalogue, you need a list first and then one file per video.

The Playlist Exporter gives you that list. It reads a playlist and returns the running order as a CSV: position, title, length, video ID and link. It does not include caption text. You then run each link through the Subtitle Downloader, or, if you write code, call the transcript endpoint of the YouTubeScribe API once per video.

Expect misses in any batch. A playlist of sixty will usually contain a few videos that come back empty, because a members-only upload slipped in, a very short clip never got a track, or a re-upload is still processing. That is normal. Note them and move on rather than assuming the run failed.

The limits, stated plainly

Since this is our tool, here is what it will not do, in the same place as what it will.

  • Private, members-only and age-restricted videos return nothing. There is no track we are permitted to read, and no workaround is being held back.
  • It does not run speech recognition. No track means no output.
  • It cannot improve a bad automatic track. Garbled audio in, garbled captions out, because the garbling happened at YouTube when the track was made.
  • No account is needed, so nothing is saved to one. The file you download is the copy to keep.
  • The Subtitle Downloader gives SRT and VTT free. On the home page TXT, SRT and PDF are free, while VTT, Markdown and JSON sit on the Pro plan.

The single most common support message we get is that the tool is broken because it returned nothing. In almost every case the video genuinely has no caption track. The fix is not a different extractor, it is speech recognition or a different source video.

YouTubeScribe support notes, June to August 2026

How to judge an extractor before you paste a link

Free extractors do not compete on accuracy, because an honest tool of any kind reads the same file. They compete on what happens to your link and your data while they work. That depends mostly on what kind of tool it is.

  • A web tool is a page you visit. You paste a link, a server fetches the track and hands you a file. No install and usually no account. This is the shape YouTubeScribe takes. The trade-off at scale is server capacity: a free tool may slow down or refuse very large batches.
  • A browser extension adds a button to the YouTube page itself. Convenient, but an extension that can read a video page can often read every page you visit unless its permissions are narrow. The install prompt tells you which, if you read it before accepting.
  • An API script is code calling the YouTube Data API. The captions download method needs OAuth and edit permission on the video, so it suits a developer working with their own channel.
  • A desktop app runs on your own machine with no server in the loop. Nothing you paste leaves your computer, which is the strongest privacy story of the four, provided you trust the publisher enough to run their program.

Before you paste anything into any of them, run through a short checklist.

  • Does it ask for your YouTube password or a Google sign-in? A public video should never need your credentials.
  • Does it say whether it logs the links you paste? Silence on this question is itself an answer.
  • Does it keep a copy of the output after handing it to you?
  • For an extension, what does the install prompt list? Read and change data on every site you visit is a broad grant for a job that needs one video page.
  • For an API script, whose credentials are doing the asking: yours, or a shared key on somebody else's server?
  • For a desktop app, is the publisher identifiable and the program signed?
  • Does the tool say plainly when a video has no caption track? A tool that returns a transcript for every link is probably inventing rather than reading.
TypeWhere it runsWhat it needs from youTrust question to ask
Web toolThe provider's serverUsually nothing but the linkDoes it log the link or keep the output?
Browser extensionInside your browserAn install and a permissions grantWhat can it read on pages that are not YouTube?
API scriptWherever you host itA Google Cloud project, OAuth, edit rightsWhose credentials are doing the asking?
Desktop appYour own machineA download and an installIs the publisher known and the program signed?

For a single public video, most of this barely matters: paste the link into any tool that does not ask for your password, and move on. The checklist earns its keep once the job repeats, such as a nightly script or a tool the whole team standardises on.

How accurate the subtitles are

The track decides the accuracy before any extractor is involved. A creator-uploaded track is as good as the person who wrote it. An automatic track is as good as YouTube's speech recognition was on that audio. Ten extractors pointed at the same automatic track return ten copies of the same errors, because none of them are writing anything.

Accuracy is also not one number. A track can get most words right and still change the meaning of a sentence. Punctuation and speaker labels are a third axis: every word can be right and the text still be hard to use as a formal quote. Any accuracy claim worth trusting says which of these it measured.

Few studies have counted rather than estimated. One that did is Becky Sue Parton's 2016 paper in the Journal of Open, Flexible and Distance Learning, which asked whether YouTube's auto-generated captions met deaf students' needs on course video. Across 68 minutes of auto-captioned video it recorded 525 phrase-level errors, an average of 7.7 a minute.

A phrase-level error is not a stray typo. It is a mistake that changes or removes part of the meaning: a wrong word standing in for the right one, a dropped clause, a name rendered as something else. For anyone relying on the text as the record of what was said, that rate is something to plan around.

A 2016 study, read honestly

Speech recognition has improved since 2016, so treat 7.7 errors a minute as evidence of how automatic tracks fail, not as today's figure for YouTube. What has not changed is the shape of the problem: an automatic track can be wrong at the level of meaning, and nothing downstream can repair that by copying it.

This is why comparing extractors on accuracy is close to meaningless. If two tools return different words for the same video, one of them is probably not reading the real track. It may be running its own speech recognition and calling it extraction, or serving an older copy from before the creator corrected the track. Different line breaks are normal. Different words are a warning sign.

What paying for a tool actually buys

If accuracy is fixed by the track, a paid tier cannot sell better words. What money usually buys is volume, more formats, batch handling or support. On YouTubeScribe the subtitle and transcript tools are free and return the same words any paid run would. Credits go on AI tools that do something with the text, such as a summary or Chat With Video. Any tool that implies a paid tier reads captions more accurately is selling something that is not for sale.

SRT, VTT, TXT: picking and fixing the file

You have a file. Now the questions change: will the thing you are feeding it to accept it, and if not, what do you edit.

SubRip, the default answer

SRT is the plain cotton of subtitle formats. Sequence number, timing line, text, blank line, repeat. It carries no positioning and no styling, and that poverty is exactly why everything accepts it: there is nothing in the file to misread. Video editors, translation platforms, social schedulers and almost every upload form take SRT.

SRT was never standardised by a standards body. It comes from SubRip, a ripping program from the early 2000s, and spread because everyone copied what already worked. YouTube's upload documentation describes it as plain UTF-8 text with no styling markup. If you open an SRT and see styling tags, some tool has been creative and a strict parser may object.

WebVTT, the web native one

WebVTT is the W3C format that browsers understand natively through the HTML track element. Same cue idea as SRT, plus cue settings for position, alignment and line placement, plus a proper header. YouTube accepts it on upload and supports its positioning, though it limits styling to bold, italic and underline.

Convert SRT to VTT by hand in three moves: add a line reading WEBVTT and a blank line at the top, change every comma in the timing lines to a full stop, and save as UTF-8 without a byte order mark. Going the other way, strip the header and swap the separator back. Cue numbers are optional in VTT, so keep or drop them.

Plain text, and what it costs you

TXT is a caption file with the timing thrown away. That is a real loss and a real gain. Prose reads like prose, so summarising, searching, quoting and pasting all get easier, and a language model wastes no attention on timecodes. But you can never put it back on video without re-timing every line.

Rule of thumb: if a human or a model is the final reader, take TXT. If a player is the final reader, take SRT or VTT. If you are not sure, take SRT, because it converts down to TXT in one step whenever you decide.

The heavier formats: TTML and EBU-TT

Timed Text Markup Language 2, a W3C Recommendation published on 8 November 2018, is XML rather than plain lines. It carries font, colour, regions and metadata that SRT and VTT have no room for, so it lives in professional captioning pipelines. EBU-TT Part 1, set out in the European Broadcasting Union's TECH 3350, is a TTML profile narrowed to what European broadcasters need.

Recognising the family is useful even if you never touch broadcast work. A plain-line file with a comma in its timestamps is SRT. The same shape with WEBVTT on the first line and a full stop is WebVTT. Angle brackets and a tt namespace mean TTML or EBU-TT, and that file will not open cleanly in a tool built only for the first two.

What YouTube accepts if you put captions back

Extraction is often only half the job. People pull a track, fix the spellings, translate it and upload it again. YouTube's supported files list is wider than most people expect, and worth knowing before you convert something unnecessarily.

FormatExtensionPositioning and styling
SubRip.srtNone, plain UTF-8 text with no styling markup
WebVTT.vttPositioning supported, styling limited to bold, italic and underline
TTML and DFXP.ttml, .dfxpStyling and positioning both supported
Scenarist Closed Caption.sccBroadcast format, YouTube's preferred choice for that class
SAMI and RealText.smi, .rtLegacy formats, still accepted
SubViewer.sbv, .subBasic timing and text
MPsub, LRC, Videotron Lambda.mpsub, .lrc, .capBasic timing and text
EBU-STL and CEA-608 family.stl and relatedBroadcast delivery formats

The practical reading of that table: if your file is already in one of those shapes, upload it as it is. Converting a TTML file down to SRT to be safe throws away positioning that YouTube would have kept. If you are writing from nothing, SRT stays the sensible default because every other tool in your chain will also accept it.

Five file problems and their fixes

  1. Accented characters appear as question marks or garbage. The file is not UTF-8. Save it again as UTF-8 and the problem disappears.
  2. A strict parser rejects a valid looking VTT. Check for a byte order mark before the WEBVTT header. Save without one.
  3. Every subtitle shows a fixed amount early or late. That is a constant offset, not a broken file. Any subtitle editor will shift the whole track by a set number of milliseconds in one action.
  4. An upload form refuses an SRT. Look for styling tags or HTML inside the cues. SRT is meant to be plain text and some parsers enforce it.
  5. Cues show on one line where you expected two. Line breaks inside a cue are part of the text and survive conversion, so if they vanished, the tool that wrote the file flattened them.

Illustrative example

A localisation team pulls SRT for 40 conference talks, fixes speaker names and product spellings once with a find and replace across all 40 files, then hands the corrected set to translators. Extraction takes minutes, the correction pass an afternoon, and no one re-times anything. These figures are illustrative, not a measured client result.

When there is nothing to extract

This is the hardest section to write honestly, because the answer is sometimes that you cannot have the thing you want. No tool on the internet can copy a file that was never written.

Google's own list of reasons

YouTube documents six reasons automatic captions may be unavailable on a video. Each one is a dead end for an extractor, not a limit of the tool you are using.

  • The captions are not available yet because the audio is complex and still processing.
  • Automatic captions do not support the language spoken in the video.
  • The video is too long.
  • The sound quality is poor, or the speech is not recognised.
  • There is a long period of silence at the beginning of the video.
  • Several people are speaking over each other, or several languages are spoken at once.

Two of those deserve a second look. The processing one is temporary: a video uploaded in the last hour may simply not have a track yet, and coming back tomorrow fixes it. The silence one catches a surprising number of creators, because a long music-only intro can be enough for the recogniser to give up before anyone speaks.

The access wall

Separate from whether a track exists is whether you are allowed near it. Public and unlisted videos are readable. Private, members-only and age-restricted videos are not, because each sits behind a check that an extractor has no business defeating. Knowing the link does not change that.

If you have a legitimate reason to need a private video's captions, ask the channel owner. They can download the file from YouTube Studio in about a minute and send it to you. That is a faster conversation than any amount of tool hunting.

The language gap

Two numbers matter here and they measure different things. YouTube's automatic captioning covers 67 languages. YouTubeScribe reads 125 languages where a track exists. The second is larger because it includes every human-uploaded and human-translated track, which can be in a language the recogniser has never supported.

So a video in a language outside the automatic list only has captions if a person made them. When creators in those languages do upload their own tracks, extraction works well and the quality is usually better than any automatic track.

Dubbing is not captioning

YouTube's automatic dubbing translates and re-voices a video into other languages. YouTube Help describes the result as translated audio tracks that viewers can switch between. It does not describe any caption file. So a video dubbed into Spanish does not, for that reason alone, have Spanish subtitles to extract.

Check the two lists separately. The caption menu lists text tracks. The audio track setting lists dubbed languages. They are rarely the same list, and assuming they match is where most disappointment in this area comes from.

A five-minute triage

Before you conclude a tool is broken, run this. It takes less time than writing a support ticket.

  1. Open the video and look at the CC button in the player. Missing or greyed out means there is no track and nothing to extract. Stop here.
  2. Open the caption menu and read the language list. Note whether the entry says auto-generated, and whether a better human track sits beside it.
  3. Open the description and click Show transcript. If the panel opens with text, a track exists and any extraction failure is a tool problem worth reporting.
  4. Check the upload date. Uploaded in the last few hours means the track may still be processing. Try again tomorrow.
  5. Check the access state. A padlock, a members badge, an age check or a video that only plays when you are signed in means the access wall, not a missing file.
  6. Only after all five, reach for speech recognition. If the video truly has no track, that is the one route left.

Captions, the law and what you may do with them

Captions did not become common because every creator decided accessibility mattered. They became common partly because law started requiring them, first for broadcasters, then more and more for the web. None of these rules mention extraction tools. They govern whoever publishes the video. What they explain is why so many videos have a track for an extractor to read.

Almost every rule points back to the Web Content Accessibility Guidelines. Success criterion 1.2.2 asks for captions on prerecorded video at Level A. Criterion 1.2.4 asks for captions on live video at Level AA. Most laws that cite WCAG ask for Level AA, which brings both into scope.

RuleWho it coversWhat it asks forStatus
WCAG 1.2.2 and 1.2.4A reference standard, not a lawCaptions for prerecorded (A) and live (AA) videoW3C guideline
ADA Title II web ruleUS state and local government sites and appsWCAG 2.1 AA, which includes captions for videoPublished 24 April 2024. Compliance from 26 April 2027 for entities serving 50,000 people or more, 26 April 2028 for smaller ones
Section 508US federal agency websites and technologyCaptions and transcripts for multimediaIn force
FCC rules under the CVAAUS TV programming later delivered onlineCaptions carried over from broadcast to internet deliveryIn force
UK accessibility regulations 2018UK public sector websites and appsWCAG 2.1 AA, including the caption criteriaIn force
European Accessibility ActProducts and services across the EUIts definitions name subtitles for the deaf and hard of hearingApplies from 28 June 2025

The ADA dates moved. The Department of Justice published the rule on 24 April 2024 with earlier deadlines, then extended them through an interim final rule published on 20 April 2026. Pages that still say the rule took effect in 2024 are out of date.

Every rule above attaches to a publisher: a council, a federal agency, a broadcaster, an EU-regulated service. None reaches an individual creator uploading from a bedroom, and YouTube does not require captions before a video goes live. That gap is why automatic captioning exists, and why caption quality across the platform is so uneven. A university or a public broadcaster usually carries a proper track because someone upstream had to make one.

The audience is large. RNID estimates that over 18 million people in the UK are deaf, have hearing loss or tinnitus. Its Subtitle It! campaign found that 80% of regulated UK on-demand services offered no subtitles when it launched in 2015. By 2024, 94.7% of the providers reporting did.

Extraction is not a licence

Getting a caption file is a technical act. Using it is a legal one. Quoting, studying, indexing and accessibility work sit comfortably within normal use. Republishing somebody's full transcript as your own content does not. YouTube's terms and the API terms govern the platform side, and both are short enough to read. This page is not legal advice.

The summary is short. For a public video with an existing track, paste the link into a free extractor and you have a real SRT in seconds. For a fast quote, YouTube's own transcript panel costs nothing. For your own uploads at scale, Studio or the API gives you the original files. And for a video with no track at all, speech recognition is the only honest answer, with all the cost and caution that implies.

Questions people ask

What is a free YouTube subtitle extractor?

It is a tool that copies the caption track a YouTube video already has and gives it to you as a file. The video and its captions are stored separately, so the extractor requests the timed text the player uses and saves it as SRT, VTT or plain text. Because it is a file transfer rather than a computation, it is fast and can be free. It does not listen to the audio and it cannot produce captions for a video that never had a track.

How do I download subtitles from a YouTube video for free?

Copy the video link, open the Subtitle Downloader on YouTubeScribe, paste it in and run it. Choose SRT or VTT, check the preview, and click Download. There is no account and no charge. The page takes the creator's English track first, then the automatic English one, then the first track listed; for another language, use the picker on the home page. Then open the file and check the last cue lines up with the end of the video.

What is the difference between an extractor, a downloader and a transcriber?

An extractor reads a caption track a video already has and copies it out as a file. A downloader is the same job under another common name. A transcriber is different in kind: it runs speech recognition on the audio and writes new text, so it works when no track exists, but it can produce fluent, wrong text for audio it misheard. Extraction can only return what is really there, including nothing.

Is extracting subtitles from YouTube legal?

Reading a caption track from a public video and keeping a copy for quoting, study, indexing or accessibility sits within ordinary use in most places. What changes the picture is what you do next. Republishing a full transcript as your own content, or building a commercial product on somebody's words without permission, is a rights question rather than a technical one. YouTube's terms of service and the YouTube API Services terms govern the platform side. This is not legal advice.

Are YouTube captions copyrighted?

Usually, yes, as part of the work they belong to. Rights in captions generally follow the video unless the creator has licensed the text separately. That does not make reading or quoting them unlawful: quotation, study, indexing and accessibility uses are widely accepted. Copying a full transcript and republishing it as your own writing is closer to reproducing an article than citing one. When in doubt, quote a portion and link to the video.

What is the difference between subtitles and captions?

Subtitles traditionally carry the spoken dialogue, often translated, and assume the viewer can hear. Captions are written for viewers who cannot hear the audio, so they add speaker labels and sound descriptions. YouTube uses the two words almost interchangeably, and its upload documentation calls the same files subtitle and caption files. For extraction it rarely matters: you get whatever text is in the track.

How many languages does YouTube support for automatic captions?

YouTube's help page on automatic captioning lists 67 languages, from Afrikaans to Zulu, checked on 2 October 2026. That is the list the speech recogniser will listen in, so a video spoken in another language will not get an automatic track. It is a different number from how many languages an extractor can read. YouTubeScribe reads 125 languages where a track exists, because people upload and translate tracks in languages the recogniser does not support.

Does a subtitle extractor work on any YouTube video?

No. It works on public and unlisted videos that already carry a caption track. Private, members-only and age-restricted videos return nothing because they sit behind an access check, and having the link does not change that. Videos with no track return nothing because there is nothing to copy. If the CC button in the player is lit, extraction should work. If it is missing or greyed out, no extractor will help.

Do I need to sign in to extract subtitles from a YouTube video?

Not for a public or unlisted video on a paste-a-link tool; the link alone is enough, and YouTubeScribe's free subtitle and transcript tools need no account. Sign-in only matters if you use the YouTube Data API's captions download method, which needs OAuth and permission to edit the video, so in practice it works on your own channel.

How do I see the transcript inside YouTube itself?

On a computer, open the description under the video and click Show transcript. YouTube Help gives the same step for the Android and iPhone apps. A panel opens with the caption text and a start time on each line, and clicking a line jumps the video to that moment. There is no download button, so getting the text out means selecting and copying it by hand.

Should I choose SRT, VTT or TXT?

Choose by destination. SRT if the text is going back onto video, into a subtitle editor or into a translation workflow. VTT if a web player using the HTML track element is the destination. TXT if a person or a language model is the final reader. On YouTubeScribe, the Subtitle Downloader gives SRT and VTT, and the Transcript Generator or home page gives TXT. When unsure, take SRT, because converting it to TXT later is one step.

Can I extract subtitles from a whole playlist at once?

Not in one click on YouTubeScribe. The Playlist Exporter returns the running order of a playlist as a CSV, with position, title, length, video ID and link, but no caption text. Run each link through the Subtitle Downloader, or call the API's transcript endpoint once per video if you write code. Expect a few misses in a large batch, such as members-only uploads or videos still processing.

Can I upload an edited subtitle file back to YouTube?

Yes, and the accepted list is wider than most people expect. YouTube takes SubRip .srt as plain UTF-8 with no styling, WebVTT .vtt with positioning and limited styling, TTML and DFXP with full styling and positioning, SCC as its preferred broadcast format, plus SAMI, RealText, SBV and SUB, MPsub, LRC, Videotron Lambda .cap, EBU-STL and several CEA-608 formats. If your file is already one of those, upload it as it is.

Why does the extractor return nothing for my video?

Almost always because the video has no caption track. YouTube lists six reasons automatic captions may be missing: the audio is still processing, the language is not supported, the video is too long, the sound is poor or the speech is not recognised, there is a long silence at the start, or several people or languages overlap. Separately, private, members-only and age-restricted videos return nothing because of access. Check the CC button first.

Does automatic dubbing count as subtitles?

No. YouTube Help describes automatic dubbing as translated audio tracks that viewers can switch between, and it does not describe any caption file. So a video dubbed into a language does not, for that reason, have subtitles in it. A text track in that language exists only if a creator uploaded one or YouTube translated an existing track.

Is it safe to paste a YouTube link into a free extractor?

For a public video, the link itself reveals nothing sensitive, since anyone can open it. The real question is what the tool does next: whether it logs the link, whether it asks for credentials a public link should not need, and whether it keeps a copy of the output. A tool that only asks for the link and gives you a file is behaving as it should. Extensions and desktop apps carry different risks, worth judging on their permissions.

How do I check an extractor is reading the real caption track?

Open the video's own transcript panel and read a sentence near the start and one near the end. Compare them with what the tool returned. A genuine extraction matches word for word. If the tool's output reads far more fluently than YouTube's panel on a video you know has a rough automatic track, it may be running its own speech recognition and calling it extraction.

Why are automatic captions wrong so often?

Because they are machine transcription run once, with no human check. Becky Sue Parton's 2016 study of YouTube's auto-generated captions on course video counted 525 phrase-level errors across 68 minutes, an average of 7.7 a minute. Speech recognition has improved since, so treat that figure as historical, but the mechanism has not changed. Every extractor copies those errors faithfully, so switching tools cannot fix them. A human track, a find and replace, or your own speech recognition can.

Why did two extractors give me different text for the same video?

They may have read different tracks, an automatic one against a human one, if the video carries both. One may have used a cached copy from before the creator corrected the track. Line breaks and light formatting also differ between tools. What should not happen is a difference in actual wording and meaning, because every honest extractor reading the same track returns the same words.

Can a free extractor create captions for a video that never had any?

No, not while it is extracting. Extraction can only copy a file that exists. If YouTube never generated or received a track, there is nothing to copy, and an honest tool returns an empty result. Anything that hands you fluent text for a track-less video has switched to speech recognition, which is a different job with its own cost and error pattern.

Why does a browser extension extractor ask for so many permissions?

Often because it was built to work broadly rather than narrowly. An extension scoped to youtube.com only needs access there. One that asks to read and change data on every site you visit is asking for far more than a captions job needs. When a narrowly scoped option exists, such as a web tool, it is the safer choice for a job this small.

What law actually requires captions to exist?

No single law covers every video. In the US, the ADA Title II web rule requires WCAG 2.1 AA, which includes captions, on state and local government sites, with compliance from 26 April 2027 or 26 April 2028 depending on size. Section 508 covers federal agencies, and FCC rules under the CVAA carry TV captions over to online delivery. The UK's 2018 regulations cover public sector sites, and the European Accessibility Act applies from 28 June 2025. None reaches an individual creator on YouTube.

The captions are out of sync with the video. How do I fix that?

If every line is wrong by about the same amount, it is a constant offset, not a damaged file. Open the SRT in any subtitle editor, shift the whole track by that number of milliseconds, then spot check the start, middle and end. If the drift grows as the video goes on, the file was probably made for a different cut of the video. In that case extract again from the exact video you are working with.

My subtitle file shows strange characters instead of accents. What went wrong?

The file is not being read as UTF-8. Accented letters, non-Latin scripts and curly quotation marks all break the same way when a tool assumes an older encoding. Open the file in a text editor that lets you choose the encoding, save it as UTF-8, and the characters come back. If a strict WebVTT parser still refuses it, check for a byte order mark before the WEBVTT header and save again without one.

Extraction worked yesterday and fails today on the same video. Why?

Check the video first. Captions can be removed or replaced by the creator, a video can be switched to private or members-only, and a re-upload keeps the title but gets a new video ID and no track yet. Open the video and look at the CC button. If it is now greyed out, the track is gone. If the CC button is lit and the transcript panel shows text, that is a genuine tool failure worth reporting.

Sources

  1. YouTube Help. Supported subtitle and caption files. support.google.com· Checked 17 August 2026.
  2. YouTube Help. Add subtitles and captions. support.google.com· Checked 17 August 2026.
  3. W3C. WebVTT: The Web Video Text Tracks Format. w3.org· Checked 17 August 2026.
  4. Google for Developers. YouTube Data API: Captions. developers.google.com· Checked 17 August 2026.
  5. YouTube Help. View video transcripts. support.google.com· Checked 2 October 2026.
  6. YouTube Help. Use automatic captioning (supported languages and reasons captions are unavailable). support.google.com· Checked 2 October 2026.
  7. YouTube Help. Use automatic dubbing. support.google.com· Checked 2 October 2026.
  8. Google for Developers. YouTube Data API: Captions: download. developers.google.com· Checked 2 October 2026.
  9. Google for Developers. YouTube Data API: Captions: list. developers.google.com· Checked 17 August 2026.
  10. Google for Developers. YouTube API Services Terms of Service. developers.google.com· Checked 17 August 2026.
  11. YouTube. Terms of Service. youtube.com· Checked 17 August 2026.
  12. Parton, B. S. (2016). Video Captions for Online Courses: Do YouTube's Auto-generated Captions Meet Deaf Students' Needs? Journal of Open, Flexible and Distance Learning 20(1), 8-18. jofdl.nz· Checked 2 October 2026.
  13. W3C WAI. Understanding SC 1.2.2 Captions (Prerecorded). w3.org· Checked 19 August 2026.
  14. W3C WAI. Understanding SC 1.2.4 Captions (Live). w3.org· Checked 19 August 2026.
  15. ADA.gov. Fact Sheet: New Rule on the Accessibility of Web Content and Mobile Apps Provided by State and Local Governments. ada.gov· Checked 2 October 2026.
  16. Section508.gov. Captions and Transcripts. section508.gov· Checked 19 August 2026.
  17. FCC. Closed Captioning of Video Programming Delivered Using Internet Protocol. fcc.gov· Checked 19 August 2026.
  18. FCC. 21st Century Communications and Video Accessibility Act (CVAA). fcc.gov· Checked 19 August 2026.
  19. legislation.gov.uk. The Public Sector Bodies (Websites and Mobile Applications) (No. 2) Accessibility Regulations 2018. legislation.gov.uk· Checked 19 August 2026.
  20. legislation.gov.uk. Directive (EU) 2019/882 (European Accessibility Act), Chapter I. legislation.gov.uk· Checked 2 October 2026.
  21. RNID. Prevalence of deafness and hearing loss. rnid.org.uk· Checked 2 October 2026.
  22. RNID. 10 years of Subtitle It! creating change together. rnid.org.uk· Checked 2 October 2026.
  23. W3C. TTML2: Timed Text Markup Language 2. w3.org· Checked 19 August 2026.
  24. EBU. TECH 3350: EBU-TT Part 1 Subtitling Format Definition. tech.ebu.ch· Checked 19 August 2026.
  25. MDN Web Docs. WebVTT API. developer.mozilla.org· Checked 17 August 2026.
  26. MDN Web Docs. The Embed Text Track element (track). developer.mozilla.org· Checked 17 August 2026.
  27. Wikipedia. SubRip. en.wikipedia.org· Checked 17 August 2026.
  28. YouTubeScribe. Extraction test log, June to August 2026. Internal measurements across public and unlisted videos. Checked 17 August 2026.
  29. YouTubeScribe. Language coverage audit, August 2026. Internal record of the 125 languages read where a caption track exists. Checked 17 August 2026.

Written by Faisal Ashfaq. Faisal Ashfaq maintains the YouTubeScribe caption extractor at Deeporax AI LTD in Oldham. He writes the under-the-hood notes on timed text.

Reviewed on 2 October 2026. If a line is wrong, email support@youtubescribe.com.

Stop reading. Start pasting.

  • No account
  • Nothing stored
  • 125 languages
youtube.com/watch?v=Get the transcript