Free YouTube subtitle extractor: every method that still works

Four ways to pull captions out of a YouTube video, what each one costs you in time, and the six official reasons a video has no track to pull.

Faisal AshfaqMaintains the extractor
Published 12 June 2026Updated 17 August 202616 min read
Free YouTube subtitle extractor: caption strips, ink pot and quill on paper
A subtitle extractor reads the caption track a video already carries. It does not listen to the audio.
What you need to know
  • 01A free YouTube subtitle extractor copies the caption track a video already has, it does not listen to the audio.
  • 02Four routes still work in 2026: a paste-a-link tool, YouTube's own transcript panel, a manual caption file download, and running speech recognition yourself.
  • 03If a video has no caption track, no extractor can produce one, and an empty result is the honest answer.

You want the words out of a YouTube video. Not a summary, not a rewrite, the actual caption text with the timings still attached so you can quote it, search it, translate it or drop it into an editor. That job has a name, subtitle extraction, and for most public videos it takes about ten seconds and costs nothing.

The confusing part is that four different things all get called a free YouTube subtitle extractor, and they behave nothing alike. One copies a file that already exists. One shows you text on screen and makes you drag your mouse over it. One is a developer endpoint that mostly refuses strangers. One transcribes the audio from scratch and charges you in minutes of compute. Pick the wrong one for your job and you either lose an hour or you get a file with no timestamps in it.

This page is the long version. It covers all four routes as they behave in August 2026, what each costs in time for a real 20-minute video, where each one quietly fails, and how to read and repair a caption file once you hold it. One disclosure first: YouTubeScribe is our product, so treat the sections about it as the vendor talking. The limits we list are the real ones, including the videos we cannot touch.

What a subtitle extractor actually does

A YouTube video and its captions are two separate things. The picture and sound live in one stream. The captions live in a timed text file that the player fetches on the side and paints over the frame when you press CC. Nothing about the words is burned into the pixels. So getting subtitles out of a YouTube video is not image processing and it is not audio processing. It is asking for a file that is already sitting there, and keeping a copy.

That single fact explains almost every behaviour people find odd about caption downloaders. Why extraction is fast: it is a file transfer, not a computation. Why it is free: no model runs. Why the accuracy is fixed and no tool can beat another on the same video: everyone is copying the identical source. And why some videos return nothing at all, no matter which of the twenty tools you try.

Two kinds of caption track

YouTube holds two species of track and an extractor will happily take either one.

  • Creator uploaded tracks. Somebody wrote or corrected these and pushed them to the video. They carry real punctuation, speaker labels where the author bothered, correct spellings of names and products, and sensible line breaks. This is the good stuff.
  • Automatic captions. YouTube's speech recognition made these when the video finished processing. No punctuation in many languages, no speaker labels, proper nouns often mangled, and line breaks placed by machine at roughly even intervals rather than at clause boundaries.

You can usually tell them apart in a second. Open the caption menu in the player. A track labelled with a language name and the word auto-generated after it is machine output. A bare language name is human work. Extractors read both, and a good one tells you which kind it handed you, because the difference decides how much editing you are about to do.

Automatic captions do not cover every language. YouTube's own help pages list roughly 70 supported languages for automatic captioning, running alphabetically from Afrikaans to Zulu. That number is worth remembering because it is smaller than most people assume, and it is the reason a perfectly clear video in a less common language sometimes has no track at all.

Extraction is not transcription

These two words get used as if they mean the same thing and they do not. Extraction reads an existing caption track and copies it. Transcription takes the audio and produces new text using speech recognition. Both end with words on your screen. Everything else about them differs.

ExtractionTranscription
InputThe caption file YouTube already storesThe audio track
Speed for 20 minutesSecondsMinutes, sometimes longer than real time on weak hardware
CostUsually freeCompute time or an API bill per minute
Accuracy ceilingExactly as good as the existing trackDepends on the model, can beat a bad auto track
Works with no caption trackNoYes
TimestampsInherited, frame accurate to the originalGenerated, quality varies by model

YouTubeScribe sits firmly on the extraction side. It reads the caption track a public or unlisted video already has. It never invents speech and it does not run speech recognition on the audio. If YouTube never wrote a track, the answer you get back is nothing, and that is the correct answer rather than a bug.

A miss is a valid result

Plenty of tools respond to a track-less video by producing plausible sentences anyway. That output has no relationship to what was said. If you are quoting a source, citing a lecture or building a dataset, a blank result you can verify beats a paragraph you cannot.

What a caption cue looks like inside

Open any SRT file in a text editor and the structure is obvious within five seconds. Each cue has four parts: a sequence number on its own line, a timing line with a start time, an arrow and an end time, one or two lines of text, then a blank line before the next cue. Times run hours, minutes, seconds, then a comma and three digits of milliseconds.

WebVTT holds the same information with two differences that trip people up. The file must begin with the literal word WEBVTT on the first line, and the milliseconds separator is a full stop rather than a comma. Sequence numbers are optional. That is genuinely most of the difference for ordinary subtitle work, which is why converting between the two is a find and replace rather than a parse and rebuild.

Once you can read a cue, you can audit any extractor's output in about thirty seconds. Check that the first cue starts near the first spoken word. Check that the last cue lands near the end of the video. Scan the middle for a gap of several minutes, which usually means the track itself is partial rather than the download failing.

The four ways to get subtitles out of a video

Here is the honest comparison. Times are for a single 20-minute talk with an existing English caption track, measured start to finish, including the fiddly bits everybody forgets to count like opening a new tab and cleaning up the paste.

MethodTime for a 20-minute videoTimestampsBest forCost
Paste a link into YouTubeScribeAbout 10 secondsYes, SRT and VTT keep them, TXT drops them on requestAnyone who wants a file, in bulk or one at a timeFree for TXT and SRT, first transcript each day free
YouTube's own transcript panel3 to 8 minutes with the copying and tidyingInline text only, not a real cue fileA quick read or one quote, no install, no third partyFree
Download the caption file by hand5 to 20 minutes, longer the first timeYes, the original fileYour own uploads, or a channel that granted you accessFree but needs an account, and API setup for scripts
Run speech recognition yourself5 to 40 minutes depending on hardwareYes, generated freshVideos with no caption track at allCompute time, or per-minute API charges

Method 1: paste the link into a free extractor

The fastest route by a wide margin. You copy the video URL, paste it, choose a format, and a file lands in your downloads folder. There is nothing to install, nothing to authorise, and no account needed to pull a transcript on the free tier.

What makes one tool better than another here is not accuracy, since everybody reads the same source track. It is the boring operational stuff: whether long videos are refused, how many languages are offered when a video carries several tracks, whether SRT costs money, whether your video URLs are logged, and whether the tool admits it found nothing instead of inventing a transcript.

On YouTubeScribe there is no hard length cap. A three-hour conference recording comes back in about twelve seconds. It reads 125 languages where a track exists, which is not the same claim as YouTube's roughly 70 automatic caption languages: those 70 are the languages the machine will listen in, while the 125 are the languages we can read back out once any track exists, including every human-uploaded translation. The Subtitle Downloader is the page built for this job, and it gives you TXT and SRT free, with VTT on Pro.

Method 2: YouTube's own transcript panel

YouTube ships a reading view for captions and most people have never opened it. On desktop, go under the video, open the description, and look for the button marked Show transcript. On mobile the same control lives in the description sheet. A scrollable panel appears beside the player with the caption text, each line prefixed by its start time, and clicking any line jumps the player there.

It is genuinely useful for reading. There is a three-dot menu inside the panel that toggles the timestamps off, and where a video carries multiple tracks a language picker appears at the top. For finding the one sentence you half remember, this beats everything else because you can scan and jump in the same motion.

The problem is getting the text out. There is no download button. You select the text with the mouse, which on a long video means a scroll-and-drag that keeps snapping to the video player, and what lands on your clipboard is a wall of text with timestamps welded into the same lines. It is not a cue file. Nothing downstream that expects SRT will accept it. For one quote it is fine. For a 90-minute podcast it is a genuinely unpleasant ten minutes.

  • Good for: one quote, a fast skim, checking whether a track exists before you bother with anything else.
  • Bad for: any file you need to feed to an editor, a translator, a search index or a script.
  • Watch out for: the panel showing the auto track when a better human track exists, since the language picker defaults are not always the best track available.

Method 3: download the caption file by hand

There are two versions of this and people conflate them constantly.

The first is YouTube Studio. If the video is on your own channel, open Studio, go to Subtitles, pick the video, open the track and use the download option. You get the real file, exactly as stored, in the format you ask for. This is the cleanest source there is. It is also useless for anybody else's video, which is the case most people are actually in.

The second is the YouTube Data API. There is a captions resource with a list method that enumerates the tracks on a video, and a download method that returns the track body. The list call is broadly available. The download call is the one that stops people: it needs OAuth rather than a plain API key, and it is scoped to captions you own or where the owner has granted third-party contributions. It was built so a creator or their caption vendor can pull their own files, not so anybody can harvest anybody's. Read the terms before you build on it.

Budget an afternoon for the API route the first time: a Google Cloud project, the API enabled, an OAuth consent screen, a client secret, a token exchange, then the actual two calls. If you own the videos and you need this to run nightly on a hundred uploads, it is the right tool. For a handful of videos you do not own, it is the wrong shape entirely.

Method 4: run speech recognition yourself

The last resort, and the only route that works when no caption track exists. You pull the audio and run it through a speech model, locally or through a paid API, and it writes fresh cues with fresh timings.

It is the only method here that can beat the original: a good modern model on clean audio will often out-punctuate and out-spell a lazy auto track. It is also the only method that can be badly wrong in ways you will not notice, because a confident model produces fluent text for mumbled audio and you have no source file to check it against.

  • Use it when: the video has no track at all, the auto track is unusable, or you need a language YouTube's automatic captioning does not cover.
  • Skip it when: a human-uploaded track already exists, because you will spend forty minutes producing something worse than a free ten second download.
  • Always check: the first and last minute, and any section with music or crosstalk. That is where invented text hides.
Extracting YouTube subtitles from a video link into SRT and TXT caption files with timestamps
Extraction copies the caption track YouTube already stores, then writes it out as SRT, VTT or plain text.

Step by step with a free extractor

This is the paste-a-link route, written out fully so you can see exactly where the seconds go. The tool is the Subtitle Downloader on YouTubeScribe, but the shape of the process is the same on any honest extractor.

  1. Copy the video URL. The share button gives you a short youtu.be link, the address bar gives you the long watch link with a v parameter. Both work. A link with a timestamp on the end also works, the timestamp is simply ignored.
  2. Open the Subtitle Downloader and paste the link into the field.
  3. Pick the language if the video carries more than one track. Where several exist, prefer the one without the auto-generated label.
  4. Pick the format. TXT for reading and writing, SRT for editing and re-uploading, VTT for the web player.
  5. Run it. A 20-minute video comes back in a couple of seconds. A three-hour recording takes about twelve.
  6. Check the file before you build on it. Open it, confirm the first cue lines up with the first words spoken, and scroll to the end to confirm it did not stop early.

Step six is the one everybody skips and it is the one that saves you. A partial track looks exactly like a complete one until you reach the end of it.

Choosing the output format

Match the format to what happens next, not to what sounds most technical.

  • TXT when a person is going to read it, or a language model is going to summarise it, or you want to paste it into a document. No timing noise, no cue numbers, just prose. Free.
  • SRT when the text is going back onto video, into a subtitle editor, into a translation workflow, or into anything that needs to know when each line appears. The most widely accepted subtitle format there is. Free.
  • VTT when the destination is a web player using the HTML track element, or a platform that specifically asks for WebVTT. Available on Pro.

One habit worth forming: keep the SRT even when you only wanted the text. Going from SRT to TXT is trivial at any point later. Going from TXT back to SRT means re-timing every line by hand, and nobody has ever enjoyed that.

More than one video at a time

One video is easy. Forty is where tools separate. If you are pulling a whole course, a conference series or a competitor's back catalogue, feeding links one by one is the slow way to spend a morning.

The Playlist Export tool takes a playlist or channel URL and works through it, so you queue the list once and collect the files at the end. If what you want is readable prose rather than cue files, the Transcript Generator is the sibling tool pointed at that job: same extraction underneath, output shaped for reading and quoting rather than for a subtitle editor.

Expect misses in any batch. A playlist of sixty will usually contain two or three videos that come back empty, because a member-only upload slipped in, or a very short clip never got a track, or a re-upload is still processing. That is normal. Note them and move on rather than assuming the run failed.

The limits, stated plainly

Since this is our tool, here is what it will not do, in the same place as what it will.

  • Private and members-only videos return nothing. There is no track we are permitted to read, and there is no clever workaround being withheld from you.
  • It does not run speech recognition. No track means no output, full stop.
  • It cannot improve a bad automatic track. Garbled audio in, garbled captions out, because the garbling happened at YouTube months ago.
  • The free tier stores nothing. That means privacy, and it also means the file is yours to keep safe, because we cannot re-send you something we never kept.
  • The first transcript each day is free. TXT and SRT are free formats. VTT is a Pro format.

The single most common support message we get is that the tool is broken because it returned nothing. In almost every case the video genuinely has no caption track. The fix is not a different extractor, it is speech recognition or a different source video.

YouTubeScribe support notes, June to August 2026

SRT, VTT, TXT: picking and fixing the file

You have a file. Now the questions change: will the thing you are feeding it to accept it, and if not, what do you edit.

SubRip, the default answer

SRT is the plain cotton of subtitle formats. Sequence number, timing line, text, blank line, repeat. It carries no positioning and no styling, and that poverty is precisely why everything accepts it: there is nothing in the file to misinterpret. Video editors, translation platforms, social schedulers and every upload form you will meet take SRT.

YouTube's own upload documentation describes SRT as plain UTF-8 text with no markup, which is the standard everyone else follows too. If you open an SRT and see styling tags in it, some tool has been creative and a strict parser downstream may object.

WebVTT, the web native one

WebVTT is the W3C format that browsers understand natively through the HTML track element. Same cue idea as SRT, plus cue settings for position, alignment and line placement, plus a proper header. YouTube accepts it on upload and supports its positioning, though it limits styling to bold, italic and underline rather than the full range the spec allows.

Convert SRT to VTT by hand in three moves: add a line reading WEBVTT and a blank line at the top, change every comma in the timing lines to a full stop, and save as UTF-8 without a byte order mark. Going the other direction, strip the header and reverse the separator. The cue numbers are optional in VTT so you can leave them or drop them.

Plain text, and what it costs you

TXT is a caption file with the timing thrown away. That is a real loss and a real gain. The gain is that prose reads like prose, so summarising, searching, quoting and pasting all get easier, and a language model handling it wastes no attention on timecodes. The loss is that you can never put it back on video without re-timing every line.

Rule of thumb: if a human or a model is the final reader, take TXT. If a player is the final reader, take SRT or VTT. If you are not sure, take SRT, because it converts down to TXT in one step whenever you decide.

What YouTube accepts if you put captions back

Extraction is often only half the job. People pull a track, fix the spellings, translate it and push it back up. YouTube's supported files list is much wider than most people expect, and worth knowing before you convert something unnecessarily.

FormatExtensionPositioning and styling
SubRip.srtNone, plain UTF-8 text with no styling markup
WebVTT.vttPositioning supported, styling limited to bold, italic and underline
TTML and DFXP.ttml, .dfxpStyling and positioning both supported
Scenarist Closed Caption.sccBroadcast format, YouTube's preferred choice for that class
SAMI and RealText.smi, .rtLegacy formats, still accepted
SubViewer.sbv, .subBasic timing and text
MPsub, LRC, Videotron Lambda.mpsub, .lrc, .capBasic timing and text
EBU-STL and CEA-608 family.stl and relatedBroadcast delivery formats

The practical reading of that table: if your file is already in one of those shapes, upload it as it is. Converting a TTML file down to SRT to be safe throws away positioning that YouTube would have honoured. If you are authoring from nothing, SRT stays the sensible default because it is the format every other tool in your chain will also accept.

Five file problems and their fixes

  1. Accented characters appear as question marks or mojibake. The file is not UTF-8. Re-save it as UTF-8 and the problem disappears.
  2. A strict parser rejects a valid looking VTT. Check for a byte order mark before the WEBVTT header. Save without a BOM.
  3. Every subtitle shows a fixed amount early or late. That is a constant offset, not a broken file. Any subtitle editor will shift the whole track by a set number of milliseconds in one action.
  4. The upload form refuses an SRT. Look for styling tags or HTML inside the cues. SRT is supposed to be plain text and some parsers enforce it.
  5. Cues render on one line where you expected two. Line breaks inside a cue are part of the text and survive conversion, so if they vanished, the tool that wrote the file flattened them.

Illustrative example

A localisation team pulls SRT for 40 conference talks, fixes speaker names and product spellings once in a find and replace across all 40 files, then hands the corrected set to translators. Extraction is minutes, the correction pass is an afternoon, and no one re-times anything. Figures here are illustrative rather than a measured client result.

When there is nothing to extract

The hardest part of writing about extraction honestly is this section, because the answer is sometimes that you cannot have the thing you want. No tool on the internet can copy a file that was never written.

Google's own list of reasons

YouTube documents six reasons automatic captions may be unavailable on a video. Every one of them is a genuine dead end for an extractor, not a limitation of the tool you are using.

  • The captions are not available yet because the audio is complex and still processing.
  • Automatic captions do not support the language spoken in the video.
  • The video is too long.
  • The sound quality is poor, or the speech is not recognised.
  • There is a long period of silence at the beginning of the video.
  • Multiple people are speaking over each other, or multiple languages are spoken at once.

Two of those are worth a second look. The processing one is temporary: a video uploaded in the last hour may simply not have a track yet, and coming back tomorrow fixes it. The silence one catches a surprising number of creators, because a long branded intro with only music at the top of the video can be enough for the recogniser to give up before anyone speaks.

The access wall

Separate from whether a track exists is whether you are allowed near it. Public and unlisted videos are readable. Private videos are not, and neither are members-only uploads, because both sit behind an access check that an extractor has no business defeating.

If a video is private and you have a legitimate reason to need its captions, the route is the channel owner: they can pull the file from YouTube Studio in about a minute and send it to you. That is a faster conversation than any amount of tool-hunting.

The language gap

Two numbers matter here and they measure different things. YouTube's automatic captioning covers roughly 70 languages. YouTubeScribe reads 125 languages where a track exists. The second number is larger because it includes every human-uploaded and human-translated track, which can be in a language the machine recogniser has never supported.

So a video in a language outside the automatic 70 will only have captions if a person made them. When creators in those languages do upload their own tracks, extraction works perfectly and the quality is usually better than any auto track, because a human wrote it.

A five-minute triage

Before you conclude a tool is broken, run this. It takes less time than opening a support ticket.

  1. Open the video and look at the CC button in the player. Greyed out or missing means there is no track and nothing to extract. Stop here.
  2. Open the caption menu and read the language list. Note whether the entry says auto-generated, and whether a better human track sits below it.
  3. Open the description and press Show transcript. If the panel appears with text, a track exists and any extraction failure is a tool problem worth reporting.
  4. Check the upload date. Uploaded in the last few hours means the track may still be processing. Try again tomorrow.
  5. Check the access state. A padlock, a members badge or a video that only plays for you when logged in means the access wall, not a missing file.
  6. Only after all five, reach for speech recognition. If the video truly has no track, that is the one route left, and it is a real one.

Extraction is not a licence

Getting a caption file is a technical act, using it is a legal one. Quoting, studying, indexing and accessibility work sit comfortably within normal use. Republishing somebody's full transcript as your own content does not. YouTube's terms and the API terms are the documents that govern this, and they are short enough to read.

The summary is short. For a public video with an existing track, paste the link into a free extractor and you are done in seconds with a real SRT. For a fast quote, YouTube's own transcript panel is right there and costs nothing. For your own uploads at scale, Studio or the API gives you the pristine originals. And for a video with no track at all, speech recognition is the only honest answer, with all the cost and all the caution that implies.

Questions people ask

What is a free YouTube subtitle extractor?

It is a tool that copies the caption track a YouTube video already has and gives it to you as a file. The video and its captions are stored separately, so the extractor requests the timed text file the player uses and saves it as SRT, VTT or plain text. Because it is a file transfer rather than a computation, it is fast and can be offered free. It does not listen to the audio and it cannot produce captions for a video that never had a track.

Is extracting subtitles from YouTube legal?

Reading a caption track from a public video is a normal technical act, and quoting, studying, indexing or making content accessible sits within ordinary use in most places. What changes the picture is what you do next. Republishing a full transcript as your own content, or building a commercial product on somebody's words without permission, is a rights question rather than a technical one. YouTube's terms of service and the YouTube API services terms are the documents that govern it, and both are short enough to read.

What is the difference between subtitles and captions?

Subtitles traditionally carry the spoken dialogue, often translated into another language, and assume the viewer can hear. Captions are written for viewers who cannot hear the audio, so they add speaker labels and sound descriptions such as door closes or laughter. YouTube uses the two words almost interchangeably across its interface and help pages, and its upload documentation calls the same files subtitle and caption files. For extraction purposes the distinction rarely matters: you get whatever text is in the track.

How many languages does YouTube support for automatic captions?

Roughly 70, running alphabetically from Afrikaans to Zulu according to YouTube's own help pages. That is the list of languages the speech recogniser will listen in, so a video spoken in a language outside it will not get an automatic track. It is a separate number from how many languages an extractor can read. YouTubeScribe handles 125 languages where a track exists, and the extra ones are covered because humans upload and translate tracks in languages the recogniser has never supported.

Does a subtitle extractor work on any YouTube video?

No, and any tool claiming otherwise is not being straight with you. It works on public and unlisted videos that already carry a caption track. Private videos and members-only uploads return nothing because they sit behind an access check. Videos with no track at all return nothing because there is nothing to copy. Roughly speaking, if the CC button in the player is lit, extraction will work. If it is greyed out or absent, no extractor will help you.

How do I download subtitles from a YouTube video for free?

Copy the video link, open the Subtitle Downloader on YouTubeScribe, paste it in, choose a language if the video has more than one track, choose TXT or SRT, and run it. A 20-minute video finishes in a couple of seconds and a three-hour recording in about twelve. TXT and SRT are free formats and the first transcript each day is free. Then open the file and check the last cue lines up with the end of the video, because partial tracks look complete until you scroll.

How do I see the transcript inside YouTube itself?

On desktop, open the video, expand the description below it, and click Show transcript. A panel appears beside the player with the caption text and a start time on each line, and clicking a line jumps the player to that moment. The three-dot menu inside the panel toggles the timestamps off, and a language picker appears when the video carries several tracks. There is no download button, so getting the text out means selecting it by hand.

Should I choose SRT, VTT or TXT?

Choose by destination. SRT if the text is going back onto video, into a subtitle editor or into a translation workflow, because it is the format every tool accepts. VTT if a web player using the HTML track element is the destination, or a platform explicitly asks for WebVTT. TXT if a person or a language model is the final reader and the timings would only be noise. When unsure, take SRT: converting SRT to TXT later is one step, while going back the other way means re-timing every line.

Can I extract subtitles from a whole playlist at once?

Yes. The Playlist Export tool takes a playlist or channel URL and works through the videos in it, so you queue the list once rather than pasting links one at a time. Expect some misses in any large batch: a playlist of sixty will usually include two or three videos that return nothing because a members-only upload slipped in, a very short clip never got a track, or a recent re-upload is still processing. Those are normal results rather than a failed run.

Can I upload an edited subtitle file back to YouTube?

Yes, and the accepted list is wider than most people expect. YouTube takes SubRip .srt as plain UTF-8 with no styling markup, WebVTT .vtt with positioning and limited styling, TTML and DFXP with full styling and positioning, SCC as its preferred broadcast format, plus SAMI, RealText, SBV and SUB, MPsub, LRC, Videotron Lambda .cap, EBU-STL and several CEA-608 broadcast formats. If your file is already in one of those, upload it as it is rather than converting down and losing positioning.

Why does the extractor return nothing for my video?

Almost always because the video has no caption track. YouTube lists six reasons automatic captions may be missing: the audio is complex and still processing, the language is not supported by automatic captioning, the video is too long, sound quality is poor or the speech is not recognised, there is a long silence at the start, or several people or languages overlap. Separately, private and members-only videos return nothing because of access rather than absence. Check the CC button in the player first.

The captions are out of sync with the video. How do I fix that?

If every line is wrong by about the same amount, it is a constant offset and not a damaged file. Open the SRT in any subtitle editor and shift the whole track by that number of milliseconds in one action, then spot check the start, middle and end. If the drift grows as the video goes on, the file was probably made for a different cut of the video, such as a version with a sponsor segment removed. In that case re-extract from the exact video you are working with.

My subtitle file shows strange characters instead of accents. What went wrong?

The file is not being read as UTF-8. Accented letters, non-Latin scripts and curly quotation marks all break the same way when a tool assumes a legacy encoding. Open the file in a text editor that lets you choose the encoding, save it as UTF-8, and the characters come back. If a strict WebVTT parser still refuses a file that looks correct, check for a byte order mark sitting before the WEBVTT header and save again without one.

The auto-generated captions are full of errors. Can a different extractor do better?

No. Every extractor copies the identical file from YouTube, so switching tools cannot change a single word. The errors were made by YouTube's speech recogniser when the video was processed and they are baked into the source. Your real options are three: check whether the creator also uploaded a human-written track, fix the file yourself with a find and replace for repeated mistakes such as names and product terms, or run speech recognition on the audio yourself with a modern model, which is the only route that can actually beat the original.

Extraction worked yesterday and fails today on the same video. Why?

Check the video first rather than the tool. Captions can be removed or replaced by the creator, a video can be switched from public to private or unlisted to members-only, and a re-upload keeps the same title while carrying a completely different video ID and no track yet. Open the video and look at the CC button: if it is now greyed out, the track is gone and no tool can retrieve it. If the CC button is lit and the transcript panel shows text, that is a genuine tool failure worth reporting.

Sources

  1. YouTube Help. Use automatic captioning. support.google.com· Checked 17 August 2026.
  2. YouTube Help. Supported subtitle and caption files. support.google.com· Checked 17 August 2026.
  3. YouTube Help. Add subtitles and captions. support.google.com· Checked 17 August 2026.
  4. W3C. WebVTT: The Web Video Text Tracks Format. w3.org· Checked 17 August 2026.
  5. Google for Developers. YouTube Data API: Captions. developers.google.com· Checked 17 August 2026.
  6. YouTube Help. View video transcripts. support.google.com· Checked 17 August 2026.
  7. YouTube Official Blog. Caption my YouTube videos. blog.youtube· Checked 17 August 2026.
  8. W3C. TTML2: Timed Text Markup Language 2. w3.org· Checked 17 August 2026.
  9. Google for Developers. YouTube Data API: Captions: download. developers.google.com· Checked 17 August 2026.
  10. Google for Developers. YouTube Data API: Captions: list. developers.google.com· Checked 17 August 2026.
  11. Google for Developers. YouTube API Services Terms of Service. developers.google.com· Checked 17 August 2026.
  12. YouTube. Terms of Service. youtube.com· Checked 17 August 2026.
  13. AssemblyAI. What is a subtitle file format? assemblyai.com· Checked 17 August 2026.
  14. Ditto Transcripts. SRT vs VTT: understanding the difference between subtitle formats for captions. dittotranscripts.com· Checked 17 August 2026.
  15. MDN Web Docs. WebVTT API. developer.mozilla.org· Checked 17 August 2026.
  16. MDN Web Docs. The Embed Text Track element (track). developer.mozilla.org· Checked 17 August 2026.
  17. Wikipedia. SubRip. en.wikipedia.org· Checked 17 August 2026.
  18. YouTubeScribe. Extraction test log, June to August 2026. Internal measurements across public and unlisted videos. Checked 17 August 2026.
  19. YouTubeScribe. Language coverage audit, August 2026. Internal record of the 125 languages read where a caption track exists. Checked 17 August 2026.

Written by Faisal Ashfaq. Faisal Ashfaq maintains the YouTubeScribe caption extractor at Deeporax AI LTD in Oldham. He writes the under-the-hood notes on timed text.

Reviewed on 17 August 2026. If a line is wrong, email support@youtubescribe.com.

Stop reading. Start pasting.

  • No account
  • Nothing stored
  • 125 languages
youtube.com/watch?v=Get the transcript