How to turn a YouTube video into a blog post
A working method for a 90-minute podcast: transcript first, then a spine, then a draft you would actually publish under your own name.

- 01A 90-minute episode transcribes to roughly 12,000 to 15,000 words. The post you publish should be 1,200 to 2,000. Most of the work is deciding what to throw away.
- 02Chapters the creator wrote beat any outline a model guesses, because the person who was in the room already decided where the hour splits.
- 03Keep timestamps on every quote until the final proofread. Strip them early and you cannot find the line again in 15,000 words.
A 90-minute podcast episode transcribes to roughly 12,000 to 15,000 words. The blog post you publish from it should be 1,200 to 2,000. That ratio is the entire job. Everything below is a method for deciding which tenth of the hour survives, and for making the surviving tenth read like something a person wrote on purpose rather than something a machine flattened out of a conversation.
Most guides on this topic are tool pages. Paste a link, get an article, done. The paste is the fast part, and it was never the hard part. The hard part is editorial judgement: which claim is the spine of the piece, which anecdote earns 200 words, which surname the automatic captions got wrong, and how much of the finished post has to be yours rather than the speaker's. That is what this guide covers, in the order you actually do it.
Disclosure: YouTubeScribe is our product, and the tools named here are ours. We have run this workflow on our own episodes and on other people's shows since 2024, and the parts we got wrong are in here too. The worked example throughout is a 90-minute interview episode, because that is the hardest common case. A 12-minute tutorial is the same method with far less cutting.
Start with the transcript, not the outline
Before you plan anything, check whether the episode has a caption track. Open the video, click the CC button, and see whether words appear. That single click decides whether the rest of this method is available to you, and it takes three seconds. People skip it, build a content calendar around a show, and find out on Thursday that the show has no captions.
YouTubeScribe reads the caption track a public or unlisted video already has. It does not run speech recognition on the audio, and it never invents speech. That is why a transcript comes back in seconds rather than minutes, and it is also why a video with no caption track returns nothing at all. We would rather tell you the track is missing than hand you a file of plausible guesses with a confident tone.
Paste the episode URL into the Transcript Generator. Download TXT if you only want to read, and SRT if you want a timestamp attached to every line, which for this workflow you do. VTT is on Pro. The first transcript each day is free, and on the free tier nothing is stored: the file lands in your browser and that is the end of it. Nothing sits on our side waiting to be indexed or sold.
What you are actually downloading
A caption track is timed text the platform already serves alongside the video. Reading it is not transcription. It is the difference between opening a file and recording one, and it explains both the speed and the limit: seconds when the track exists, nothing at all when it does not.
If there is no caption track, this method stops
There is one hard dependency in the whole workflow, and that is it. If the CC button shows nothing, you are in speech recognition territory, which is a different job with a different budget and a different error profile. You have two honest options from there. Ask the show for their own transcript, since most podcast hosts generate one automatically for the feed and many creators will send it if you say what you are writing. Or run the audio through a speech recognition service and plan to check every proper noun by hand, because that is where the errors will be.
How accurate is the file you just downloaded
Good, not perfect. Industry guidance commonly reports automatic captions landing somewhere around 85 to 95 percent accuracy, and those are figures reported across the industry rather than anything we measured ourselves. Do the arithmetic anyway. Five percent of a 14,000 word transcript is roughly 700 wrong words. Most of them are harmless: a dropped article, a contraction expanded, a filler noise written out. The dangerous ones cluster in exactly the places a reader notices.
- Proper nouns, especially surnames and company names that are not ordinary English words.
- Statistics, percentages and dates, where one wrong digit changes the claim you are publishing under your name.
- Technical terms and acronyms, which get mapped to the nearest common word that sounds similar.
- Speaker turns, because a plain caption track carries no speaker labels at all.
- Punctuation and sentence boundaries, which are inferred from pauses rather than from grammar.
None of that makes the transcript unusable. It makes the transcript a source document rather than a manuscript. Treat it the way a reporter treats an interview recording: reliable for the shape of what was said, checked by hand for anything specific.
Which source videos repurpose well
Not every video wants to be an article. The test is simple: does the information live in the speech? If a viewer could follow the video with their eyes closed and lose nothing, it will repurpose cleanly. If the value sits on the screen, or in the edit, or in a face reacting, you will fight the format the whole way and the result will read thin.
| Source type | How well it repurposes | What it becomes | Where it breaks |
|---|---|---|---|
| Interview podcast | Very well | A themed piece built on two or three claims, with quoted exchanges as evidence | The guest recaps their own book for 20 minutes and none of it is quotable |
| Webinar or training | Very well | A step by step guide with the live Q and A becoming an FAQ section | Slides carry numbers the speaker never says out loud, so the transcript loses them |
| Tutorial or screencast | Well, with extra work | A written procedure with screenshots you capture yourself | Every 'click here' needs an image the caption track cannot give you |
| Conference talk or lecture | Very well | An argument piece that follows the speaker's own structure | Audience questions are off-microphone and the captions drop them entirely |
| Panel discussion | Moderately | A round-up of positions, one section per panellist | Cross-talk makes attribution hard when there are no speaker labels |
| Sermon or lesson | Very well | A reflective piece with the passage or set text as the spine | Long quoted readings need their own citation and permission check |
| Product demo | Moderately | A feature explainer with a comparison table you build yourself | Reads as a brochure unless you add context the vendor did not give |
| Vlog, reaction or gameplay | Poorly | Very little worth publishing as prose | The value is visual and the speech is reactive filler around it |
Keep the timestamps until the last edit
Download the timestamped file even though timestamps look ugly in a working document. While editing you will need to jump back to the audio to check a name, a figure, or whether a quote is really what was said. In a 14,000 word file, finding a line again without a timestamp means reading until you hit it. We stripped timestamps early exactly once, and spent an afternoon learning not to.
One more job before you leave the raw file alone: give it a structural pass. Automatic caption tracks arrive as a wall of short lines with no paragraphs, and a wall is hard to skim. Merge the lines into paragraphs at the obvious topic changes, keeping the timestamp of the first line in each paragraph. Ten minutes of that makes the next hour of reading much faster, and it costs you nothing in accuracy because you are not changing a single word of what was said.
Cut 15,000 words into a spine
The spine is the shape of the article before any sentences exist. It is a list of section headings and, under each one, a single sentence saying what that section proves. Build it before you write anything, because the spine is where you throw material away, and throwing material away is far cheaper in a list than in finished prose you have grown fond of.
Chapters, when the creator wrote them, beat anything a model guesses
Check the description and the player scrubber for chapters. If the creator wrote them, they have already made the decision you are about to make: this hour splits here, then here, then here. A model reading a transcript is inferring that structure from word patterns. A human who was in the room is not inferring anything. Pull the list with the Chapters Extractor and start your spine from it.
YouTube requires at least three timestamps for chapters to appear, the first starting at 00:00, with each segment running at least ten seconds. That floor is worth knowing because it tells you what you are looking at. A creator who wrote eleven chapters for a 90-minute show has done real structural work and you should respect it. A show with three chapters called 'Intro', 'Interview' and 'Outro' has told you nothing, and you are building the structure yourself.
With no chapters, read for first appearances
Skim the transcript looking for the moment a new idea enters, not the moment it gets discussed at length. Interviews circle. The same point surfaces at 12 minutes, again at 41, again in the wrap-up. The first appearance is usually the cleanest statement of it, and the later passes are elaboration and agreement. Mark first appearances. Those marks become your headings.
The Summary Generator helps here, but as a map rather than as copy. Run it, read the summary, then go back into the transcript and find where each summarised point actually lives so you can quote it properly. Never paste a summary into the draft. It reads like a summary, because it is one, and readers can tell within a paragraph.
Building the spine
- List the chapters, or your marked first appearances, in the order they occur in the episode.
- Write one sentence under each that says what a reader learns there. If you cannot write that sentence, it is not a section.
- Delete any line whose sentence duplicates another line's sentence. Two chapters are often one argument in different clothes.
- Cut what remains to five or six sections. A 1,500 word post cannot carry eleven headings without turning into a list of stubs.
- Pick one quote per surviving section, four lines at most, and write its timestamp beside it.
- Reorder for the reader, not for the recording. Nothing obliges you to keep the episode's running order.
Before the spine is finished, write one sentence at the top of the document saying what this post argues. Not what the episode covered. What the post argues. An hour of conversation holds four or five arguable claims, and you are picking one to carry the piece while the rest become supporting material or a second post next month. Any section that does not serve that sentence gets cut, however good the quote inside it is. This one line does more cutting than every other step combined, and writing it is uncomfortable, which is how you know it is working.
What goes in the first cut
- Sponsor reads in full, including the ones woven into the conversation so they sound spontaneous.
- The opening minutes of greetings, weather and how everyone has been since last time.
- The guest summarising their own book or product. Link to the thing instead.
- 'Does that make sense', 'right, right, right', and every other conversational nod that carried meaning with a face attached.
- Cross-talk where two people speak at once and the caption track returns word salad.
- The outro: subscribe, rate, review, where to find us, see you next week.
- Any anecdote that needs three minutes of setup to land a small point.
On an interview show, that first cut usually removes 60 to 70 percent of the file. What remains is still too long, and that is normal. The second cut happens while drafting, when you discover which sections actually have something to say and which ones only felt substantial because the speakers were enjoying themselves.
If a section only exists because it happened, cut it. The recording is a record of an hour. The post is an argument about that hour.
Faisal Ashfaq, YouTubeScribe

Write the draft a person would publish
Now you can generate. The Video to Blog Post tool at /tools/video-to-blog-generator takes the episode and returns a structured first draft with headings and pull quotes already placed. It is a Pro tool. Be clear with yourself about what arrives: a first draft. Headings that are roughly right, quotes that are roughly placed, paragraphs that hold together and say nothing anyone will remember on Tuesday. It removes the blank page and the mechanical work of shaping 14,000 words into sections. It does not remove the writing.
Put the generated draft next to your spine. Where the two agree, you have confirmation that the structure is obvious from the material. Where they disagree, trust your spine, because you listened to the episode and the tool read a text file. Then start deleting from the generated draft rather than adding to it.
The 40 percent rule
Common guidance in content teams is that a repurposed post should be at least 40 percent original material, and as floors go it is a sensible one. The transcript is the foundation, not the finished product. Your introduction, your conclusion, the context the speakers assumed everyone shared, the examples they did not give, and the analysis of why any of it matters: that is the part that makes the post worth reading instead of worth skipping in favour of the audio.
In practice this is less abstract than it sounds. Here is the part only you can write.
- The opening. The episode opened with hellos. Your post opens with the claim.
- The context. Speakers assume a shared world. A reader arriving from search does not know who the guest is or why the argument matters this year.
- The bridges. Conversation jumps between ideas and nobody minds. Prose has to carry the reader across the gap on purpose.
- The counterpoint. If a guest made a claim you think is wrong or partial, say so, in your own voice, with your name on it.
- The ending. Podcasts end when the time runs out. Articles end on a point.
One practical way to hold that line: when the draft is done, highlight every sentence that could only have come from the transcript. If the highlighting covers more than half the page, you have published a transcript with headings on it. The check takes two minutes and it is much harder to argue with than a word count.
Structure that survives contact with a reader
The generated draft will give you a competent structure that mirrors the episode. Replace it with a structure that serves someone who has not heard the episode and may never hear it. The two are rarely the same shape, because a conversation builds toward its best material and an article has to lead with it.
- Open with the single most interesting thing anyone said, in your words, in two sentences.
- By the third paragraph, say what the post covers and who it is for.
- One idea per section, with the heading naming the idea rather than repeating the chapter title.
- Quotes as evidence, never as filler. If the quote does not say it better than you can, paraphrase and move on.
- Close with what changes for the reader, not with a summary of what the episode covered.
Write the title after you know the argument
Titles lifted from the episode's own title inherit the show's framing, which serves people who already follow the show and nobody else. A reader landing from a results page needs to know what they get and why it is worth six minutes of their evening. Say the useful thing plainly, put the concrete noun early, and keep the guest's name in the title only when that name is the reason anyone would click. Then write the meta description as a promise you can keep, not as a summary of the episode.
Do not write it in the host's voice
Publish the piece as your write-up of the episode, in your own voice, with the speakers quoted properly. Ghost-writing a personality out of a transcript is a trap, and we walked into it, which is covered in the last section. The version that works is honest about what it is: you listened, you thought the interesting parts through, and here they are with the quotes attached and the timestamps linked.
Rights
A public transcript is not a licence to republish someone's episode as your article. Quote fairly, attribute by name, and link the episode. If the show is not yours, do not present the post as an official write-up. If you are on the same team as the show, say so in the post rather than leaving readers to work it out.
Edit against the audio
This is the pass that separates a publishable post from transcript soup, and it is the pass every tool landing page leaves out. Budget more time for it than for the draft. On a 90-minute episode we spend roughly twice as long editing as drafting, and the ratio has never gone the other way.
Check every proper noun, number and technical term
Open the timestamped file and the video side by side. Every name, figure and acronym in your draft gets checked against the audio at its timestamp. Given the caption accuracy commonly reported across the industry, this is not optional work. The failures are boring and consistent: a surname rendered phonetically, a company name split into two ordinary words, a figure in basis points losing a digit, an acronym expanded into a phrase the speaker never used.
- Every person's name, spelled the way they spell it, checked against their own site or profile rather than the caption.
- Every company, product and version number mentioned in the sections you kept.
- Every statistic, with its unit and its time period, listened to rather than read.
- Every quoted claim, played at its timestamp, because a misheard negative reverses the meaning.
- Every acronym, expanded once on first use in your own words rather than the speaker's shorthand.
Quote fidelity
A quote in your post should match what a reader hears when they click the timestamp. You may remove filler and false starts silently, since spoken English is full of both and nobody expects a court record. You may not reorder clauses, join two separate answers into one quote, or tidy a hedge into a clean claim. Cut from the middle with an ellipsis. Add a word for grammar inside brackets. If a speaker corrected themselves ten minutes later, quote the correction, not the slip.
Read the whole thing out loud
The fastest way to catch transcript residue is your own voice. Written-down speech has a particular rhythm: too many clauses, sentences that restart in the middle, connective tissue that only worked because a face was attached to it. Anything you stumble over while reading aloud was spoken rather than written, and it needs rewriting as prose. This usually takes 20 minutes and finds more problems than any checker.
Before you publish
- Embed the episode near the top, and link individual quotes to their timestamps so a reader can check you in one click.
- Credit the show and the guest by name inside the first hundred words.
- Write the title and meta description for a reader in a results page, not for the transcript's vocabulary.
- Link to two or three related posts you have already published, and out to the show's own site.
- Add Article structured data with the real author and the real publish date.
- Set a canonical URL, especially if the piece also goes to a newsletter or a syndication partner.
Internal links are the item everyone leaves for later and never does. Decide before publishing where this post sits on your site: which existing piece it supports, which one a new reader should follow for background, and which older post now needs a link pointing at this one. Two or three each way is plenty. A post built from an episode usually pairs well with something more procedural you have already written, because the conversation supplies the argument and the older piece supplies the steps.
People ask whether to publish the full transcript as its own page for search traffic. Our answer is usually no. A page whose entire body is machine-produced timed text has nothing on it that a person wrote, and search guidance has been explicit for years about rewarding content made for people rather than made for ranking. If you have a genuine accessibility or reference reason to publish the transcript, publish it, label it clearly as a transcript, and do not present it as an article.
The question is not whether a machine wrote part of it. The question is whether anyone useful read it before it went out.
YouTubeScribe publishing notes
What we stopped doing, and the sequence we run now
Four things we stopped
We stopped asking a model for a full rewrite in the host's voice. It drifted. Three paragraphs in, the voice had become a generic warm podcast voice that belonged to nobody, and worse, the quotes stopped matching the audio, because a rewrite treats quoted speech as text to improve. Anything inside quotation marks now comes straight from the timestamped file and is never regenerated, ever, for any reason.
We stopped stripping timestamps before editing. Clean text is nicer to draft in, and then you need to check one number and there is no route back to the line without reading 15,000 words with your eyes. Timestamps now stay in the working file until the final proofread, and they come out in the last five minutes.
We stopped chasing length. An early post ran to 3,400 words because the episode genuinely had that much in it. Nobody finished it. The version we would write today is 1,600 words with the episode embedded for anyone who wants the full hour. Length is not a proxy for value when the source is a conversation, because conversations are padded by design.
We stopped letting the generated draft write the opening. The introduction is the only part every reader reads, and a generated one always opens by describing the episode instead of making a claim. We now write the first two paragraphs by hand before looking at any draft, which also forces us to decide what the post is about before a tool decides for us.
How long this takes, stage by stage
| Stage | Time for a 90-minute episode | What you produce | What goes wrong |
|---|---|---|---|
| Get the transcript | 2 to 5 minutes | A timestamped file of 12,000 to 15,000 words, plus the chapter list if one exists | The video has no caption track, and the method stops there |
| Build the spine | 30 to 45 minutes | Five or six headings, one sentence each, one timestamped quote per heading | Keeping eleven sections because the episode had eleven chapters |
| Draft | 20 to 40 minutes | A 1,200 to 2,000 word draft with your own opening and ending already written | Pasting generated paragraphs that describe the episode instead of arguing something |
| Edit against the audio | 45 to 90 minutes | Checked names, numbers and quotes, and prose that reads as written rather than spoken | Skipping the audio check and shipping a misspelled guest surname |
| Publish | 15 to 20 minutes | Embed, timestamped quote links, credits, internal links, structured data | Publishing the raw transcript as a second thin page for extra keywords |
That totals three to four hours for an episode you have not heard, and under two if you were on the call and already know where the good parts are. Your first one will take longer than the table says. Your fifth will take less, mostly because you stop rescuing sections that were never going to work.
The sequence, start to finish
- Open the episode and click CC. If no captions appear, stop and get a transcript another way before planning anything.
- Paste the URL into the Transcript Generator and download the timestamped file. Keep it open for the rest of the job.
- Run the Chapters Extractor. If the creator wrote chapters, save the list. If not, note that you are building the structure yourself.
- Read the transcript once at speed, marking the first appearance of each new claim. Do not edit anything on this pass.
- Write the spine: five or six headings, with one sentence under each saying what the reader learns there.
- Choose one quote per heading, four lines at most, and write the timestamp beside every one of them.
- Delete every heading whose sentence duplicates another. Six sections becomes four more often than you expect.
- Write the opening two paragraphs by hand, before you generate anything at all.
- Run the Video to Blog Post tool for the middle sections, place its draft beside your spine, and keep only what matches.
- Write the bridges, the context and the counterpoint yourself, aiming for at least 40 percent of the finished post being material that is not in the transcript.
- Write the ending. It says what changes for the reader, not what the episode covered.
- Check every name, number and acronym against the audio at its timestamp, using the show's own site for spellings.
- Read the whole post out loud and rewrite every sentence you stumble over.
- Embed the episode, link the quotes to their timestamps, credit the show and the guest, and add internal links and Article structured data.
- Strip the working timestamps out of the body text, publish, and keep the transcript file for corrections later.
An illustrative example
Illustrative, not measured. A 92-minute founder interview transcribes to about 13,800 words with eight chapters. The spine cuts it to five sections: two chapters turn out to be the same argument in different clothes and merge, one is a sponsor read and disappears. The finished post runs 1,650 words and carries four quotes totalling roughly 180 words, which means about nine tenths of the published text is written rather than transcribed. The episode gets embedded twice, once at the top and once beside the section where the best story lives. Treat these numbers as a worked example of the shape, not as a measured result from a specific post.
The method is unglamorous, and that is the point of it. Tools get you a transcript in seconds and a structured draft in a minute, and both are genuinely useful. The three hours afterwards are the reason your post is worth reading, the reason the guest is happy to be quoted in it, and the reason anyone comes back for the next one.
Questions people ask
How do I turn a YouTube video into a blog post?
Download the video's caption track, cut it into a spine of five or six sections, draft from that spine, then edit against the audio. The transcript is raw material rather than a draft. For a 90-minute podcast you are reducing 12,000 to 15,000 words down to 1,200 to 2,000, which means the real work is selection and editing, not generation. A tool handles the first two minutes of that and none of the last two hours.
How long should the finished post be?
For a 90-minute episode, aim for 1,200 to 2,000 words. That is roughly a tenth of the transcript. Longer posts from conversational sources almost always carry padding, because conversation is padded by design and cutting it is what you are being paid for. If the episode genuinely holds more than one article, split it into two focused posts rather than publishing one long one that nobody finishes.
Do I need permission to repurpose someone else's podcast?
A public transcript is not a licence to republish the episode as your article. Quoting fairly with attribution and a link is normal editorial practice. Reproducing the bulk of someone's episode as your own content is not, whatever the transcript's availability suggests. Name the show and the guest, link the episode, keep quotes short and clearly marked, and if the show is yours or your employer's, disclose that in the post.
Which kinds of video repurpose well?
Anything where the information lives in the speech: webinars, tutorials, interviews, lectures, sermons, trainings and panels. The test is whether a listener with their eyes closed would lose anything. Vlogs, reactions and gameplay repurpose poorly because the value is visual and the speech is filler around it. Screencasts sit in the middle: the words work, but you will need to capture screenshots the caption track cannot give you.
Can I just publish the transcript?
You can, but it will not perform and it is not an article. A transcript is timed text with no structure, no context and no argument, and search guidance has been consistent about rewarding content made for people rather than for rankings. If you have an accessibility or reference reason to publish the full transcript, publish it, label it plainly as a transcript, and keep it separate from the piece you wrote.
What is the fastest workflow for a 90-minute episode?
Roughly three hours total: five minutes to pull the transcript and chapters, 45 minutes to build the spine, 30 minutes to draft, an hour to edit against the audio, and 20 minutes to publish. The order matters more than the speed. Building the spine before drafting is what stops you rescuing sections that should have been cut, which is where the time actually disappears on a first attempt.
Should I use the creator's chapters as my headings?
Use them as your starting outline, then cut and rename. Chapters written by the creator reflect how the person in the room thought the hour divided, which is better information than a model inferring structure from word patterns. But chapters serve viewers scrubbing a timeline, not readers scanning a page. Merge duplicates, drop sponsor segments, and rewrite each heading so it names an idea rather than a segment.
TXT, SRT or VTT: which do I download?
Take SRT for this workflow, because you need timestamps during editing to jump back to the audio and check names, figures and quotes. TXT is fine if you only want to read the episode. Both are free on YouTubeScribe, with VTT available on Pro for anyone who needs it for players or captioning pipelines. The first transcript each day is free and nothing is stored on the free tier.
How much of the post has to be original writing?
Common guidance puts the floor at 40 percent original material, and in practice a good repurposed post ends up higher. Your introduction, conclusion, context, examples and analysis are the original part. The transcript supplies quotes and the raw sequence of ideas. If you strip out the quotes and what remains does not stand as a piece of writing, you have published a summary of an episode rather than an article about it.
Should I embed the video, and where?
Yes, near the top, and a second time beside the section carrying the best material. Link individual quotes to their timestamps as well, so a reader who doubts a quote can verify it in one click. Embedding is good for the reader, good for the show, and good for your credibility, since it signals that you are writing about the episode rather than quietly replacing it.
The video has no captions. What now?
This method stops there, because we read a caption track and never run speech recognition on audio. Two options remain. Ask the show for their own transcript, since most podcast hosts generate one for the feed and creators often share it. Or run the audio through a speech recognition service and accept the error rate, which will show up first in guest surnames, company names and any technical term the speaker uses as shorthand.
The transcript has no punctuation or speaker labels. How do I fix it?
Expect this from automatic captions, since they infer sentence boundaries from pauses and carry no speaker information. Add speaker labels by hand for the passages you plan to quote, working from the audio at the timestamps rather than guessing from the text. Do not label the whole file. You only need attribution for the four or five quotes that survive the spine, which takes minutes rather than an afternoon.
Names, numbers and technical terms come out wrong.
That is the expected failure mode, and it is why the audio check exists. Automatic captions are commonly reported at 85 to 95 percent accuracy, which on 14,000 words is hundreds of wrong words concentrated in proper nouns, statistics and jargon. Check every name against the person's own site or profile, every figure against the audio at its timestamp, and every acronym against how the speaker actually used it.
The AI draft does not sound like the episode.
It will not, and asking it to is the mistake we made. We tried full rewrites in the host's voice, and the voice drifted into something generic within three paragraphs while the quotes quietly stopped matching the audio. Publish the piece in your own voice as your write-up of the episode, with the speakers quoted verbatim from the timestamped file. Anything inside quotation marks should never be regenerated.
The post reads like a summary and nobody finishes it.
Usually the opening is the problem. A generated draft opens by describing the episode, and readers leave because a description of a conversation is not a reason to keep reading. Write the first two paragraphs by hand before generating anything, starting with the single most interesting claim in your own words. Then check that every section makes a point rather than reporting that a topic came up.
Sources
- YouTube Help. Use automatic captioning. support.google.com· Checked 17 August 2026.
- YouTube Help. Supported subtitle and caption files. support.google.com· Checked 17 August 2026.
- YouTube Help. Add subtitles and captions. support.google.com· Checked 17 August 2026.
- W3C. WebVTT: The Web Video Text Tracks Format. w3.org· Checked 17 August 2026.
- Google for Developers. YouTube Data API: Captions. developers.google.com· Checked 17 August 2026.
- YouTube Help. Add video chapters to your videos. support.google.com· Checked 17 August 2026.
- YouTube Help. View video transcripts. support.google.com· Checked 17 August 2026.
- Google Search Central. Creating helpful, reliable, people-first content. developers.google.com· Checked 17 August 2026.
- Google Search Central. Google Search's guidance about AI-generated content. developers.google.com· Checked 17 August 2026.
- Google Search Central. Spam policies for Google web search. developers.google.com· Checked 17 August 2026.
- Google Search Central. Article structured data. developers.google.com· Checked 17 August 2026.
- W3C. Web Content Accessibility Guidelines (WCAG) 2.2. w3.org· Checked 17 August 2026.
- MDN Web Docs. WebVTT API. developer.mozilla.org· Checked 17 August 2026.
- U.S. Copyright Office. More Information on Fair Use. copyright.gov· Checked 17 August 2026.
- Ryan Robinson. How to Turn a YouTube Video into a Blog Post. ryrob.com· Checked 17 August 2026.
- RightBlogger. YouTube to Blog Post tool. rightblogger.com· Checked 17 August 2026.
- YouTubeScribe. Publishing notes, 2024 to 2026.
Written by Faisal Ashfaq. Faisal Ashfaq maintains the YouTubeScribe caption extractor at Deeporax AI LTD in Oldham. He writes the under-the-hood notes on timed text.
Reviewed on 17 August 2026. If a line is wrong, email support@youtubescribe.com.