Captions include speech and non-speech sound. Subtitles usually carry dialogue only, often as a translation.
Also called: closed captions, CC, SDH
On YouTube both words get used for the same timed-text track. In broadcasting they are not the same thing.
Captions assume the viewer may not hear the audio. They include speaker labels, music cues, and sound effects. Subtitles assume the viewer can hear, and mainly carry dialogue. When the dialogue is translated, that track is still a subtitle track.
YouTube stores one or more caption tracks next to the video. A track is a list of cues. Each cue has a start time, an end time, and a line of text. The file you download as SRT or VTT is that list in a standard wrapper.
Automatic captions are generated from speech. Creator-uploaded captions are typed or reviewed by a person. Both are tracks. A tool that extracts a transcript is reading a track, not listening to the audio.
If you need a file for a deaf or hard-of-hearing audience, you want captions: speech plus the sounds that change the meaning. If you need a translation under a video the viewer can already hear, you want subtitles.
Most YouTube auto tracks sit in the middle. They capture speech and miss most non-speech sound. Treat them as a draft, not as a finished accessibility file.
Captions include speech and non-speech sound. Subtitles usually carry dialogue only, often as a translation.
On YouTube both words get used for the same timed-text track. In broadcasting they are not the same thing.
Start with transcript-generator, subtitle-downloader.