Ask three people in a marketing team what a subtitle is and you will get three answers. Ask a broadcaster, a university lecturer and a solicitor and you will get three more.
The confusion is understandable — all three formats turn speech into text — but the distinction matters commercially and legally. Order subtitles when you needed closed captions and your video will still be inaccessible to a deaf viewer. Commission captions when you needed a transcript and you will have a timecoded file nobody can read as a document. Assume a transcript satisfies your accessibility obligations and you may find it does not.
This guide sets out exactly what each format is, what it contains, which file types it uses, and how to work out which one you actually need.
The short answer
| Closed captions | Subtitles | Transcripts | |
|---|---|---|---|
| Assumes the viewer | Cannot hear the audio | Can hear, but does not understand the language | Is reading instead of watching or listening |
| Language | Same as the audio | Usually translated | Same as the audio (or translated separately) |
| Includes non-speech sound | Yes — [door slams], [laughter], [music] | No | Sometimes, if requested |
| Identifies speakers | Yes | Rarely | Yes |
| Timecoded | Yes, synchronised to the frame | Yes, synchronised to the frame | Optional |
| Can be switched off | Yes (closed) / No (open) | Yes, usually | N/A — separate document |
| Typical formats | SRT, VTT, SCC, TTML | SRT, VTT, ASS | DOCX, PDF, TXT |
| Primary purpose | Accessibility | Translation and comprehension | Record, reference, search |
If you remember only one thing: captions are for people who cannot hear; subtitles are for people who cannot understand the language; transcripts are for people who want to read.

What are closed captions?
Closed captions are a timecoded text version of everything meaningful in the audio, displayed in sync with the video and written on the assumption that the viewer cannot hear the soundtrack at all.
That last point is what separates captions from subtitles. Because the viewer is receiving no audio information, captions must carry more than dialogue. A properly produced caption file includes:
- All spoken dialogue, in the same language as the audio
- Speaker identification, so it is clear who is talking when the speaker is off-screen or when several people are involved
- Non-speech audio that carries meaning — [phone ringing], [applause], [engine starting], [indistinct shouting]
- Music and sound cues where they matter to the content — [sombre music], or song lyrics where relevant and licensed
- Manner of speech where it changes the meaning — [whispering], [sarcastic]
Captions also have to obey reading-rate limits. A caption that is technically accurate but flashes past in half a second has not communicated anything, so professional captioning involves decisions about line breaks, character counts per line and minimum display duration. This is craft work, and it is the reason automated caption tracks so often feel unusable even when the words are broadly right.
“Closed” versus “open” captions
The word closed refers to how the captions are delivered, not what they contain:
- Closed captions are supplied as a separate file (or a separate data track) that the player renders on demand. The viewer can turn them on and off, and search engines and platforms can read the text.
- Open captions are burned permanently into the video frame. They cannot be switched off, cannot be translated later and cannot be indexed — but they will always display, which makes them useful for social media feeds that autoplay muted, or where you cannot rely on the player supporting caption files.
Most organisations should default to closed captions and burn in open captions only for specific distribution channels.
What are subtitles?
Subtitles assume the opposite starting point: the viewer can hear the audio perfectly well but does not understand the language being spoken.
Because sound is available, subtitles carry dialogue only. There is no need to describe a door slamming — the viewer hears it. Subtitles are typically a translation of the spoken content into a target language, and that translation involves compression: spoken sentences are longer than what a viewer can comfortably read in the time available, so a subtitler condenses meaning rather than transcribing word for word.
Good subtitling is therefore closer to translation than to transcription. It requires judgement about idiom, register, cultural reference and humour, and it requires a subtitler who understands both languages well enough to know what can safely be cut. This is why we handle it as a distinct discipline within our captioning and subtitling services, supported by our transcription and translation and audio translation teams.
Subtitles can also be same-language, which is common in education and in content aimed at second-language audiences — the text supports comprehension rather than replacing sound.
SDH: the format that sits in between
SDH stands for Subtitles for the Deaf and Hard of Hearing, and it is where the two categories overlap. SDH is delivered through the subtitle track of a player, so it behaves technically like subtitles, but it contains caption-style information: speaker labels and non-speech audio descriptions.
You will encounter SDH on streaming platforms and Blu-ray, where the delivery pipeline has a subtitle track but no dedicated caption track. If a platform asks you for “SDH,” it wants caption content in a subtitle wrapper. Supplying plain translated subtitles instead is one of the most common accessibility failures in on-demand distribution.
What are transcripts?
A transcript is a complete text document of the audio, produced to be read rather than displayed over video. It has no obligation to be synchronised to the frame, and it is not constrained by reading rate or character limits — which means it can be more complete than any caption file.
Transcripts come in several styles, and the style you choose changes both the price and the usefulness of the result:
- Verbatim transcription captures everything: false starts, repetitions, stutters, filler words, interruptions, and often non-verbal sounds. It is the standard for evidential and research work, where how something was said matters as much as what was said. See our verbatim transcription services.
- Intelligent verbatim (sometimes called clean verbatim) removes filler and false starts without altering meaning or paraphrasing. This is the most commonly requested style for business, academic and media work.
- Edited or clean-read transcription tidies grammar and structure for publication.
- Time-stamped transcripts add timecodes at intervals or at each speaker change, so a reader can jump straight to the relevant point in the recording. This is the format that bridges most usefully into captioning.
Transcripts do work that captions cannot. They can be read at speed, quoted, annotated, coded, searched, redacted, translated, indexed by search engines and filed as a record. For a great deal of professional work they are the deliverable that matters:
- Legal transcription, court transcription and hearing transcription, where the transcript is the record
- Medical transcription and GP transcription, where it becomes part of the clinical documentation
- Market research and focus group transcription, where analysts code the text
- Academic and research transcription, where the transcript is the dataset
- Interview transcription and meeting minutes for business use
File formats: what to ask for
Getting the format right is half of a smooth delivery.
Caption and subtitle formats
- SRT (SubRip) — the universal workhorse. Plain text, simple timecodes, accepted almost everywhere. Minimal styling.
- VTT (WebVTT) — the web standard, used with HTML5 video and most modern platforms. Supports positioning and basic styling, and is the usual choice for a website.
- SCC (Scenarist Closed Caption) — the broadcast format, carrying CEA-608 caption data. Required by many broadcasters.
- TTML / DFXP / IMSC — XML-based, richer styling and positioning; common in streaming delivery specifications.
- ASS / SSA — heavy styling control, mostly used in specialist and fan-subtitling contexts.
- Burned-in (open captions) — delivered as a new video file rather than a text file.
Transcript formats
- DOCX — the default for professional work: editable, comment-friendly, template-compatible.
- PDF — for fixed, distributable records.
- TXT / CSV — for import into analysis software such as NVivo or Atlas.ti.
If you are commissioning work, name the format and the platform it is going to. “SRT for YouTube” and “SCC to broadcast spec” are very different jobs.
Which one do you actually need?
Work backwards from the person who will use the output.
You need closed captions if:
- You publish video or synchronised audio to a website, intranet, learning platform or social channel
- Deaf or hard-of-hearing viewers form part of your audience — which, for any public-facing content, they do
- You have accessibility obligations to meet (see below)
- You want your video content to be indexable and searchable
You need subtitles if:
- Your audience does not share the language of the audio
- You are localising content for international markets
- You are distributing through a platform that requires language subtitle tracks — often alongside SDH
You need a transcript if:
- The text is the record: legal proceedings, clinical notes, recorded statements, board minutes
- Someone will analyse, code or quote the content
- You want a readable page to accompany a podcast, webinar or video
- The source is audio-only, where a transcript is the appropriate text alternative
You need more than one — which is the usual answer. A recorded webinar published on a UK university site realistically needs closed captions on the player, a downloadable transcript on the page, and translated subtitles if the audience is international.
The production order that saves money
Here is the practical point most buyers miss: the transcript is the parent document.
An accurate, speaker-identified, time-stamped transcript is the raw material from which caption files, SDH and translated subtitles are all derived. Commission the transcript properly once, and every downstream format becomes a formatting and translation exercise rather than a fresh piece of work.
Do it the other way round — start with an auto-generated caption track and try to build a usable transcript out of it — and you pay twice: once for the original, and again for the correction pass. This is where our video transcription and audio transcription work usually begins.
Where accuracy stops being a preference
None of these formats is any use if the words are wrong, and this is where the distinction between automated output and human transcription becomes concrete rather than philosophical.
Automatic speech recognition performs respectably on a single clear speaker in a quiet room. It degrades sharply on exactly the material professional clients send us: overlapping speech in conference calls and focus groups, regional and non-native accents, clinical terminology and drug names, case citations, proper nouns, and recordings made in courtrooms, consulting rooms and on hand-held devices. We have written separately about whether AI will replace medical transcriptionists and about why human legal transcription still matters.
There is a compliance dimension too. Accessibility standards do not simply ask whether a caption file exists; they require captions to be accurate and synchronised, and require a text alternative for audio-only content. Under the Equality Act 2010, service providers and employers owe an anticipatory duty to make reasonable adjustments, and UK public sector bodies are held to WCAG level AA under the Public Sector Bodies (Websites and Mobile Applications) Accessibility Regulations 2018. A caption track that mishears a name or drops a clause does not discharge that duty — and in evidential settings, an inaccurate transcript creates a problem of its own, as we discuss in our piece on recorded statements as evidence.
Five mistakes worth avoiding
- Supplying translated subtitles where SDH was required. Deaf viewers get dialogue with no speaker labels and no sound cues, and the platform’s accessibility requirement is unmet.
- Treating a transcript as a caption substitute for video. A transcript is the correct text alternative for audio-only content; synchronised media needs captions.
- Burning in captions too early. Open captions cannot be switched off, translated or indexed. Keep a clean master.
- Ignoring reading rate. Accurate text at 400 words per minute is unreadable text.
- Publishing auto-captions unchecked. They are a starting point, not a deliverable — and on public-facing content they are a documented, dated record of what you published.
Frequently asked questions
Are closed captions and subtitles the same thing? No. Closed captions are in the same language as the audio and include speaker identification and non-speech sounds, because they are written for viewers who cannot hear. Subtitles are usually a translation and contain dialogue only, because they are written for viewers who can hear but do not understand the language.
Is a transcript enough for accessibility? For audio-only content, a transcript is the appropriate text alternative. For video and other synchronised media, captions are required, and a transcript is a valuable addition rather than a replacement.
What is the difference between CC and SDH? Both carry caption-style information. The difference is delivery: CC is a dedicated caption track that the player renders, while SDH is supplied through the subtitle track, which is how most streaming and disc-based platforms are built.
Should I choose SRT or VTT? VTT for websites and modern web players; SRT where you need the broadest possible compatibility; SCC or TTML where a broadcaster or streaming platform specifies it.
Can you produce all three from one recording? Yes, and that is the efficient route. We produce the transcript first, then derive caption files, SDH and translated subtitles from the approved text.
Get the right format, first time
Imperial Intelligence provides 100% human transcription, captioning, subtitling and translation. We do not publish machine output and call it a deliverable, because our clients’ work is used as record, as evidence and as clinical documentation — contexts in which “mostly right” is not a standard.
We work from hubs in London and New York across more than twelve countries, with GDPR-aligned workflows and a documented data security commitment.
If you are not certain which format your project needs, tell us where the content is going and we will tell you what to order.
- See our captioning and subtitling services
- View pricing or request a free trial
- Contact us — UK: 020 8146 3222 | US: +1 347 295 4572 | info@imperialintelligence.co.uk
Related services
Audio transcription · Video transcription · Subtitling services near you · Multilingual transcription · Text translation · Media transcription · Educational transcription · Zoom transcription · YouTube video transcription · Certified transcription · Voice over services















