There is a gap at the centre of most international content strategies, and it is measurable.
English accounts for 49.5% of all web content whose language is known — the next largest, Spanish, sits at 6.0%. Yet only around 1.52 billion people speak English at all, native or otherwise: roughly one person in five. Nearly half the internet is written in a language that four out of five people on the planet do not speak.
Video makes that gap wider, not narrower. A written page can at least be run through a browser translation. A video without a text layer is a sealed box: search engines cannot read it, translation tools cannot reach it, and a viewer who does not speak the language has nothing to work with.
Multilingual video transcription is what opens the box. It is the step that turns speech into structured, timecoded, translatable text — and everything else in a localisation strategy is built on top of it. This guide explains how the process works, where it fails, and how to commission it so the output is worth having.

The commercial case, in numbers
The reluctance to invest in localisation usually rests on an assumption that international audiences will make do with English. The research does not support it.
RWS surveyed 6,500 consumers globally in 2023 and found:
- More than 81% would not purchase from a brand that does not offer support in their own language
- 89% believe they should have the option to interact with companies in their preferred language
- 93% say companies need to communicate in preferred languages across all channels, at all times
- 44% are actively frustrated by the dominance of English online
On the platform side, the effect is now visible in the numbers. When YouTube began testing multi-language audio, early participants reported over 25% of watch time coming from dubbed versions rather than the original. Jamie Oliver’s channel tripled its views after adding Spanish, Portuguese and Hindi tracks to existing content. Roughly 80% of YouTube’s content is already in languages other than English.
The audience is there. The content simply is not reaching it.
What multilingual video transcription actually means
The terminology in this space is used loosely, which leads to organisations buying the wrong thing. Four distinct services sit in a chain:
1. Transcription converts spoken audio into text in the same language as the recording. With timecodes and speaker identification, this becomes the master asset. See video transcription and audio transcription.
2. Translation converts that text into another language. Applied to a transcript, this is transcription and translation work; applied directly to a recording, audio translation.
3. Subtitling takes translated text and fits it to the screen — condensed to a readable rate, timed to the frame, broken across lines that respect grammar. This is a distinct craft from translation, handled through our captioning and subtitling services.
4. Dubbing or voice-over takes a translated script and performs it, matched to the original timing. See voice over services.
“Multilingual video transcription” properly describes steps one and two working together: an accurate transcript of the source audio, plus text versions in every target language, timecoded and ready to drive whatever comes next. If a supplier offers you “multilingual transcription” and delivers untimed prose, you will pay again to make it usable.
There is also genuine multilingual transcription in the narrower sense: recordings where several languages appear in the same audio — international conference calls, cross-border interviews, multinational focus groups. Those require a transcriber fluent in each language present, working with code-switching mid-sentence. That is specialist work, and it is what our multilingual transcription services and language transcription teams are built for.
The workflow that scales
Localisation gets expensive when it is done in parallel — a separate pipeline per language, each starting from the raw video. It gets efficient when it is done as a hub and spoke, with one master transcript at the centre.
Step 1 — Produce a verified source transcript. Full accuracy, speaker identification, timecodes at speaker changes or fixed intervals, and a terminology list agreed with you: product names, personal names, technical terms, acronyms.
Step 2 — Sign off the master. This is the single highest-leverage moment in the entire project, and it is the one most often skipped. Every error still present here will be translated faithfully into every target language. One mistake in the source becomes twelve mistakes across twelve markets, and each has to be found and corrected separately.
Step 3 — Translate from the approved text. Human translators work from clean, contextualised, speaker-attributed text rather than guessing at audio. Terminology stays consistent because it was fixed in step one.
Step 4 — Adapt per language and per output. The same translated content is then shaped for its destination: subtitle files with reading-rate limits and line breaks, captioning or SDH tracks for accessibility, dubbing scripts with timing notes, or readable transcript pages for the web.
Step 5 — Publish the text layer as well as the video. Translated transcripts are indexable pages. Video files are not.
Do it in this order and adding a thirteenth language is an incremental cost. Do it the other way round — one bespoke project per market — and it is a thirteenth project.
Where the machine-only route breaks
Automated pipelines chain speech recognition to machine translation. On clean, single-speaker, scripted audio in a major language, the result can be serviceable. On the material organisations actually hold, it fails in specific and predictable ways.
Errors compound multiplicatively. If recognition is 90% accurate and translation is 90% accurate, the output is not 90% accurate — it is closer to 81%, and the second stage cannot detect that it is translating a mistake. It translates the error confidently and fluently, which makes it harder to spot than a garbled word would be.
Named entities have no fallback. People’s names, company names, place names, product names and acronyms are exactly what recognition gets wrong and what translation has no basis for repairing. A misheard surname propagates into every market.
Specialist vocabulary collapses. Clinical terminology, drug names, legal citations, engineering and financial terms. This is why sector experience matters more than raw language capability in healthcare, legal and finance work.
Speaker attribution disappears. Multi-party recordings — conference calls, Zoom meetings, focus groups — depend on knowing who said what. Automated diarisation degrades sharply with crosstalk and accented speech, and a translation cannot restore attribution it never received.
Register and idiom are lost. Machine translation renders words, not tone. Humour, hedging, formality, sarcasm, politeness levels — these carry meaning in the original and are exactly what a global audience judges you on.
Text expansion breaks the screen. Translating English into Spanish, French, Portuguese or German typically lengthens the text by around 15–30%. A subtitle that fitted comfortably in English overflows, or gets cut. Someone has to make an editorial judgement about what to compress, and that decision requires understanding what the sentence is for.
Script and layout rules differ. Arabic and Hebrew run right-to-left. Chinese, Japanese and Korean break lines by character, not by word, and have their own punctuation conventions. Thai does not use spaces between words. These are not cosmetic details; get them wrong and the subtitle is unreadable.
We have written elsewhere about why human legal transcription still matters and whether AI will replace medical transcriptionists. The multilingual case is the same argument with the stakes multiplied by the number of markets.
Choosing which languages to start with
Most organisations should not localise into twelve languages at once. Sequence it with data.
Start with your own analytics. Look at where traffic already comes from despite the language barrier, which markets have high impressions but low engagement, and which countries your sales pipeline already touches. Demand that exists in spite of English-only content is the safest bet.
Then weigh the web-content picture. The current distribution of content languages online tells you where the competitive space is:
| Language | Share of web content | Speakers worldwide |
|---|---|---|
| English | 49.5% | 1.52bn |
| Spanish | 6.0% | 560m |
| German | 5.9% | ~135m |
| Japanese | 4.9% | ~125m |
| French | 4.5% | 321m |
| Portuguese | 4.1% | 264m |
| Russian | 3.4% | 255m |
| Italian | 2.8% | ~65m |
Web content share: W3Techs, September 2026. Speaker totals include second-language speakers.
The interesting cases are where speaker numbers are large relative to available content — Hindi, Arabic, Mandarin and Portuguese all have far more speakers than the volume of content addressing them.
For most UK and European businesses, the practical starting set is: Spanish, French, German and Italian, extending into Portuguese, Arabic, Chinese, Japanese, Korean, Russian and Turkish as demand appears. We also work regularly into Dutch, Nordic and other European markets, and support clients operating in Germany, France, Spain and the United States.
The search advantage nobody claims
There is a second return on multilingual transcription that has nothing to do with accessibility or courtesy: search engines index text, not speech.
A video published with no text layer competes on its title and description alone. The same video published with a full transcript gains a page of substantive, keyword-rich, naturally-written content. Publish that transcript in six languages and you have six indexable pages where you previously had none — each one able to rank in its own market, in its own language, against far less competition than the English original faces.
Practical points worth getting right:
- Publish the translated transcript as on-page text, not as a downloadable file only
- Use
hreflangannotations so search engines understand which version serves which market - Translate the metadata too — title, description, chapter markers, thumbnails and tags
- Where the platform supports it, upload proper subtitle tracks rather than relying on auto-translation of auto-captions
This applies to YouTube content, podcasts, webinars and any media library. It is the reason transcription frequently pays for itself before the accessibility or localisation benefits are counted at all.
Where multilingual transcription earns its keep
Corporate and internal communications. All-hands recordings, policy briefings and compliance training delivered to a workforce spread across multiple countries. See corporate transcription and business transcription.
E-learning and higher education. Recorded lectures and course materials for international cohorts, where a transcript also serves students studying in a second language. See educational transcription and academic transcription.
Market and academic research. Multi-country studies where analysts need every interview in a common working language while retaining the original for verification. See market research transcription and research transcription.
Media and broadcast. Programme localisation, archive exploitation and access services. See media transcription and television transcription.
Regulated and evidential settings. Cross-border investigations, arbitration and multilingual proceedings, where an accurate original and an accurate translation are both part of the record. See international arbitration transcription and certified transcription.
There is a compliance dimension in Europe too. The European Accessibility Act has applied since June 2025, and organisations serving EU markets face accessibility requirements for audiovisual content — obligations that a multilingual text layer helps discharge rather than duplicate.
A checklist for commissioning the work
- Specify the source transcript style — verbatim, intelligent verbatim or clean read. See verbatim transcription.
- Supply a glossary before work starts. Names, products, acronyms, preferred terminology per language. This single document prevents most downstream inconsistency.
- Ask for timecodes on the master, at speaker changes at minimum.
- Confirm who signs off the source text, and build that approval into the schedule before translation begins.
- Name the output formats per language — SRT, VTT, SCC, TTML, DOCX — and the platform each is going to.
- Decide subtitles versus dubbing per market, not globally. They serve different audiences and cost very differently.
- Ask how multi-speaker and mixed-language audio is handled, and by whom.
- Check data handling. Where is audio processed, who has access, are subcontractors used, how long are files retained, and is your content used to train models? Our data security commitment sets out our position.
- Plan the text layer’s publication, not just the video’s.
Work with a human team across twelve countries
Imperial Intelligence provides 100% human transcription, translation, captioning and subtitling. We do not run your recordings through automated recognition and pass the output on, because the compounding-error problem described above is not theoretical — it is what we are hired to fix.
We work from hubs in London and New York, across more than twelve countries, with GDPR-aligned workflows, speaker-identified and timecoded masters, and translators working in their own native languages. If your video library is currently reaching one audience in one language, we can tell you exactly what it would take to reach the rest.
- Explore multilingual transcription services and transcription with translation
- See video transcription and captioning and subtitling
- View pricing, or send us one file as a free trial
- Contact us — UK: 020 8146 3222 | US: +1 347 295 4572 | info@imperialintelligence.co.uk







