Flat diagram of five stacked bars with decreasing amber fill, representing how the same vocabulary covers less and less across reading registers.

Why Is 95% Coverage in Subtitles Not 95% in the News?

Quick answer

Why Is 95% Coverage in Subtitles Not 95% in the News?

Coverage is genre-specific, not a single number. The same complete HSK vocabulary that recognises about 78 percent of TV-subtitle words covers only 69 percent of news, 54 percent of a classic novel, and 34 percent of a classical text. So 95 percent coverage in one register is nowhere near 95 percent in another.

Last reviewed 2026-08-15 by the Hanzi Coverage editorial team

Almost every “how many words to read Chinese” answer quotes one number from one corpus, usually film-subtitle dialogue, and then implies it holds for everything you might read. It does not. Coverage is not a property of your vocabulary alone. It is a property of your vocabulary against a specific kind of text, and the same words you know cover wildly different shares of subtitles, news, encyclopedia articles, novels, and classical prose. This guide is the honest cross-genre picture, built from five separate corpora and one identical method, so the registers are directly comparable. It is the explainer behind the coverage-by-genre matrix.

The single fact that breaks the single number

Here is the whole point in one row. Take a reader who has learned the complete HSK 3.0 vocabulary, all 10,969 words through the level 7 to 9 band, and measure what share of running words that vocabulary recognises in each register:

RegisterWhat it isCoverage by full HSK 3.0Coverage at HSK 4 (3,181 words)
SubtitlesFilm and TV dialogue77.5%69.8%
NewsJournalism and web news69.4%54.4%
EncyclopedicWikipedia and reference61.6%48.2%
LiteratureClassic vernacular novels54.4%41.7%
ClassicalWenyanwen (文言文)34.0%20.0%

One reader, one vocabulary, and the coverage swings from 77.5 percent down to 34 percent purely because of what they chose to read. The person who feels fluent watching a subtitled drama and then bounces off a newspaper has not gotten worse at Chinese between the two. They have moved from a register where their vocabulary is dense to one where it is thin. That gap, not any change in the learner, is what a one-number answer hides. The figures come from the coverage-by-genre matrix; the subtitles column reproduces this site’s long-standing single-corpus figures exactly, which is how we validate the method.

Why the “95 percent rule” is a moving target

Reading-acquisition research is usually summarised as: you read comfortably at roughly 95 percent known-word coverage and independently at roughly 98 percent. Those thresholds, discussed with sources in the 95% and 98% coverage rules and the 98% research, honestly, are stated as a coverage of the text in front of you, whatever it is. That is exactly why they cannot be converted into a single vocabulary size. The vocabulary you need to reach 95 percent coverage is different for every register:

So “aim for 95 percent” is sound advice about a text, and useless as a word count until you say which register. Notice that in the by-the-HSK-vocabulary method above, no register reaches 95 percent even with the entire HSK list. That is deliberate and honest: the last stretch to comfortable reading is proper nouns, names, and specialised vocabulary that live outside the exam syllabus, so it is earned from reading itself, not from finishing HSK.

Why the registers differ so much

The ranking, spoken > news > encyclopedic > literary > classical, is not an accident of which corpora we picked. It falls out of how each register uses vocabulary:

A special case worth naming: social media is not simply “harder” or “easier.” Its plain sentences are close to the spoken register and land early, but its slang layer, yyds, 666, 绝绝子, homophone puns, is invisible to every coverage number because it is absent from HSK lists and from the segmenter’s dictionary. Coverage can measure four of our five registers cleanly; social media is the honest reminder that a percentage cannot capture everything.

Coverage is recognition, not comprehension

One caveat travels with every figure on this page. These numbers are the share of word tokens you recognise, not the share of meaning you understand. The unknown words are disproportionately the content words that carry the point of a sentence, and grammar, idiom, and reference sit on top of vocabulary. A 70 percent recognition figure on news means the missing 30 percent is largely the who, where, and what of the story, so recognition overstates comprehension. Treat the whole matrix as a map of where your vocabulary is dense and where it is thin, not as a comprehension score.

What to do with this

The practical move is to stop reading averages and measure the register you actually care about:

  1. Pick your real target and paste a genuine sample into the coverage analyzer. It segments the text, checks every word against HSK 2.0 and 3.0, and reports your coverage plus a ranked list of the words you are missing, most common first.
  2. Grade a candidate text with the difficulty grader to see the HSK level at which it crosses the comfortable-reading mark, then compare texts before you commit to one.
  3. Compare registers side by side in the coverage-by-genre matrix, so you can see precisely how much harder the front page is than the subtitles you practise with, and set expectations before you start.
  4. Learn the gap in frequency order. The best order to learn Chinese words explains why the common words you are missing pay off most, and spaced repetition, for example with Wordbrush, the app I built, keeps them from leaking back out. Every tool here is free and runs in your browser.

The bottom line: there is no single “how much Chinese” number, because coverage is genre-specific. Decide what you want to read, measure your coverage of a real sample of it, and let the register, not a blog’s round figure, tell you how far you have to go.

Frequently asked questions

Why does the same vocabulary cover so much less of the news than of subtitles?

Spoken dialogue reuses a small set of very common words, so a modest vocabulary covers most of it. News spreads meaning across far more distinct words and is packed with proper nouns and a formal reporting register that no HSK list contains. At full HSK 3.0 the same vocabulary covers about 78 percent of subtitle words but only about 69 percent of news words.

Does reaching 95 percent coverage mean the same amount of study for every genre?

No. The 95 percent comfort mark is a coverage of the specific text, so the vocabulary needed to reach it differs sharply by register: modest for dialogue, much larger for news, and possibly unreachable from any fixed wordlist for literary and classical prose, where names and rare vocabulary live outside the exam syllabus.

Are these coverage figures measured or estimated?

They are computed directly from five open, attributed corpora, film subtitles, a news crawl, Wikipedia, public-domain novels, and classical works, by one identical method. The subtitles column reproduces this site’s long-standing single-corpus figures cell for cell, which validates the method. Token coverage is still recognition, not full comprehension.