The coverage matrix: how far do the same Chinese words go?

Quick answer

Do the top N Chinese words cover every genre equally?

No. The HSK vocabulary through Level 4, about 3,181 words, reaches roughly 69.8 percent of TV subtitle words but only about 54.4 percent of news, 41.7 percent of a classic novel, and 20 percent of a classical text. Spoken dialogue leans on a few very common words, so the same study level goes much further there.

Last reviewed 24 July 2026 by the Hanzi Coverage editorial team

Most coverage tools report one number from one corpus. Drag the slider below to pick how many of the most common words you know, then watch the same vocabulary hit a different share of five reading registers at once. It is the honest answer to "how many words do I need," genre by genre.

HSK scheme
Focus register

The most common 3,181 words (HSK 3.0 through Level 4) reach 69.8% of TV subtitles, 54.4% of news and web text, 48.2% of Chinese Wikipedia, 41.7% of classic novels, and 20% of classical texts.

  • TV subtitles 69.8%
  • News and web 54.4%
  • Chinese Wikipedia 48.2%
  • Classic novels 41.7%
  • Classical texts 20%

That is a 49.8-point gap between the easiest register (subtitles) and the hardest (classical) at the same study level.

The shape of the spread

The slider gives you one level at a time. This is the whole curve: cumulative coverage against words known, one line per register. The lines fan apart as your vocabulary grows, so the last thousand words you learn buy far more subtitle coverage than classical coverage. It is the same data as the tables below, drawn.

HSK words known versus coverage, by reading register Line chart: cumulative HSK 3.0 coverage against words known, one line per reading register. At the full vocabulary the same words reach 77.5% of subtitles but only 34% of classical text. 0% 20% 40% 60% 80% 02k4k6k8k10k cumulative HSK words known Subtitles 77.5% News 69.4% Wikipedia 61.6% Novels 54.4% Classical 34%
Same words, five registers. Every point is the cumulative share of running Chinese word tokens a learner recognises at that HSK milestone, from the same open corpora behind the coverage matrix. Token coverage is recognition, not comprehension.
Cumulative HSK 3.0 coverage by register
Words knownSubtitlesNewsWikipediaNovelsClassical
506 50.4%27.9%24.3%27.9%9.3%
1,256 60.7%37.8%33.3%34.9%13.3%
2,209 66.8%48.6%42.4%38.8%16.9%
3,181 69.8%54.4%48.2%41.7%20%
4,240 72%58.4%51.9%44.7%22.4%
5,363 73.9%62.4%55.4%48.2%25.3%
10,969 77.5%69.4%61.6%54.4%34%

The full coverage matrix (HSK 3.0)

Each cell is the share of running Chinese word tokens in that register recognised by a learner who knows every HSK word up to and including that level. Read across a register row to see how far each study milestone stretches into that kind of text.

Share of each register reached by the top-N HSK 3.0 vocabulary
Register HSK 1506 wordsHSK 21,256 wordsHSK 32,209 wordsHSK 43,181 wordsHSK 54,240 wordsHSK 65,363 wordsHSK 7 to 910,969 words
TV subtitles 50.4%60.7%66.8%69.8%72%73.9%77.5%
News and web 27.9%37.8%48.6%54.4%58.4%62.4%69.4%
Chinese Wikipedia 24.3%33.3%42.4%48.2%51.9%55.4%61.6%
Classic novels 27.9%34.9%38.8%41.7%44.7%48.2%54.4%
Classical texts 9.3%13.3%16.9%20%22.4%25.3%34%

The full coverage matrix (HSK 2.0)

The same computation against the older six-level HSK 2.0 word lists, for learners studying to that standard.

Share of each register reached by the top-N HSK 2.0 vocabulary
Register HSK 1150 wordsHSK 2297 wordsHSK 3595 wordsHSK 41,193 wordsHSK 52,491 wordsHSK 64,991 words
TV subtitles 37.4%46.3%52.5%57.8%61.6%64.5%
News and web 19.2%25.4%30.7%38.2%46.3%51.7%
Chinese Wikipedia 17.7%22.3%27.1%34.3%41.1%45.1%
Classic novels 18.6%23.7%27.3%31.4%34.5%36.6%
Classical texts 4.7%8.2%10.6%16.1%18.5%20.6%

Why one number is never enough

Word frequency in Chinese is steeply skewed, so the first thousand words you learn do far more work than the next thousand. That is the whole case for learning common words in order. But how much work they do depends on what you read: dialogue reuses a tiny core vocabulary, while news, reference prose, and literature spread probability across far more words and proper nouns. A single subtitle-corpus figure quietly overstates how ready you are for the news or a novel. This matrix shows the gap directly, so you can plan against the genre you actually want to read.

Honest caveats

For each register, the cumulative share of running Chinese-character word tokens covered by knowing every HSK word up to and including a given level. This is coverage BY THE HSK VOCABULARY, not by the corpus's own top-N words, so it is lower than a raw frequency-rank coverage figure: HSK ordering is not identical to any single corpus's frequency ordering, and proper nouns are never in the HSK list. Denominator per register is every Chinese-character-only word token in that corpus. Latin and punctuation tokens are excluded. Matching is against the simplified HSK surface form only; the two Project Gutenberg registers are converted traditional-to-simplified before counting so the same simplified match applies to every register. Corpora: film and TV subtitles (OpenSubtitles 2018, CC BY-SA 4.0), a news and web crawl and Chinese Wikipedia (Leipzig Corpora Collection, derived statistics only), and public-domain classic and classical works (Project Gutenberg). Full attribution and rebuild method: coverage by genre.

Next

Paste your own text into the coverage analyzer to see your recognition at each HSK level for that exact passage, use the difficulty grader to find the level a text needs, or read how many words you need to read Chinese and the best order to learn them. For the full cross-register comparison as a sourced data study, see the statistics page.