The coverage matrix: how far do the same Chinese words go?
Quick answer
Do the top N Chinese words cover every genre equally?
No. The HSK vocabulary through Level 4, about 3,181 words, reaches roughly 69.8 percent of TV subtitle words but only about 54.4 percent of news, 41.7 percent of a classic novel, and 20 percent of a classical text. Spoken dialogue leans on a few very common words, so the same study level goes much further there.
Most coverage tools report one number from one corpus. Drag the slider below to pick how many of the most common words you know, then watch the same vocabulary hit a different share of five reading registers at once. It is the honest answer to "how many words do I need," genre by genre.
The most common 3,181 words (HSK 3.0 through Level 4) reach 69.8% of TV subtitles, 54.4% of news and web text, 48.2% of Chinese Wikipedia, 41.7% of classic novels, and 20% of classical texts.
That is a 49.8-point gap between the easiest register (subtitles) and the hardest (classical) at the same study level.
The shape of the spread
The slider gives you one level at a time. This is the whole curve: cumulative coverage against words known, one line per register. The lines fan apart as your vocabulary grows, so the last thousand words you learn buy far more subtitle coverage than classical coverage. It is the same data as the tables below, drawn.
| Words known | Subtitles | News | Wikipedia | Novels | Classical |
|---|---|---|---|---|---|
| 506 | 50.4% | 27.9% | 24.3% | 27.9% | 9.3% |
| 1,256 | 60.7% | 37.8% | 33.3% | 34.9% | 13.3% |
| 2,209 | 66.8% | 48.6% | 42.4% | 38.8% | 16.9% |
| 3,181 | 69.8% | 54.4% | 48.2% | 41.7% | 20% |
| 4,240 | 72% | 58.4% | 51.9% | 44.7% | 22.4% |
| 5,363 | 73.9% | 62.4% | 55.4% | 48.2% | 25.3% |
| 10,969 | 77.5% | 69.4% | 61.6% | 54.4% | 34% |
The full coverage matrix (HSK 3.0)
Each cell is the share of running Chinese word tokens in that register recognised by a learner who knows every HSK word up to and including that level. Read across a register row to see how far each study milestone stretches into that kind of text.
| Register | HSK 1506 words | HSK 21,256 words | HSK 32,209 words | HSK 43,181 words | HSK 54,240 words | HSK 65,363 words | HSK 7 to 910,969 words |
|---|---|---|---|---|---|---|---|
| TV subtitles | 50.4% | 60.7% | 66.8% | 69.8% | 72% | 73.9% | 77.5% |
| News and web | 27.9% | 37.8% | 48.6% | 54.4% | 58.4% | 62.4% | 69.4% |
| Chinese Wikipedia | 24.3% | 33.3% | 42.4% | 48.2% | 51.9% | 55.4% | 61.6% |
| Classic novels | 27.9% | 34.9% | 38.8% | 41.7% | 44.7% | 48.2% | 54.4% |
| Classical texts | 9.3% | 13.3% | 16.9% | 20% | 22.4% | 25.3% | 34% |
The full coverage matrix (HSK 2.0)
The same computation against the older six-level HSK 2.0 word lists, for learners studying to that standard.
| Register | HSK 1150 words | HSK 2297 words | HSK 3595 words | HSK 41,193 words | HSK 52,491 words | HSK 64,991 words |
|---|---|---|---|---|---|---|
| TV subtitles | 37.4% | 46.3% | 52.5% | 57.8% | 61.6% | 64.5% |
| News and web | 19.2% | 25.4% | 30.7% | 38.2% | 46.3% | 51.7% |
| Chinese Wikipedia | 17.7% | 22.3% | 27.1% | 34.3% | 41.1% | 45.1% |
| Classic novels | 18.6% | 23.7% | 27.3% | 31.4% | 34.5% | 36.6% |
| Classical texts | 4.7% | 8.2% | 10.6% | 16.1% | 18.5% | 20.6% |
Why one number is never enough
Word frequency in Chinese is steeply skewed, so the first thousand words you learn do far more work than the next thousand. That is the whole case for learning common words in order. But how much work they do depends on what you read: dialogue reuses a tiny core vocabulary, while news, reference prose, and literature spread probability across far more words and proper nouns. A single subtitle-corpus figure quietly overstates how ready you are for the news or a novel. This matrix shows the gap directly, so you can plan against the genre you actually want to read.
Honest caveats
- Coverage is recognition, not comprehension. Recognising a share of the words is a planning estimate, not a claim you will understand every sentence. The comfortable-reading thresholds people cite (around 95 to 98 percent) are about known-word coverage, and even those assume you can infer the rest.
- This is coverage by the HSK vocabulary, not by a corpus top-N. It is lower than the usual "1000 words gives you X percent" figures, which count a text's own most frequent words. HSK order is not the same as any one text's frequency order, and proper nouns never appear on the HSK list.
- The classical figure understates the difficulty by design. Classical Chinese uses a different lexicon and grammar, and a shared character usually carries a different classical meaning, so even the low number here is optimistic.
- The subtitles column is the validation. It reproduces our long-standing single-register coverage figures cell for cell, so the four other columns are the same method applied to their own corpus, not a different measure.
- This is a static, replicable computed asset, never proprietary data. Every corpus is CC-licensed or public domain and fully attributed. See the full per-register sources and method.
For each register, the cumulative share of running Chinese-character word tokens covered by knowing every HSK word up to and including a given level. This is coverage BY THE HSK VOCABULARY, not by the corpus's own top-N words, so it is lower than a raw frequency-rank coverage figure: HSK ordering is not identical to any single corpus's frequency ordering, and proper nouns are never in the HSK list. Denominator per register is every Chinese-character-only word token in that corpus. Latin and punctuation tokens are excluded. Matching is against the simplified HSK surface form only; the two Project Gutenberg registers are converted traditional-to-simplified before counting so the same simplified match applies to every register. Corpora: film and TV subtitles (OpenSubtitles 2018, CC BY-SA 4.0), a news and web crawl and Chinese Wikipedia (Leipzig Corpora Collection, derived statistics only), and public-domain classic and classical works (Project Gutenberg). Full attribution and rebuild method: coverage by genre.
Next
Paste your own text into the coverage analyzer to see your recognition at each HSK level for that exact passage, use the difficulty grader to find the level a text needs, or read how many words you need to read Chinese and the best order to learn them. For the full cross-register comparison as a sourced data study, see the statistics page.