How much Chinese do you need to read the news, novels, or classical texts?
Quick answer
Does the genre change how much Chinese I need?
Yes, sharply. The HSK 4 vocabulary that recognises about 69.8 percent of TV subtitle words covers only about 54.4 percent of news text, 41.7 percent of a classic novel, and 20 percent of a classical text. Spoken dialogue leans on a few very common words; written and older registers spread across far more vocabulary, so the same study level goes much less far.
Most coverage tools report one number from one corpus. This matrix computes the same thing, the cumulative share of running word tokens a learner recognises at each HSK level, across five reading registers, so you can see how far your vocabulary actually stretches into each kind of text. Paste your own text into the coverage analyzer for a per-text answer.
The matrix at a glance
The same numbers as the table below, as a heat grid: brighter amber is higher coverage. The grid is bright in the top-left (easy register, high level) and fades to almost nothing in the bottom-right, which is the whole finding in one picture.
| Register | HSK 1 | HSK 2 | HSK 3 | HSK 4 | HSK 5 | HSK 6 | HSK 7 to 9 |
|---|---|---|---|---|---|---|---|
| Subtitles | 50.4% | 60.7% | 66.8% | 69.8% | 72% | 73.9% | 77.5% |
| News | 27.9% | 37.8% | 48.6% | 54.4% | 58.4% | 62.4% | 69.4% |
| Wikipedia | 24.3% | 33.3% | 42.4% | 48.2% | 51.9% | 55.4% | 61.6% |
| Novels | 27.9% | 34.9% | 38.8% | 41.7% | 44.7% | 48.2% | 54.4% |
| Classical | 9.3% | 13.3% | 16.9% | 20% | 22.4% | 25.3% | 34% |
The coverage matrix (HSK 3.0)
Each cell is the share of running Chinese word tokens in that register recognised by a learner who knows every HSK word up to and including that level. Read across a row to see how one study level performs in different genres.
| Level | Words known | Film and TV subtitles | News and web journalism | Encyclopedic (Wikipedia) | Classic literature (Ming-Qing vernacular fiction) | Classical Chinese (wenyanwen) |
|---|---|---|---|---|---|---|
| HSK 1 | 506 | 50.4% | 27.9% | 24.3% | 27.9% | 9.3% |
| HSK 2 | 1,256 | 60.7% | 37.8% | 33.3% | 34.9% | 13.3% |
| HSK 3 | 2,209 | 66.8% | 48.6% | 42.4% | 38.8% | 16.9% |
| HSK 4 | 3,181 | 69.8% | 54.4% | 48.2% | 41.7% | 20% |
| HSK 5 | 4,240 | 72% | 58.4% | 51.9% | 44.7% | 22.4% |
| HSK 6 | 5,363 | 73.9% | 62.4% | 55.4% | 48.2% | 25.3% |
| HSK 7 to 9 | 10,969 | 77.5% | 69.4% | 61.6% | 54.4% | 34% |
The coverage matrix (HSK 2.0)
The same computation on the older six-level HSK 2.0 word lists, for learners studying against that standard.
| Level | Words known | Film and TV subtitles | News and web journalism | Encyclopedic (Wikipedia) | Classic literature (Ming-Qing vernacular fiction) | Classical Chinese (wenyanwen) |
|---|---|---|---|---|---|---|
| HSK 1 | 150 | 37.4% | 19.2% | 17.7% | 18.6% | 4.7% |
| HSK 2 | 297 | 46.3% | 25.4% | 22.3% | 23.7% | 8.2% |
| HSK 3 | 595 | 52.5% | 30.7% | 27.1% | 27.3% | 10.6% |
| HSK 4 | 1,193 | 57.8% | 38.2% | 34.3% | 31.4% | 16.1% |
| HSK 5 | 2,491 | 61.6% | 46.3% | 41.1% | 34.5% | 18.5% |
| HSK 6 | 4,991 | 64.5% | 51.7% | 45.1% | 36.6% | 20.6% |
What each register is
- Film and TV subtitles Spoken / dialogue
-
OpenSubtitles 2018 (Chinese, Simplified) word-frequency list.
77,769,662 Chinese-character word tokens. Source: hermitdave/FrequencyWords (content/2018/zh_cn/zh_cn_full.txt) (CC BY-SA 4.0). Segmentation: pre-segmented word-frequency list (source-provided).
- News and web journalism Journalistic / written-web
-
Leipzig zho_news_2020 (300K sentences). The Leipzig news crawl mixes news portals (Sing Tao, NYTimes Chinese, Nikkei Chinese) with forum text (bbs.voc.com.cn), so this register is online journalism plus discussion, not print-only news.
7,238,799 Chinese-character word tokens. Source: Leipzig Corpora Collection (Wortschatz, Universitaet Leipzig) (CC BY-NC 4.0 (Leipzig Corpora Collection); only derived coverage statistics are published here, never the corpus). Segmentation: Leipzig ASV tokenizer (source-provided word list).
- Encyclopedic (Wikipedia) Expository / reference
-
Leipzig zho_wikipedia_2018 (300K sentences).
5,099,888 Chinese-character word tokens. Source: Leipzig Corpora Collection (Wortschatz, Universitaet Leipzig) (CC BY-NC 4.0 (Leipzig Corpora Collection); only derived coverage statistics are published here, never the corpus). Segmentation: Leipzig ASV tokenizer (source-provided word list).
- Classic literature (Ming-Qing vernacular fiction) Literary narrative
-
Six public-domain vernacular novels: Dream of the Red Chamber, Journey to the West, Romance of the Three Kingdoms, The Scholars, Wonders Old and New, Stories to Caution the World. Classic (Ming-Qing) vernacular fiction, not contemporary novels. Character names and archaic vocabulary are outside the HSK list, so HSK coverage of literary prose is genuinely lower than of speech.
1,933,371 Chinese-character word tokens. Source: Project Gutenberg (public domain) (Public domain). Segmentation: OpenCC t2s (traditional to simplified) then jieba 0.42.1.
- Classical Chinese (wenyanwen) Classical literary
-
Eight public-domain classical works: Analects, Dao De Jing, Zuo Zhuan, Mencius, Records of the Grand Historian, Book of Han, History of Song, Xu Xiake's Travels. Classical Chinese uses a different lexicon and grammar from the modern HSK vocabulary. The low coverage figure is the point: a modern HSK vocabulary does not unlock classical texts, and even this figure overstates comprehension because shared characters carry different classical meanings.
3,448,284 Chinese-character word tokens. Source: Project Gutenberg (public domain) (Public domain). Segmentation: OpenCC t2s then jieba 0.42.1 (approximate: classical Chinese is largely monosyllabic and jieba is trained on modern Chinese, so segmentation is heuristic).
How to read this
- Coverage is recognition, not comprehension. Recognising 54.4 percent of the words in a news article is a planning estimate, not a claim you will understand it.
- This is coverage by the HSK vocabulary, not by a corpus top-N. It is lower than the usual "1000 words gives you X percent" figures, because those count a text's own most frequent words, and HSK order is not the same as any one text's frequency order. Proper nouns never appear on the HSK list.
- The classical figure understates the difficulty by design. Classical Chinese uses a different lexicon and grammar, and a shared character usually carries a different classical meaning, so even the low number here is optimistic. See the frequency coverage curve for the single-register view.
- The subtitles column is the validation. It reproduces our long-standing single-corpus coverage figures cell for cell, so the four other columns are the same method applied to their own corpus, not a different measure.
For each register, the cumulative share of running Chinese-character word tokens covered by knowing every HSK word up to and including a given level. This is coverage BY THE HSK VOCABULARY, not by the corpus's own top-N words, so it is lower than a raw frequency-rank coverage figure: HSK ordering is not identical to any single corpus's frequency ordering, and proper nouns are never in the HSK list. Denominator per register is every Chinese-character-only word token in that corpus. Latin and punctuation tokens are excluded. Matching is against the simplified HSK surface form only; the two Project Gutenberg registers are converted traditional-to-simplified before counting so the same simplified match applies to every register. All figures are computed from the corpora named above and are static, replicable assets, never proprietary data. Rebuild method and full per-corpus attribution: see the site's data provenance notes.
Next
Paste your own text into the coverage analyzer to see your recognition at each HSK level, use the difficulty grader to find the level a text needs, or read how many words you need to read Chinese.