How much Chinese do you need to read the news, novels, or classical texts?

Quick answer

Does the genre change how much Chinese I need?

Yes, sharply. The HSK 4 vocabulary that recognises about 69.8 percent of TV subtitle words covers only about 54.4 percent of news text, 41.7 percent of a classic novel, and 20 percent of a classical text. Spoken dialogue leans on a few very common words; written and older registers spread across far more vocabulary, so the same study level goes much less far.

Last reviewed 24 July 2026 by Nathanael Desmond

Most coverage tools report one number from one corpus. This matrix computes the same thing, the cumulative share of running word tokens a learner recognises at each HSK level, across five reading registers, so you can see how far your vocabulary actually stretches into each kind of text. Paste your own text into the coverage analyzer for a per-text answer.

The matrix at a glance

The same numbers as the table below, as a heat grid: brighter amber is higher coverage. The grid is bright in the top-left (easy register, high level) and fades to almost nothing in the bottom-right, which is the whole finding in one picture.

Coverage heat grid: reading register by HSK 3.0 level Heat grid of cumulative HSK 3.0 coverage: five reading registers (rows) by study level (columns). Coverage is highest for subtitles at high levels and lowest for classical text, fading from about 77.5% to 9.3%. 1234567-9 HSK 3.0 level Subtitles 50.4% 60.7% 66.8% 69.8% 72% 73.9% 77.5% News 27.9% 37.8% 48.6% 54.4% 58.4% 62.4% 69.4% Wikipedia 24.3% 33.3% 42.4% 48.2% 51.9% 55.4% 61.6% Novels 27.9% 34.9% 38.8% 41.7% 44.7% 48.2% 54.4% Classical 9.3% 13.3% 16.9% 20% 22.4% 25.3% 34%
Cumulative HSK 3.0 word coverage across five registers. Brighter amber is higher coverage. The same study level goes far in speech and almost nowhere in classical text. Figures match the table below; token coverage is recognition, not comprehension.
Cumulative HSK 3.0 coverage by register and level
RegisterHSK 1HSK 2HSK 3HSK 4HSK 5HSK 6HSK 7 to 9
Subtitles 50.4%60.7%66.8%69.8%72%73.9%77.5%
News 27.9%37.8%48.6%54.4%58.4%62.4%69.4%
Wikipedia 24.3%33.3%42.4%48.2%51.9%55.4%61.6%
Novels 27.9%34.9%38.8%41.7%44.7%48.2%54.4%
Classical 9.3%13.3%16.9%20%22.4%25.3%34%

The coverage matrix (HSK 3.0)

Each cell is the share of running Chinese word tokens in that register recognised by a learner who knows every HSK word up to and including that level. Read across a row to see how one study level performs in different genres.

Cumulative HSK 3.0 word coverage by reading register
Level Words known Film and TV subtitlesNews and web journalismEncyclopedic (Wikipedia)Classic literature (Ming-Qing vernacular fiction)Classical Chinese (wenyanwen)
HSK 1 506 50.4%27.9%24.3%27.9%9.3%
HSK 2 1,256 60.7%37.8%33.3%34.9%13.3%
HSK 3 2,209 66.8%48.6%42.4%38.8%16.9%
HSK 4 3,181 69.8%54.4%48.2%41.7%20%
HSK 5 4,240 72%58.4%51.9%44.7%22.4%
HSK 6 5,363 73.9%62.4%55.4%48.2%25.3%
HSK 7 to 9 10,969 77.5%69.4%61.6%54.4%34%

The coverage matrix (HSK 2.0)

The same computation on the older six-level HSK 2.0 word lists, for learners studying against that standard.

Cumulative HSK 2.0 word coverage by reading register
Level Words known Film and TV subtitlesNews and web journalismEncyclopedic (Wikipedia)Classic literature (Ming-Qing vernacular fiction)Classical Chinese (wenyanwen)
HSK 1 150 37.4%19.2%17.7%18.6%4.7%
HSK 2 297 46.3%25.4%22.3%23.7%8.2%
HSK 3 595 52.5%30.7%27.1%27.3%10.6%
HSK 4 1,193 57.8%38.2%34.3%31.4%16.1%
HSK 5 2,491 61.6%46.3%41.1%34.5%18.5%
HSK 6 4,991 64.5%51.7%45.1%36.6%20.6%

What each register is

Film and TV subtitles Spoken / dialogue

OpenSubtitles 2018 (Chinese, Simplified) word-frequency list.

77,769,662 Chinese-character word tokens. Source: hermitdave/FrequencyWords (content/2018/zh_cn/zh_cn_full.txt) (CC BY-SA 4.0). Segmentation: pre-segmented word-frequency list (source-provided).

News and web journalism Journalistic / written-web

Leipzig zho_news_2020 (300K sentences). The Leipzig news crawl mixes news portals (Sing Tao, NYTimes Chinese, Nikkei Chinese) with forum text (bbs.voc.com.cn), so this register is online journalism plus discussion, not print-only news.

7,238,799 Chinese-character word tokens. Source: Leipzig Corpora Collection (Wortschatz, Universitaet Leipzig) (CC BY-NC 4.0 (Leipzig Corpora Collection); only derived coverage statistics are published here, never the corpus). Segmentation: Leipzig ASV tokenizer (source-provided word list).

Encyclopedic (Wikipedia) Expository / reference

Leipzig zho_wikipedia_2018 (300K sentences).

5,099,888 Chinese-character word tokens. Source: Leipzig Corpora Collection (Wortschatz, Universitaet Leipzig) (CC BY-NC 4.0 (Leipzig Corpora Collection); only derived coverage statistics are published here, never the corpus). Segmentation: Leipzig ASV tokenizer (source-provided word list).

Classic literature (Ming-Qing vernacular fiction) Literary narrative

Six public-domain vernacular novels: Dream of the Red Chamber, Journey to the West, Romance of the Three Kingdoms, The Scholars, Wonders Old and New, Stories to Caution the World. Classic (Ming-Qing) vernacular fiction, not contemporary novels. Character names and archaic vocabulary are outside the HSK list, so HSK coverage of literary prose is genuinely lower than of speech.

1,933,371 Chinese-character word tokens. Source: Project Gutenberg (public domain) (Public domain). Segmentation: OpenCC t2s (traditional to simplified) then jieba 0.42.1.

Classical Chinese (wenyanwen) Classical literary

Eight public-domain classical works: Analects, Dao De Jing, Zuo Zhuan, Mencius, Records of the Grand Historian, Book of Han, History of Song, Xu Xiake's Travels. Classical Chinese uses a different lexicon and grammar from the modern HSK vocabulary. The low coverage figure is the point: a modern HSK vocabulary does not unlock classical texts, and even this figure overstates comprehension because shared characters carry different classical meanings.

3,448,284 Chinese-character word tokens. Source: Project Gutenberg (public domain) (Public domain). Segmentation: OpenCC t2s then jieba 0.42.1 (approximate: classical Chinese is largely monosyllabic and jieba is trained on modern Chinese, so segmentation is heuristic).

How to read this

For each register, the cumulative share of running Chinese-character word tokens covered by knowing every HSK word up to and including a given level. This is coverage BY THE HSK VOCABULARY, not by the corpus's own top-N words, so it is lower than a raw frequency-rank coverage figure: HSK ordering is not identical to any single corpus's frequency ordering, and proper nouns are never in the HSK list. Denominator per register is every Chinese-character-only word token in that corpus. Latin and punctuation tokens are excluded. Matching is against the simplified HSK surface form only; the two Project Gutenberg registers are converted traditional-to-simplified before counting so the same simplified match applies to every register. All figures are computed from the corpora named above and are static, replicable assets, never proprietary data. Rebuild method and full per-corpus attribution: see the site's data provenance notes.

Next

Paste your own text into the coverage analyzer to see your recognition at each HSK level, use the difficulty grader to find the level a text needs, or read how many words you need to read Chinese.