How many Chinese words do you actually need?
A small number of high-frequency words carries most of everyday Chinese. This tool lets you drag N and see how much of typical running text the top N words cover, for both the HSK 2.0 and 3.0 schemes. It is the clearest case for learning the common words first.
Knowing the most common 2,209 words (up to HSK 3.0 Level 3) covers about 66.8% of the running words in everyday spoken Chinese.
| Level | Words known (cumulative) | Coverage of everyday speech |
|---|---|---|
| HSK 3.0 Level 1 | 506 | 50.4% |
| HSK 3.0 Level 2 | 1,256 | 60.7% |
| HSK 3.0 Level 3 | 2,209 | 66.8% |
| HSK 3.0 Level 4 | 3,181 | 69.8% |
| HSK 3.0 Level 5 | 4,240 | 72% |
| HSK 3.0 Level 6 | 5,363 | 73.9% |
| HSK 3.0 Levels 7 to 9 | 10,969 | 77.5% |
Source: OpenSubtitles 2018 (Chinese, Simplified) word-frequency list, via hermitdave/FrequencyWords (content/2018/zh_cn/zh_cn_full.txt) (CC BY-SA 4.0). Coverage is the cumulative share of running words a learner who knows every word up to a level would recognize. It is token coverage in a spoken register, not full comprehension.
Why frequency order matters
Because word frequency in Chinese is steeply skewed, the first thousand or so words you learn do far more work than the next thousand. That is the entire argument for learning common words in order and drilling them until they stick. See the best order to learn Chinese words.
Honest caveats
- The coverage figures come from a film and television subtitle corpus, a spoken register. A newspaper, a contract, or a poem has a different profile.
- Covering a share of words is not the same as understanding: it is a planning estimate.
- The curve is a static, replicable computed asset built from a public frequency list, not proprietary data. Its source and method are documented on-site.