How Much Chinese Can You Already Read?
Quick answer
How Much Chinese Can You Already Read?
Paste any Chinese text and estimate your word-recognition coverage at each HSK level, then see which words to learn next. A guide to reading your own coverage honestly.
The most useful thing you can know before you sit down to read a piece of Chinese is how much of it you can already read. Not your HSK level in the abstract, but your coverage of this text: what share of its words you would recognize, and which specific words you would not. That is a concrete, answerable question, and the analyzer answers it in your browser without uploading anything.
What the analyzer actually measures
Paste any Simplified or Traditional Chinese and the tool does three things, all locally:
- Segments the text into words. Chinese is written without spaces, so it uses a dictionary-based longest-match pass to split the string into words rather than loose characters.
- Looks up each word. Every word is checked against the combined HSK 2.0 and 3.0 vocabulary and its corpus frequency.
- Computes coverage at each level. Assuming you know every word at or below a given HSK level, it reports what share of the text’s words that covers, for every level in both schemes.
It then lists the words above your chosen level, ranked most-common-first, so the words at the top of the list are the highest-value ones to learn next.
A worked example
Say you paste a short news paragraph and set your level to HSK 3.0 Level 4. The tool might report 88 percent coverage. That single number tells you a lot: you are below the roughly 95 percent comfort threshold, so this paragraph will feel like work rather than reading, and roughly one word in eight is unknown to you. The ranked gap list then shows you which words those are, with the common ones first. If the top few are words you keep meeting elsewhere, they are your next study batch.
Bump the level selector to HSK 3.0 Level 5 and the coverage number jumps, because the level-5 vocabulary fills in many of those gaps. That is the coverage curve made concrete for your own text.
Reading the result honestly
Coverage is a genuinely useful estimate, but it is an estimate, and treating it as an exact score will mislead you:
- Segmentation is imperfect. Any automatic segmenter mis-splits names, ambiguous strings, and words outside its dictionary. Words the dictionary does not recognize are shown as “outside the HSK list,” and they inflate or deflate the raw count a little. Read the split as a close approximation.
- Coverage is not comprehension. Recognizing most words is not understanding every sentence. The unknown words are often the ones carrying the meaning, and grammar and idiom are invisible to a word count.
- HSK membership is a proxy for “known.” The tool assumes you know every word up to your selected level, which no real learner does exactly. It is a planning estimate of the text, not a test of you.
What to do with it
Once you know your coverage and your gap list, the move is the same one the whole site points at: learn the common words you are missing, in frequency order, and review them so they stick. The best order to learn Chinese words explains why frequency order is the efficient path, and spaced-repetition drilling, for example with Wordbrush from our publisher Urban Algorithm, keeps the words you add from leaking back out. The analyzer itself is free, ungated, and never sends your text anywhere.