Chinese text coverage analyzer

Paste any Simplified or Traditional Chinese below. HanziCoverage segments the text, matches each word against the combined HSK 2.0 and 3.0 vocabulary (over 11,000 words, ordered by corpus frequency), and estimates how much of the text a learner at each HSK level would already recognize. It runs entirely in your browser: your text is never uploaded.

Your pasted text is never uploaded. The tool records only anonymous aggregate statistics, a coarse text-length band and a coverage band, never the text itself and no accounts or identifiers. Full detail is on the privacy page.

What "coverage" means

Coverage is the share of the words in a text that you already recognize. If you know every word in a text except a handful, your coverage is high and the text will feel readable. Reading research commonly cites roughly 95 percent known-word coverage for comfortable extensive reading and roughly 98 percent for comfortable independent reading. Those thresholds are discussed, with sources, on the comprehensible-input guide.

How far your HSK level reaches, by genre

The coverage of a single text depends heavily on what kind of text it is. Spoken dialogue leans on a small set of very common words, so an HSK level covers a lot of it; news, reference, fiction, and classical writing spread across far more vocabulary and proper nouns, so the same level covers much less. The figures below are the cumulative share of running words a learner recognizes at a given HSK level in five reading registers, computed from open corpora (not from your pasted text). Change the level to see how far your vocabulary stretches into each genre.

Scheme

At HSK 4 (knowing about 3,181 HSK words), a learner recognizes:

Film and TV subtitles 69.8%
News and web journalism 54.4%
Encyclopedic (Wikipedia) 48.2%
Classic literature (Ming-Qing vernacular fiction) 41.7%
Classical Chinese (wenyanwen) 20%

These are corpus-wide averages for each register, not a reading of your own text. See the full coverage-by-genre matrix for every level and both HSK schemes, or paste a specific passage above for a per-text answer and use the difficulty grader to find the level a text needs.

A worked example

Say you have studied through HSK 4. Paste a chatty message or a TV-subtitle line into the analyzer and you will often land near 69.8 percent coverage, close to what the subtitle corpus predicts for that level, so it reads comfortably. Paste a news article and the same vocabulary typically covers only about 54.4 percent, and a passage of classical Chinese drops to roughly 20 percent. That is not a bug in your studying: it is why a learner who breezes through dialogue can still stall on the front page. The analyzer shows the number for the exact text in front of you; the genre table above shows what to expect before you start.

How the tool works

Four-stage diagram of the analyzer pipeline: a block of pasted text, then the same text split into bracketed word-chunks (segmentation), then those chunks matched against a dictionary card with some filled amber to show a match, then a circular coverage gauge with an amber covered arc as the output.
  1. It segments your pasted text into words using a dictionary-based longest-match pass.
  2. It looks up each word's HSK level (both the 2.0 and 3.0 schemes) and a corpus frequency.
  3. It computes your recognized share at each level, counting only Chinese-character words.
  4. It lists the words above your chosen level, ranked by how common they are, so the highest-value words to learn are at the top.
  5. It recognizes conservative proper names, a common surname followed by a given name, or an interpunct foreign name like a transliteration, and keeps them out of the words-to-learn list, since you memorize a name once rather than studying it as vocabulary.
  6. Any character it still cannot place in the dictionary is grouped as beyond the dictionary, not padded into your study list, so a text full of names or specialist terms does not read as a wall of vocabulary you must memorize.

Simplified, Traditional, or a mix

You can paste Simplified characters, Traditional characters, or a document that mixes both. HSK words are matched in either script against the same vocabulary, using a variant map from CC-CEDICT, so a Traditional-only text recognizes the same HSK vocabulary as its Simplified equivalent and a mixed-script document is handled word by word. Where a surface form exists in both scripts, the Simplified reading is used. This is a character-matching step, not a converter: it does not rewrite your text. The variant map covers HSK vocabulary, so a non-HSK word or a name written only in Traditional may fall through to the beyond-the-dictionary bucket like any other unrecognized character.

What to trust, and what not to

HanziCoverage is independent and is not affiliated with HSK, Hanban, or Chinese Testing International. Word data is derived from the HSK 2.0 and 3.0 vocabulary lists and CC-CEDICT; the frequency ordering is from the OpenSubtitles 2018 Chinese frequency list (hermitdave/FrequencyWords, CC BY-SA 4.0). The per-genre figures are computed from open corpora (film subtitles, a news crawl, Chinese Wikipedia, and public-domain fiction and classical works), each attributed on the coverage-by-genre page. Token coverage is not comprehension. These figures show what share of words a learner recognises at each HSK level in each register, not how much they understand. Registers differ in how frequency-skewed they are: subtitle dialogue is dominated by a small set of very common words, so HSK coverage there is high; news, encyclopedic, literary and classical text spread probability over far more vocabulary and proper nouns, so the same HSK level covers much less.