Chinese text difficulty grader
Choosing a reading for a class or a tutee? Paste a candidate text and the grader returns the HSK level a reader needs to reach comfortable and independent coverage, a plain-language difficulty band, which reading register the passage resembles, and the words that sit above level or outside the HSK list entirely. It runs in your browser: no accounts, no student data, and your text is never uploaded.
Quick answer
How do I tell if a Chinese text is the right level for my class?
Grade it by word coverage. Reading research links comfortable reading to knowing about 95 percent of the running words and independent reading to about 98 percent. Paste the passage below and the grader finds the lowest HSK level that clears those marks, names the register it reads like, and flags words above level or outside HSK.
Your pasted text is graded here in your browser and is never uploaded, and no student data is involved. The tool records only anonymous aggregate statistics, a coarse text-length band and a coverage band, never the text itself and no accounts or identifiers. Full detail is on the privacy page.
What the grader reports
- Minimum HSK level for 95% and 98% coverage. The lowest level at which a reader recognizes about 95 percent (comfortable reading) and 98 percent (independent reading) of the Chinese words in your text.
- A plain-language band: beginner, lower-intermediate, intermediate, or advanced, set from where the 95 percent mark lands.
- Which register it resembles. The passage's coverage curve is matched against five reading registers so you can see whether it reads like dialogue, news, reference, fiction, or classical writing.
- An above-level and out-of-register breakdown. The words a reader would not yet know split into HSK words above the comfortable level and words outside the HSK list entirely, with detected proper nouns kept separate so a text full of names does not read as a wall of vocabulary.
- The share of vocabulary outside HSK, plus a warning when a passage reads as classical or highly technical, where a modern HSK coverage number understates the real difficulty.
Why one HSK number is never enough: 95% and 98% by register
A single coverage figure is only true for the kind of text it came from. The same HSK vocabulary that covers most of a TV subtitle covers far less of a newspaper, and less again of a classical essay, because news, reference, fiction, and classical writing spread across far more vocabulary and proper nouns. In fact, HSK vocabulary alone does not reach 95 percent coverage of any ordinary register: a real reader closes the last few percent with the proper nouns and common non-HSK words this grader also surfaces. The table shows how far the full HSK list stretches into each register, computed from open corpora (not from your pasted text).
| Register | HSK 1 to 9 ceiling | Reaches 95%? | Reaches 98%? |
|---|---|---|---|
| Film and TV subtitlesSpoken / dialogue | 77.5% | No | No |
| News and web journalismJournalistic / written-web | 69.4% | No | No |
| Encyclopedic (Wikipedia)Expository / reference | 61.6% | No | No |
| Classic literature (Ming-Qing vernacular fiction)Literary narrative | 54.4% | No | No |
| Classical Chinese (wenyanwen)Classical literary | 34% | No | No |
Ceilings are the HSK 1 to 9 figures from the coverage-by-genre matrix. Because no register clears 95 percent on HSK words alone, the grader above answers the per-text question directly: it counts the proper nouns and out-of-list words in your passage so you can see whether the gap is a handful of names or a genuine vocabulary wall.
How the difficulty band is set
The band comes from where the 95 percent comfortable-reading mark lands in your text. If a reader clears it by HSK 1 or 2 the passage is beginner; by HSK 3 or 4, lower-intermediate; by HSK 5 or 6, intermediate; only at HSK 7 to 9, or not at all within HSK, advanced. The share of words outside HSK is shown alongside so you can tell an advanced text apart from an ordinary one that simply leans on names or specialist terms.
How to read the result honestly
- Coverage is not comprehension. The grade is a word-recognition estimate. Grammar, idiom, and rare content words carry meaning that word coverage cannot see, so use the band to compare texts and set expectations, not as an exact reading score.
- Segmentation is imperfect. Chinese is written without spaces, so any automatic segmenter mis-splits ambiguous strings and words outside its dictionary. Treat the split, and the coverage, as a close estimate.
- Name detection is conservative. The grader tags common surnames with an out-of-list given name, common place and country names, and interpunct foreign names, so it keeps most names out of the vocabulary list, but it can miss a name built from ordinary characters. Out-of-list words it cannot place are shown honestly as beyond the dictionary, not sorted into rare versus specialist.
- Classical Chinese is out of register. The lexicon here is modern HSK; classical text uses a different vocabulary and grammar, so its low coverage number understates the difficulty. Read a classical warning as "this is not modern Chinese," not as a study target.
Related tools: the coverage analyzer gives a learner the same passage from their own point of view (what they can already read at their level), the coverage-by-genre matrix shows every level in every register, and the guide on estimating the HSK level of a text explains the method behind this grade.
HanziCoverage is independent and not affiliated with HSK, Hanban, or Chinese Testing International. Word data is derived from the HSK 2.0 and 3.0 vocabulary lists and CC-CEDICT; the frequency ordering is from the OpenSubtitles 2018 Chinese frequency list (hermitdave/FrequencyWords, CC BY-SA 4.0). The per-register figures are computed from open corpora, each attributed on the coverage-by-genre page. Token coverage is not comprehension. These figures show what share of words a learner recognises at each HSK level in each register, not how much they understand. Registers differ in how frequency-skewed they are: subtitle dialogue is dominated by a small set of very common words, so HSK coverage there is high; news, encyclopedic, literary and classical text spread probability over far more vocabulary and proper nouns, so the same HSK level covers much less.