Chinese text difficulty grader

Choosing a reading for a class or a tutee? Paste a candidate text and the grader returns the HSK level a reader needs to reach comfortable and independent coverage, a plain-language difficulty band, which reading register the passage resembles, and the words that sit above level or outside the HSK list entirely. It runs in your browser: no accounts, no student data, and your text is never uploaded.

Quick answer

How do I tell if a Chinese text is the right level for my class?

Grade it by word coverage. Reading research links comfortable reading to knowing about 95 percent of the running words and independent reading to about 98 percent. Paste the passage below and the grader finds the lowest HSK level that clears those marks, names the register it reads like, and flags words above level or outside HSK.

Your pasted text is graded here in your browser and is never uploaded, and no student data is involved. The tool records only anonymous aggregate statistics, a coarse text-length band and a coverage band, never the text itself and no accounts or identifiers. Full detail is on the privacy page.

What the grader reports

Why one HSK number is never enough: 95% and 98% by register

A single coverage figure is only true for the kind of text it came from. The same HSK vocabulary that covers most of a TV subtitle covers far less of a newspaper, and less again of a classical essay, because news, reference, fiction, and classical writing spread across far more vocabulary and proper nouns. In fact, HSK vocabulary alone does not reach 95 percent coverage of any ordinary register: a real reader closes the last few percent with the proper nouns and common non-HSK words this grader also surfaces. The table shows how far the full HSK list stretches into each register, computed from open corpora (not from your pasted text).

Coverage of each reading register by the full HSK vocabulary, and whether it reaches 95 or 98 percent.
Register HSK 1 to 9 ceiling Reaches 95%? Reaches 98%?
Film and TV subtitlesSpoken / dialogue 77.5% No No
News and web journalismJournalistic / written-web 69.4% No No
Encyclopedic (Wikipedia)Expository / reference 61.6% No No
Classic literature (Ming-Qing vernacular fiction)Literary narrative 54.4% No No
Classical Chinese (wenyanwen)Classical literary 34% No No

Ceilings are the HSK 1 to 9 figures from the coverage-by-genre matrix. Because no register clears 95 percent on HSK words alone, the grader above answers the per-text question directly: it counts the proper nouns and out-of-list words in your passage so you can see whether the gap is a handful of names or a genuine vocabulary wall.

How the difficulty band is set

The band comes from where the 95 percent comfortable-reading mark lands in your text. If a reader clears it by HSK 1 or 2 the passage is beginner; by HSK 3 or 4, lower-intermediate; by HSK 5 or 6, intermediate; only at HSK 7 to 9, or not at all within HSK, advanced. The share of words outside HSK is shown alongside so you can tell an advanced text apart from an ordinary one that simply leans on names or specialist terms.

How to read the result honestly

Related tools: the coverage analyzer gives a learner the same passage from their own point of view (what they can already read at their level), the coverage-by-genre matrix shows every level in every register, and the guide on estimating the HSK level of a text explains the method behind this grade.

HanziCoverage is independent and not affiliated with HSK, Hanban, or Chinese Testing International. Word data is derived from the HSK 2.0 and 3.0 vocabulary lists and CC-CEDICT; the frequency ordering is from the OpenSubtitles 2018 Chinese frequency list (hermitdave/FrequencyWords, CC BY-SA 4.0). The per-register figures are computed from open corpora, each attributed on the coverage-by-genre page. Token coverage is not comprehension. These figures show what share of words a learner recognises at each HSK level in each register, not how much they understand. Registers differ in how frequency-skewed they are: subtitle dialogue is dominated by a small set of very common words, so HSK coverage there is high; news, encyclopedic, literary and classical text spread probability over far more vocabulary and proper nouns, so the same HSK level covers much less.