Reference

Frequently asked questions

The questions we hear most about coverage analysis, answered plainly.

Can the tool read classical Chinese (文言文)?

You can paste classical Chinese and get a coverage number, but read it as a difficulty signal, not a comprehension score. A full HSK vocabulary recognises only about a quarter to a third of classical words, and even that overstates understanding, because familiar characters carry different classical meanings. Segmentation is heuristic on monosyllabic classical text.

Does the genre of a text change how much coverage I get?

Yes, sharply. The HSK 4 vocabulary that recognises about 70 percent of TV subtitle words covers only about 54 percent of news, 42 percent of a classic novel, and 20 percent of a classical text. Spoken dialogue leans on a few common words, while written and older genres spread across far more vocabulary and names.

Does HanziCoverage upload the text I paste?

No. Segmentation, classification, and coverage math all run in your browser over data bundled with the page, so your pasted text is never uploaded or transmitted. We receive only anonymous aggregate statistics, such as a coarse text-length band and a coverage band, never the text itself and no accounts or identifiers.

How accurate is the coverage estimate?

It is a close estimate, not an exact count. Automatic segmentation mis-splits names and words outside the dictionary, and coverage assumes you know every word up to your chosen level. Use it to compare texts and plan study, not as a precise reading score.

How does the analyzer handle names and proper nouns?

Proper nouns are never HSK words, so they always count as unknown for coverage. To keep the missing-words list honest, the tool labels common surnames, given-name characters, and place names as names rather than as rare vocabulary, so a sentence full of people and places does not read as full of hard words.

How many characters do you need to read a Chinese newspaper?

Comfortable newspaper reading usually wants roughly 2,000 to 3,000 characters and 3,000 to 4,000 words, but the honest answer is your coverage of the specific article. Paste a real sample into the analyzer to see where you land rather than trusting a round number.

Does the analyzer work with Traditional characters?

Yes. You can paste Simplified or Traditional Chinese, or a mix. The dictionary maps common Traditional forms back to their Simplified entry before matching, so coverage and the gap list work either way. Traditional variants that share a Simplified word are treated as the same word.

Is word coverage the same as understanding the text?

No. Coverage measures how many words you recognize, not whether you understand every sentence. Grammar, idiom, and the rare content words that carry a topic's meaning are exactly the hard part, and word coverage cannot see them. Treat a coverage figure as a planning estimate for how readable a text is likely to be, not a comprehension score.

What coverage do I need to read Chinese comfortably?

Reading research commonly cites roughly 95 percent known-word coverage for comfortable extensive reading and roughly 98 percent for comfortable independent reading without constant dictionary help. Below about 90 percent, a text is usually a hard slog. These are guideline figures from the reading literature, not exact thresholds, and coverage is not the same as full comprehension.

What is the segmentation lexicon the tool uses?

Chinese is written without spaces, so the tool first splits your text into words using a dictionary of about 50,000 common word forms drawn from an open subtitle corpus. Keeping frequent non-HSK words whole, instead of shattering them into single characters, makes the missing-words list accurate. It only aids splitting, never the coverage math.