Reference

Chinese reading and coverage glossary

The words behind the numbers: coverage, segmentation, and the vocabulary schemes the tools use, defined in plain English. 22 terms.

Character vs word
A distinction that trips up beginners: a character (hanzi) is a single written symbol, while a word can be one, two, or more characters together. HSK counts are word counts, not character counts, and the two totals are different, so a "2,500 characters" figure should never be compared directly with a word count.
Classical Chinese / 文言文
The written language of the Chinese canon, from the Analects to the dynastic histories, largely one character per word and built on function particles like 之, 其, and 也. It is a different register from modern Chinese, so a full modern HSK vocabulary recognises only about a quarter to a third of its words, and even that overstates real understanding because familiar characters carry shifted classical meanings. Coverage here reads as a difficulty signal, not a comprehension score.
Corpus
A large, structured collection of real text used to measure how words actually behave. The frequency ranks and the genre coverage figures here are computed from corpora such as film subtitles, a news crawl, encyclopedia articles, and public-domain novels, so the numbers reflect real usage rather than intuition. Every corpus is attributed with its size and license, and the exact percentage depends on which corpus a figure was measured against.
Coverage
The share of the words in a text that a reader already recognizes. If you know every word except a few, your coverage is high and the text feels readable. Reading research commonly cites roughly 95 percent known-word coverage for comfortable extensive reading and roughly 98 percent for comfortable independent reading. Coverage is a word-recognition estimate, not a guarantee that you understand every sentence.
Extensive vs intensive reading
Two study modes coverage helps you choose between. Extensive reading is reading a large volume of easy text for flow, ideally near 98 percent known-word coverage so you rarely stop. Intensive reading is working slowly through a hard text, looking words up and studying grammar, where low coverage is expected and fine. A coverage figure tells you which mode a given text actually suits, rather than whether the text is good or bad.
Frequency rank
A word's position in a corpus by how often it appears in real Chinese, where a lower rank means more common. The word tables here sort by frequency so the words you will actually meet most are at the top, which gives the fastest payoff per word learned.
Genre
A type of writing defined by its purpose and conventions: news, fiction, social media, encyclopedic reference, classical prose. Genre is the practical face of register, and it moves coverage sharply. The same study path that reads a TV subtitle comfortably can stall on a newspaper or a novel, because each genre draws on a different slice of the vocabulary and a different weight of names and specialist words.
Hanzi (Chinese character)
A single Chinese character, the basic written unit of the language. One character maps to roughly one syllable and often to a unit of meaning, but most everyday words are built from two or more characters. On this site the 简 column shows the simplified hanzi for each word.
HSK
The Hanyu Shuiping Kaoshi, the official standardized test of Mandarin proficiency for non-native speakers, run by China's education authorities. HSK is both an exam you can sit for a certificate and a graded vocabulary syllabus that tells you which words to learn in what order, which is useful even if you never take the test.
HSK 2.0
The six-level version of the HSK exam in use since 2010, topping out at about 5,000 words at HSK 6. Many test centers still administer it, which is why this site keeps a 2.0-vs-3.0 toggle on every word list so you can study against the exam your center actually offers.
HSK 3.0
The nine-level HSK standard announced in 2021, grouped into three bands and reaching a much larger vocabulary than the old exam. It also makes speaking mandatory from Level 3. Adoption has been staggered, so both HSK 2.0 and HSK 3.0 are in real use right now.
Named entity
A word that names a specific person, place, organisation, or product, such as 北京 or 王伟. Named entities sit outside the HSK vocabulary by definition, so they always count as unknown for coverage even when a reader recognises them instantly. To keep the missing-words list honest, the analyzer labels common surnames, given-name characters, and place names as names rather than as rare specialist vocabulary.
Pinyin
The official system for writing Mandarin sounds in the Latin alphabet, with tone marks over the vowels. It tells you how a character is pronounced. "nǐ hǎo" is the pinyin for 你好 (hello). Every word list on this site shows tone-marked pinyin next to the hanzi so you can read a word before you can write it.
Proper noun
A name for one specific thing, a person, a place, a country, as opposed to a common noun for a whole class of things. Proper nouns are never in the HSK word list, so they always score as unknown, which is one reason a real coverage figure sits a little below a perfectly clean estimate. The analyzer flags the common ones so a sentence full of people and places is not mistaken for a sentence full of hard vocabulary.
Register
The level and style of language a text uses: casual speech, journalism, formal writing, or classical prose. Register matters for coverage because the same vocabulary reaches far into one register and stalls in another. An HSK 4 vocabulary recognises about 70 percent of spoken subtitle words but only about 20 percent of a classical text, so a coverage figure only means something once you know which register the text belongs to.
Segmentation
Splitting a run of Chinese characters into words. Chinese is written without spaces, so a tool has to decide where one word ends and the next begins before it can score coverage. Any automatic segmenter mis-splits ambiguous strings, names, and words outside its dictionary, which is why coverage figures are close estimates rather than exact counts.
Simplified characters
The streamlined character set used in mainland China and Singapore, standardized in the 1950s and 1960s to have fewer strokes. The HSK exam uses simplified characters, so they are the default (the 简 column) on every word list here.
Spaced repetition
A review method that resurfaces each word just before you would forget it, stretching the gap between reviews as the word sticks. It is far more efficient than rereading a list, because you spend time only on words at risk, and it is the standard tool serious vocabulary learners rely on.
Token
A single counted unit of text. After segmentation, a Chinese sentence becomes a list of word tokens, and coverage is computed over the tokens that are actual Chinese words. Punctuation, spaces, and Latin letters are excluded from the count, so the percentage reflects the Chinese vocabulary in the text rather than its overall length.
Traditional characters
The older, fuller character forms still used in Taiwan, Hong Kong, and many overseas communities. Many characters are identical in both sets; where a word's traditional form differs, this site shows it in the 繁 column so you can recognize it when reading outside the mainland.
Type-token ratio
A measure of how varied a text's vocabulary is: the number of distinct words (types) divided by the total number of running words (tokens). A low ratio means heavy repetition, which makes a text easier to read once you know its core words; a high ratio means the writing keeps introducing new vocabulary. Spoken dialogue tends to repeat, while news and classical prose spread across far more distinct types.
Word count vs character count
Two totals that measure different things and should never be swapped. A character count tallies individual written symbols; a word count tallies segmented words, which may be one, two, or more characters each. A text's word count is lower than its character count, and vocabulary standards like HSK are sized in words, so comparing a word figure against a character figure misstates the real learning load. Coverage here is measured over words, after segmentation.