The 98% Comprehension Research, Honestly
Quick answer
The 98% Comprehension Research, Honestly
Where the 95% and 98% reading-coverage figures actually come from, what the studies did and did not show, and how to apply them to Chinese without overclaiming.
The claim that you need “98 percent coverage to read comfortably” is repeated everywhere in language learning, usually with no source and a lot of false precision. It rests on real research, but the research is narrower and more careful than the slogan. This guide lays out what the studies actually found, and how to use the figures honestly for Chinese.
Where the numbers come from
The coverage-and-comprehension figures trace mainly to vocabulary research in second-language reading:
- Hu and Nation (2000), Unknown vocabulary density and reading comprehension, is the study most often cited for the 98 percent figure. Working with English learners reading fiction, it found that adequate unassisted comprehension generally required knowing around 98 percent of the running words, and that comprehension dropped off as coverage fell below that.
- Laufer (1989 and later work) is associated with the lower, roughly 95 percent, threshold as a point below which reading comprehension tends to break down.
- Nation (2006) estimated the vocabulary size needed to reach those coverage levels for English text, reinforcing that the thresholds are about coverage of running words, not a fixed word count.
These are the anchors. The community discussion that brought them into Chinese learning, including at Hacking Chinese, generally cites this same lineage.
What the studies did and did not show
It is worth being precise, because the slogan flattens several caveats:
- The figures are about word coverage, not comprehension directly. The studies relate the percentage of words known to comprehension outcomes. They do not say you understand 98 percent of the meaning at 98 percent coverage. The unknown 2 percent is often exactly the words that carry the point.
- The thresholds are ranges, not switches. Different studies, text types, and readers give somewhat different numbers. “95 and 98” are convenient landmarks, not hard cutoffs, and treating them as exact is a misreading of the research.
- Most of the original work was on English. Applying the figures to Chinese is reasonable and common, because the underlying relationship between coverage and comprehension is not English-specific, but it is an extrapolation. Chinese adds wrinkles the English studies did not face, such as word segmentation and a character layer beneath the word layer.
- “Adequate comprehension” was defined by the studies, not by you. The comprehension bar in the research is a specific measured outcome, not “enjoyed the book.” Your personal comfort may sit at a different coverage than the study average.
How to apply it to Chinese without overclaiming
The responsible way to use these figures is as planning landmarks:
- Treat roughly 95 percent as “readable for pleasure with some guessing” and roughly 98 percent as “readable without a dictionary.” Both are approximate.
- Measure coverage on your own material rather than trusting a global number. The analyzer estimates your coverage of any text you paste, and the difficulty grader reports the level at which a text crosses these marks.
- Remember that coverage is a proxy. A high coverage number means you will hit fewer unknown words, not that you will understand everything between them. This is covered further in the 95% and 98% coverage rules.
Why we never claim these as our own numbers
On this site the 95 and 98 percent figures are always presented as figures from the reading-acquisition literature, cited to that literature, and never as a HanziCoverage finding. Our tools compute your coverage of your text; the interpretation thresholds are borrowed, sourced, and hedged. That distinction is the whole point of using them honestly: the math is ours, the thresholds are the field’s, and neither is a guarantee about any individual reader.