中文圣经

New Words Appear at N^0.492 While Reading

Answering: “how fast new words appear when reading”

Updated 2026-09-26 · plain text · JSON

Heaps' law describes how a text's vocabulary grows as you read: distinct words ≈ K·N^β, where N is words read. Fitted over this text in canonical order, sampled every thousand words, β comes out at 0.492 with K 16.64. Below 1 means what every reader hopes for — the further you read, the fewer surprises per page.

The curve in concrete terms: Genesis alone introduces 2,848 distinct words, 24.3% of the Bible's whole vocabulary, in its first 32,144 words of text. By the end of the Old Testament you have met 10,398 (88.6%), and Matthew — a whole book — adds only 248 more.

That asymmetry is the argument for finishing something rather than sampling: the hard work is front-loaded. It is also why the site's decks are built per book in frequency order — the first two hundred words of a book carry most of its repetitions, and the rest of it is largely vocabulary you already met.

Vocabulary met while reading in canonical order
AfterWords of text readDistinct words metShare of the Bible's vocabulary
Genesis32,1442,84824.3%
Deuteronomy124,9645,27945%
Psalms342,4268,84475.3%
Isaiah392,2139,70682.7%
Malachi494,74810,39888.6%
Matthew514,42710,64690.7%
John565,21510,94493.2%
Revelation650,62811,739100%

Questions

Does reading order change this?
The fit is for canonical order. Reading the Gospels first meets a smaller vocabulary faster — their 3,971 words already cover 88.7% of the whole Bible's word tokens.
What is K?
The constant in K·N^β, fitted here at 16.64. It has no meaning on its own; it scales the curve.
How was the fit made?
Least squares on log(distinct words) against log(words read), sampled every 1,000 words across all 650,628 tokens.

Related reference pages

Take it further

How these numbers were produced

  • Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
  • Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
  • HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
  • Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
  • Computed: 2026-09-26, from the text on this site.

These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.

Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).