New Words Appear at N^0.492 While Reading
Answering: “how fast new words appear when reading”
Updated 2026-09-26 · plain text · JSON
Heaps' law describes how a text's vocabulary grows as you read: distinct words ≈ K·N^β, where N is words read. Fitted over this text in canonical order, sampled every thousand words, β comes out at 0.492 with K 16.64. Below 1 means what every reader hopes for — the further you read, the fewer surprises per page.
The curve in concrete terms: Genesis alone introduces 2,848 distinct words, 24.3% of the Bible's whole vocabulary, in its first 32,144 words of text. By the end of the Old Testament you have met 10,398 (88.6%), and Matthew — a whole book — adds only 248 more.
That asymmetry is the argument for finishing something rather than sampling: the hard work is front-loaded. It is also why the site's decks are built per book in frequency order — the first two hundred words of a book carry most of its repetitions, and the rest of it is largely vocabulary you already met.
| After | Words of text read | Distinct words met | Share of the Bible's vocabulary |
|---|---|---|---|
| Genesis | 32,144 | 2,848 | 24.3% |
| Deuteronomy | 124,964 | 5,279 | 45% |
| Psalms | 342,426 | 8,844 | 75.3% |
| Isaiah | 392,213 | 9,706 | 82.7% |
| Malachi | 494,748 | 10,398 | 88.6% |
| Matthew | 514,427 | 10,646 | 90.7% |
| John | 565,215 | 10,944 | 93.2% |
| Revelation | 650,628 | 11,739 | 100% |
Questions
- Does reading order change this?
- The fit is for canonical order. Reading the Gospels first meets a smaller vocabulary faster — their 3,971 words already cover 88.7% of the whole Bible's word tokens.
- What is K?
- The constant in K·N^β, fitted here at 16.64. It has no meaning on its own; it scales the curve.
- How was the fit made?
- Least squares on log(distinct words) against log(words read), sampled every 1,000 words across all 650,628 tokens.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-26, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).