New Words Appear at N^0.482 While Reading
Answering: “how fast new words appear when reading”
Updated 2026-09-22 · plain text · JSON
Heaps' law describes how a text's vocabulary grows as you read: distinct words ≈ K·N^β, where N is words read. Fitted over this text in canonical order, sampled every thousand words, β comes out at 0.482 with K 17.74. Below 1 means what every reader hopes for — the further you read, the fewer surprises per page.
The curve in concrete terms: Genesis alone introduces 2,781 distinct words, 24.9% of the Bible's whole vocabulary, in its first 33,134 words of text. By the end of the Old Testament you have met 9,897 (88.6%), and Matthew — a whole book — adds only 234 more.
That asymmetry is the argument for finishing something rather than sampling: the hard work is front-loaded. It is also why the site's decks are built per book in frequency order — the first two hundred words of a book carry most of its repetitions, and the rest of it is largely vocabulary you already met.
| After | Words of text read | Distinct words met | Share of the Bible's vocabulary |
|---|---|---|---|
| Genesis | 33,134 | 2,781 | 24.9% |
| Deuteronomy | 127,160 | 5,126 | 45.9% |
| Psalms | 353,239 | 8,362 | 74.9% |
| Isaiah | 403,268 | 9,211 | 82.5% |
| Malachi | 507,198 | 9,897 | 88.6% |
| Matthew | 526,966 | 10,131 | 90.7% |
| John | 577,923 | 10,421 | 93.3% |
| Revelation | 663,876 | 11,168 | 100% |
Questions
- Does reading order change this?
- The fit is for canonical order. Reading the Gospels first meets a smaller vocabulary faster — their 3,986 words already cover 89.6% of the whole Bible's word tokens.
- What is K?
- The constant in K·N^β, fitted here at 17.74. It has no meaning on its own; it scales the curve.
- How was the fit made?
- Least squares on log(distinct words) against log(words read), sampled every 1,000 words across all 663,876 tokens.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).