80.2% of New Testament Words Are in the Old
Answering: “vocabulary shared by both testaments”
Updated 2026-09-22 · plain text · JSON
The two testaments were written centuries apart in different languages, but they are read here in one Chinese translation, which means their vocabularies can be compared directly. 5,153 words appear in both. Measured as a Jaccard index — shared words over total distinct words — the overlap is 46.1%.
The asymmetry matters more than the overlap. The Old Testament uses 9,897 distinct words against the New Testament's 6,424, so almost all of the New Testament's vocabulary is contained in the Old, while the reverse is far from true. A reader who works through the Gospels and the epistles first will still meet 4,744 genuinely new words on turning back to Genesis.
Set against the exclusive counts published separately — 4,744 words only in the Old Testament, 1,271 only in the New — this gives the whole picture: a large shared core, a modest New Testament extension, and a long Old Testament tail of names, offerings and measurements.
| Measure | Value |
|---|---|
| Old Testament distinct words | 9,897 |
| New Testament distinct words | 6,424 |
| Words in both | 5,153 |
| Jaccard overlap | 46.1% |
| New Testament vocabulary found in the Old | 80.2% |
| Old Testament only | 4,744 |
| New Testament only | 1,271 |
Questions
- Is the overlap unusual?
- Not for one translation of one canon: a shared translator's vocabulary is exactly what a single Chinese edition produces. The figure describes this edition, not the original languages.
- Which is the bigger vocabulary jump?
- Old Testament first: 4,744 words that the New Testament never uses, against 1,271 in the other direction.
- Where do the exclusive-word figures come from?
- The same dataset, counted by occurrence: a word counts as exclusive when every one of its occurrences falls on one side.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).