中文圣经

80.2% of New Testament Words Are in the Old

Answering: “vocabulary shared by both testaments

Updated 2026-09-22 · plain text · JSON

The two testaments were written centuries apart in different languages, but they are read here in one Chinese translation, which means their vocabularies can be compared directly. 5,153 words appear in both. Measured as a Jaccard index — shared words over total distinct words — the overlap is 46.1%.

The asymmetry matters more than the overlap. The Old Testament uses 9,897 distinct words against the New Testament's 6,424, so almost all of the New Testament's vocabulary is contained in the Old, while the reverse is far from true. A reader who works through the Gospels and the epistles first will still meet 4,744 genuinely new words on turning back to Genesis.

Set against the exclusive counts published separately — 4,744 words only in the Old Testament, 1,271 only in the New — this gives the whole picture: a large shared core, a modest New Testament extension, and a long Old Testament tail of names, offerings and measurements.

Vocabulary of the two testaments
MeasureValue
Old Testament distinct words9,897
New Testament distinct words6,424
Words in both5,153
Jaccard overlap46.1%
New Testament vocabulary found in the Old80.2%
Old Testament only4,744
New Testament only1,271

Questions

Is the overlap unusual?
Not for one translation of one canon: a shared translator's vocabulary is exactly what a single Chinese edition produces. The figure describes this edition, not the original languages.
Which is the bigger vocabulary jump?
Old Testament first: 4,744 words that the New Testament never uses, against 1,271 in the other direction.
Where do the exclusive-word figures come from?
The same dataset, counted by occurrence: a word counts as exclusive when every one of its occurrences falls on one side.

Related reference pages

Take it further

How these numbers were produced

  • Text: 和合本 (Chinese Union Version), simplified script — public domain66 books, 1,189 chapters, 31,021 verses.
  • Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
  • HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
  • Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
  • Computed: 2026-09-22, from the text on this site.

These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.

Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).