4,744 Words Are Old Testament Only
Answering: “words only in the Old Testament”
Updated 2026-09-22 · plain text · JSON
The two halves of the Bible do not share a vocabulary. The Old Testament is four times the length and brings with it the words of sacrifice, kingship, tribal geography and law; the New Testament, shorter and later, adds a smaller set of its own.
For a learner the asymmetry is good news, because it makes the New Testament a much smaller vocabulary problem: 1,271 words that appear nowhere else, against 4,744 on the other side.
That is one reason reading plans start in the Gospels, and it is visible in the difficulty ranking too — the books with the highest share of elementary vocabulary are almost all in the New Testament.
The halves are not the same size either: 699,012 characters of Old Testament against 222,137 of New, which is most of why the two vocabulary counts differ by a factor of 3.7.
| Part | Books | Characters | Verses | HSK 1-3 tokens |
|---|---|---|---|---|
| Old Testament | 39 | 699,012 | 23,081 | 58% |
| New Testament | 27 | 222,137 | 7,940 | 65.2% |
| Words found only here | 4,744 (OT) | 1,271 (NT) | — | — |
| Shared vocabulary | 5,153 | — | — | — |
Questions
- Which words are shared?
- 5,153 words occur in both testaments — the common core of the text.
- Should I learn them separately?
- That is what the book decks do in effect: each is built from its own book's frequency list, so reading the Gospels never asks you to learn the vocabulary of Leviticus first.
- How is a word counted as belonging to one testament?
- By occurrence: a word counts as Old-Testament-only if every one of its occurrences is in the 39 Old Testament books.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).