The Gospels' 3,986 Words Cover 89.6%
Answering: “do the Gospels cover the Bible's vocabulary”
Updated 2026-09-22 · plain text · JSON
Reading plans start with the Gospels for reasons that have nothing to do with vocabulary, but the vocabulary happens to agree. Matthew, Mark, Luke and John together use 3,986 distinct words — 35.7% of the Bible's full vocabulary — and because those words are the common ones, they carry 89.6% of all its word tokens.
Inside the New Testament the figure is 95.7%, counted directly over New Testament text rather than inferred. What is left is mostly the epistles' theological vocabulary and the names in Acts and Revelation, which a reader picks up book by book.
The two testaments overlap more than their sizes suggest: 5,153 words appear in both, so 80.2% of the New Testament's vocabulary is already present in the Old. Starting in the Gospels therefore costs nothing later — it front-loads exactly the words the rest of the Bible reuses.
| Measure | Value |
|---|---|
| Distinct words in the four Gospels | 3,986 |
| Share of the Bible's vocabulary | 35.7% |
| Share of the Bible's word tokens they cover | 89.6% |
| Share of New Testament tokens they cover | 95.7% |
| New Testament vocabulary also in the Old | 80.2% |
| Words shared by both testaments | 5,153 |
Questions
- Which Gospel should I read first?
- John, by vocabulary: 72.3% of its word tokens are HSK 1-3, and it is 1st of all 66 books by that measure.
- How much of the Old Testament do they cover?
- Less, and the table says why: the Old Testament runs to 9,897 distinct words against the New Testament's 6,424.
- Is there a deck for exactly these words?
- Each Gospel has its own deck, ordered by how often each word occurs in that Gospel.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).