中文圣经

22.5% of Bible Words Are Not in HSK

Answering: “Bible words not in HSK

Updated 2026-09-22 · plain text · JSON

HSK is a proficiency syllabus for contemporary Mandarin, so it has no reason to include 耶和华 or 祭司. The gap it leaves is measurable: 22.5% of the 663,876 word tokens in the text, drawn from a vocabulary of 11,168 distinct words.

Most of it is names. 662 of the words here read as proper names by their dictionary gloss — people, peoples and places, transliterated character by character for sound. The rest is the vocabulary of worship, sacrifice and law.

The practical consequence: a course gets you to 71.4% and no further, and the remainder has to come from the text. That is what a Bible-specific deck is for, and it is a small list compared to an HSK level — the words repeat constantly once you meet them.

Where the words come from
SourceShare of word tokens
HSK 1-359.7%
HSK 4-611.7%
HSK 7-96.1%
Not in HSK22.5%

Questions

Are the non-HSK words hard?
Not individually — they are just unlisted. They repeat far more often inside Scripture than an HSK 6 word does, so they are learned quickly from reading.
Which HSK list is this?
HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
Is there a deck for exactly these words?
The book and topic decks are built from Bible frequency, not from HSK, so they cover this vocabulary directly.

Related reference pages

Take it further

How these numbers were produced

  • Text: 和合本 (Chinese Union Version), simplified script — public domain66 books, 1,189 chapters, 31,021 verses.
  • Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
  • HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
  • Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
  • Computed: 2026-09-22, from the text on this site.

These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.

Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).