中文圣经

The Chinese Bible in Numbers

Updated 2026-09-22 · plain text · JSON

Every page in this section answers one question about the Chinese text of the Bible with a number, and shows the table the number came from. The counts are computed from the 和合本 (Chinese Union Version), simplified script text published on this site — characters from the simplified script, words from its own segmentation, HSK levels from the official lists — and the whole dataset is downloadable below.

The headline figures
MeasureValue
Books / chapters / verses66 / 1,189 / 31,021
Chinese characters921,149
Distinct characters2,996
Distinct words11,168
Characters for 95% coverage971
Words for 95% coverage3,253
HSK 1-3 word tokens59.7%
Not in HSK 1-922.5%
Words occurring once2,579
Distinct pinyin syllables1,117
Easiest book by vocabularyJohn (72.3% HSK 1-3)
Reading time, whole Bible59 hours at 260 chars/min

Questions answered with one number

Reading coverage by HSK level

Vocabulary profile of every book

All 66 books, easiest first by the share of word tokens at HSK 1-3.

Use the dataset

Everything on these pages comes from one file: /raw/corpus-stats.json. It carries the global figures, a profile for each of the 66 books and one for each of the 1,189 chapters. It answers cross-origin, so it can be fetched from a page or a notebook.

Citation: Chinese Bible corpus statistics, chinese-bible.com, generated 2026-09-22. Reuse is free with attribution; the word glosses are CC-CEDICT, CC BY-SA.

Take it further

How these numbers were produced

  • Text: 和合本 (Chinese Union Version), simplified script — public domain66 books, 1,189 chapters, 31,021 verses.
  • Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
  • HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
  • Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
  • Computed: 2026-09-22, from the text on this site.

These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.

Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).