971 Characters Cover 95% of the Bible
Answering: “how many characters to read the Bible”
Updated 2026-09-22 · plain text · JSON
Character frequency in Chinese follows a steep curve, and Scripture is no exception: a small core carries most of the text. This page states where that curve sits for the Chinese Bible specifically, computed over every character occurrence in the simplified text.
What 95% coverage feels like in practice: roughly one character in twenty is unfamiliar, which is about one per verse at an average verse length of 30 characters. That is comfortable with a tap-to-reveal reader and uncomfortable without one. For unassisted reading, the 1,699-character mark is the one to aim at.
Coverage is not the same as comprehension. Knowing every character in a verse does not mean knowing every word: 22.5% of the Bible's word tokens sit outside the HSK lists, largely names and terms specific to Scripture.
| Coverage | Characters needed | Words needed |
|---|---|---|
| 50% | 73 | 99 |
| 80% | 355 | 825 |
| 90% | 648 | 1,900 |
| 95% | 971 | 3,253 |
| 99% | 1,699 | 6,802 |
Both columns are computed over token occurrences, not over the vocabulary list.
Questions
- Which characters are they?
- The most frequent are 的 他 你 我 们 人 在 和 是 耶 说 不. The site publishes the ranked list, and the 32 characters that occur in all 66 books have a page of their own.
- Where should a beginner start reading?
- John has the highest share of HSK 1-3 vocabulary (72.3%), and the shortest books — 2 John, 3 John, Philemon — use the fewest distinct characters.
- Does pinyin help or hinder?
- Both, which is why it is a column you can switch off. Reading with pinyin keeps you moving; turning it off is how you find out which characters you actually know.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).