The Fourth Tone Is 32.5% of the Chinese Bible
Answering: “tone distribution in the Chinese Bible”
Updated 2026-09-22 · plain text · JSON
Tones are counted from the dictionary pronunciation of every word in the text, one entry per syllable, weighted by how often the word occurs. Nothing here is an estimate of how a particular reader speaks; it is the distribution of the tones the text asks for.
The fourth tone leads at 32.5%, then second (á) at 20%. The neutral tone takes 9.9%, which is high for a written text and comes almost entirely from grammatical particles — 的 alone is 48,163 occurrences of neutral-toned syllable.
For a learner the practical reading is about drilling priorities: tone pairs involving the fourth tone are the ones that will come up most often when reading Scripture aloud, and the third tone — the one most learners find hardest — is not rare either at 18.7%. Any verse on this site can be played back with word-by-word highlighting to hear the pattern.
| Tone | Share of syllables |
|---|---|
| fourth (à) | 32.5% |
| second (á) | 20% |
| first (ā) | 18.9% |
| third (ǎ) | 18.7% |
| neutral | 9.9% |
Counted from dictionary pinyin, so a two-syllable word contributes two syllables.
Questions
- Does this match spoken Mandarin generally?
- Roughly, and the differences are what make it interesting: Scripture is formal written prose read aloud, with a high share of particles and a vocabulary of names that a conversation does not use.
- How many distinct syllables are involved?
- 1,117 distinct tone-marked syllables, which is close to the whole inventory of Mandarin — and the reason so many characters sound alike.
- Is the pinyin checked by a person?
- It is generated from the segmented text and the dictionary. That is accurate for ordinary vocabulary and can put the wrong tone on a rare transliterated name.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).