中文圣经

The Fourth Tone Is 32.5% of the Chinese Bible

Answering: “tone distribution in the Chinese Bible

Updated 2026-09-22 · plain text · JSON

Tones are counted from the dictionary pronunciation of every word in the text, one entry per syllable, weighted by how often the word occurs. Nothing here is an estimate of how a particular reader speaks; it is the distribution of the tones the text asks for.

The fourth tone leads at 32.5%, then second (á) at 20%. The neutral tone takes 9.9%, which is high for a written text and comes almost entirely from grammatical particles — 的 alone is 48,163 occurrences of neutral-toned syllable.

For a learner the practical reading is about drilling priorities: tone pairs involving the fourth tone are the ones that will come up most often when reading Scripture aloud, and the third tone — the one most learners find hardest — is not rare either at 18.7%. Any verse on this site can be played back with word-by-word highlighting to hear the pattern.

Share of syllables by tone
ToneShare of syllables
fourth (à)32.5%
second (á)20%
first (ā)18.9%
third (ǎ)18.7%
neutral9.9%

Counted from dictionary pinyin, so a two-syllable word contributes two syllables.

Questions

Does this match spoken Mandarin generally?
Roughly, and the differences are what make it interesting: Scripture is formal written prose read aloud, with a high share of particles and a vocabulary of names that a conversation does not use.
How many distinct syllables are involved?
1,117 distinct tone-marked syllables, which is close to the whole inventory of Mandarin — and the reason so many characters sound alike.
Is the pinyin checked by a person?
It is generated from the segmented text and the dictionary. That is accurate for ordinary vocabulary and can put the wrong tone on a rare transliterated name.

Related reference pages

Take it further

How these numbers were produced

  • Text: 和合本 (Chinese Union Version), simplified script — public domain66 books, 1,189 chapters, 31,021 verses.
  • Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
  • HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
  • Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
  • Computed: 2026-09-22, from the text on this site.

These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.

Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).