18 Characters Share One Sound: yì
Answering: “homophones in the Chinese Bible”
Updated 2026-09-22 · plain text · JSON
Mandarin has only about 1,200 tone-marked syllables to spell out thousands of characters, so homophony is unavoidable — and this page measures how much of it a reader of the Chinese Bible actually meets. Each character that appears in the text is looked up in the dictionary, grouped by its single-syllable reading, and the groups are counted.
The result: 997 syllables cover the text's characters, 2.34 characters per syllable on average, and the crowded readings run to 18. The top of the list — yì (18), shì (17), jiàn (14), xī (12), fù (12) — is exactly where listening comprehension breaks down and where the written character carries the meaning that the sound cannot.
This is the practical argument for reading with characters visible rather than with pinyin alone: 64.8% of the word tokens in this text are single characters, so a homophone group of a dozen is a genuine ambiguity in speech that the page resolves instantly.
| Syllable | Distinct characters in the text |
|---|---|
| yì | 18 |
| shì | 17 |
| jiàn | 14 |
| xī | 12 |
| fù | 12 |
| fú | 11 |
| bì | 11 |
| hé | 10 |
| jì | 10 |
| zhì | 10 |
Across all 997 syllables the average is 2.34.
Questions
- Does tone disambiguate them?
- Only partly — these groups are already tone-marked, so the 18 characters read yì share both the sound and the tone.
- How do readers cope in speech?
- Context, and two-syllable words: pairing characters up is how Chinese removes most ambiguity, which is why a third of the tokens here are two-character words.
- Where can I see a character's readings?
- Every word page shows pinyin, tone colours and the character breakdown, and any verse can be played aloud.
Related reference pages
Take it further
- 97 ready-made vocabulary decks — by HSK level, by book or by topic, exportable to Anki, Pleco or Quizlet.
- Printable 田字格 worksheets — passages with pinyin, meanings and tracing rows.
- The parallel reader — Chinese, pinyin and your own language side by side, every word tappable.
- The dataset behind these numbers — plain JSON, free to reuse with attribution.
How these numbers were produced
- Text: 和合本 (Chinese Union Version), simplified script — public domain — 66 books, 1,189 chapters, 31,021 verses.
- Words: forward maximum matching against CC-CEDICT (this site's own segmentation). A different segmenter gives slightly different word counts; character counts are unaffected.
- HSK levels: HSK 3.0 levels from the complete-hsk-vocabulary list; band 7 covers HSK 7-9.
- Reading times: 260 characters a minute — an assumption about a fluent adult reader, not a measurement.
- Computed: 2026-09-22, from the text on this site.
These are counts anyone can reproduce from the published dataset. They are not a peer-reviewed linguistic study, and the difficulty rankings are one stated formula rather than a validated readability score.
Definitions: CC-CEDICT, licensed CC BY-SA (https://creativecommons.org/licenses/by-sa/4.0/). If you share this file onward, keep this notice and share alike. Character data: Make Me a Hanzi (Arphic Public License / LGPL).