Translation note. Translations are by the author unless otherwise credited.

The theory in one proposition

The square outline of a Chinese character is not the smallest functional unit of the writing system. Inside it, recurring graphemes, 部件 bùjiàn, combine according to graphotactic rules. These graphemes can indicate a semantic field, a family of possible sounds, a structural function or more than one of these at once.

The resulting character does not operate as a picture that communicates an idea without language. Nor is it an indivisible sign whose internal form is irrelevant to reading. It works as a structured mnemonic trigger. A semantic cue narrows the field of meaning; a phonetic cue activates a region of phonological space; knowledge of the language and the textual context selects the relevant morpheme.

Some writing-system typologies call Chinese morphosyllabic because a complete character commonly writes a morpheme and a syllable. That observation concerns the relation between the completed graph and a linguistic unit. It does not explain the internal mechanism by which the graph is built or recognised. The phonosemantic account proposed here is a theory of that mechanism.

A mature system at its first attestation

The earliest extensive Chinese writing currently known appears in late Shang oracle-bone inscriptions from around 1250 BCE. It is already a fully mature writing system. The inscriptions can record names, dates, questions, events and complete propositions. They employ phonetic borrowing, semantic differentiation, recurring components and conventional spatial arrangements. Nothing essential to the writing-system type is waiting to be invented.

Earlier development must have occurred, perhaps on materials that did not survive. Neolithic marks, clan signs and later traditions concerning knotted cords are relevant to the search for precursors, but their relationship to the mature script remains unproved. That unresolved prehistory does not turn the attested oracle-bone system into an immature stage.

The core principle visible there has remained stable. Seal, clerical and regular forms altered how graphs were written. Qin and later authorities standardised inventories and shapes. Cursive usage, printing, type design and twentieth-century simplification changed graphic forms and writing practices. New characters and new conventions were added. These developments are comparable to spelling reform, font history, abbreviation and standardisation in other scripts. They did not transform an incomplete system into a fundamentally different one.

This is the central historical claim of the article: the documented system begins fully mature, and no later development has replaced its phonosemantic and graphotactic foundation.

Graphemes, strokes and the quadriform space

Strokes are the physical movements or graphic segments used to write. Graphemes are functional units. 水 and its left-position allograph 氵 illustrate the distinction. The form compresses under spatial pressure, but its identity within the system remains recognisable. The same applies to 心 and 忄, 火 and 灬, 刀 and 刂.

Components are constrained by position and proportion. Some favour the left, top or enclosing position. They compress or expand to preserve the conventional square. The character is therefore not a bag of parts. It is a two-dimensional construction in which position can help identify a grapheme's role.

The traditional category 形声 xíngshēng recognised the combination of a meaning-related element and a sound-related element. The postface to Shuowen jiezi, compiled around 100 CE, gives 江 and 河 as examples:

形声者,以事为名,取譬相成,江河是也。

In phonosemantic formation, a category is named from the matter concerned and a corresponding sound is taken to complete it, as in 江 and 河.
Translation by the author.

The translation remains subject to the author's final philological approval. The passage matters because it records an early analysis of the same structural fact. The received Shuowen assigns 7,697 of its 9,353 analysed entries to 形声. Modern percentages vary with the corpus and criteria, but the numerical disagreement does not change the predominance of phonosemantic construction.

Sound is familial, not phonemically exact

The phonetic grapheme does not function like a letter. It need not prescribe one modern pronunciation. It identifies a phonological family through relations in initials, rhymes, tones, historical voicing or combinations of these features.

青 provides a familiar modern example.

Character Graphemic structure Standard Chinese reading Broad semantic field
氵 + 青 qīng water, clarity, cleanliness
日 + 青 qíng sun, light, clear weather
忄 + 青 qíng feeling, mental or relational state
讠 + 青 qǐng speech, request
米 + 青 jīng grain, refinement, essence

Table 1. Authorial explanatory table. It presents a synchronic structural family, not a complete etymology of each graph.

An annotated phonetic series built around the component 青.
Figure 1. Authorial explanatory diagram. 青 activates a sound family while the accompanying grapheme differentiates broad semantic fields. Open full size

Modern readings do not exhaust the evidence. Sound change in Chinese has been highly patterned. Members of a phonetic series may separate predictably within a variety while remaining related across historical stages and present-day Chinese varieties. The component points into a network; it does not transcribe one prestige pronunciation fixed for all time.

This familial model also accommodates derivation and polysemy. A graph can hold related readings or functions because its phonetic cue activates a morphophonological family rather than a single phoneme string. The reader does not mechanically pronounce the components. Recognition proceeds through the interaction of visual structure, the living lexicon and context.

Why the classification changes teaching

Treating every complete character as an independent symbol makes literacy appear to require several thousand unrelated acts of visual memorisation. The phonosemantic model identifies a smaller recurring inventory and the rules by which it combines.

Mastering a few hundred high-value graphemes does not by itself provide vocabulary, grammar or textual knowledge. It provides structural access to the great majority of commonly encountered character forms. That is the intended claim. It distinguishes learning the system from learning every word written with it.

A grapheme-based curriculum would therefore begin with functional components, their allographs and their legal positions. Character families would then teach:

  • semantic fields built around recurring graphemes;
  • fuzzy phonetic families rather than isolated readings;
  • correct component choice and placement, which together constitute Chinese spelling;
  • recursive structures in which a compound unit can function as a grapheme inside a larger character.

Research on component awareness supports parts of this pedagogy. Experiments have found that explicit instruction in semantic components can improve learners' inference of unfamiliar character meanings, while phonetic-component awareness contributes to character acquisition. These results do not adjudicate the whole theory. They show that readers use information below the complete-character level.

Cognitive efficiency as a testable proposition

The article's cognitive claim is stronger than the observation that characters have components. A phonosemantic character packages several cues into one spatial unit. Semantic field, familial sound and configuration become available together rather than through a purely linear sequence. This may create mnemonic and processing advantages in contexts where associative retrieval and morphemic distinction matter.

That proposition does not declare one script universally superior. Alphabetic and phonosemantic systems optimise different relations between sound, meaning, sequence and space. The claim is that Chinese writing possesses its own efficiencies and must not be measured as an unsuccessful approximation of alphabetic transcription.

The strongest versions of the cognitive argument still require direct testing. Suitable studies would compare grapheme-aware and whole-character teaching, track recognition of lawful and unlawful component arrangements, and test how phonetic families behave across Chinese varieties. Until then, cognitive efficiency remains a prediction of the theory supported by partial experimental evidence, not a settled numerical ranking of scripts.

Unicode records an inventory, not the generative system

Unicode 17.0 contains more than 100,000 unified Chinese, Japanese and Korean ideographs. This is an extraordinary achievement in textual interchange. It is also an encoded inventory of abstract characters. Each accepted character receives an interoperable identity; a font supplies one or more glyphs for it.

That design leaves the generative structure of the script outside the encoded character. Unicode can store 清 as U+6E05. It does not encode 清 as the lawful composition of 氵 and 青. A component-aware database can add that analysis, but ordinary Unicode text does not derive the character's identity from it.

Ideographic Description Sequences do not close the gap. Unicode 17.0 states that a sequence such as ⿰氵青 is a description, not a formal encoding. There is no canonical description, no assigned semantics, no defined equivalence with the encoded character and no requirement that a renderer create the intended graph. IDS can be useful for search and analysis, yet it does not turn an unencoded graph into a stable, interoperable character.

Three panels distinguish a Unicode code point, contrasting Noto Serif SC and KaiTi glyphs, and a left-right component analysis.
Figure 2. Authorial explanatory diagram. A code point identifies an encoded unit, fonts supply differing glyphs and a component analysis describes structure. The analysis does not itself create an interoperable encoded character. Open full size

The limitation is visible whenever scholarship encounters a graph outside the current repertoire. A private-use code point can display a local glyph when cooperating users share a font and an agreement. The same code point can mean something else elsewhere. Sending the text without its private agreement does not send a standardised character identity. An image preserves appearance but loses ordinary textual search, collation and interchange. A local font solves display for one environment, not the model or interoperability problem.

Four provisional treatments of an encountered unencoded graph and the interoperability gap that remains.
Figure 3. Authorial explanatory diagram based on Unicode 17.0 technical documentation. An IDS, a private-use assignment and an image can each preserve some information, but none by itself supplies a standard interoperable character identity. Open full size

Tangut material

Tangut demonstrates that a large completed repertoire does not end discovery. Unicode 9.0 encoded 6,125 Tangut ideographs and 755 components. Later work identified additional ideographs, components and glyph corrections. Unicode and ISO working papers from 2023 to 2025 describe newly identified, explicitly unencoded Tangut material and propose further code points. One 2025 proposal reconstructs a graph from a damaged lexicographic witness precisely because no encoded character matched it.

Tangut IDS also exposes structural limits. Technical feedback submitted to Unicode has requested additional description operators for arrangements that the existing IDS vocabulary cannot express adequately. Even a successful description would still not confer an interoperable character identity. The research cycle remains: encounter a graph, analyse it, document it, propose it, wait for standardisation and only then exchange it as an ordinary encoded character.

Jianzipu

Guqin tablature, 减字谱 jiǎnzìpǔ, combines reduced fingering signs in dense two-dimensional clusters. The possible written units are productively assembled rather than selected from a small fixed list of whole symbols. A 2019 encoding proposal introduced format controls and hundreds of precomposed musical symbols. Subsequent technical discussion continued because the central issue is the relation between components, layout and an interoperable textual unit.

As of August 2026, jianzipu is still not an ordinary encoded notation in Unicode 17.0. Unicode working documents describe proposed component blocks and format controls, with standardisation discussed for a later version. In practice, projects use images, private-use assignments, ASCII or Chinese input strings and custom OpenType behaviour. These solutions can render music within a prepared environment. They do not establish one portable textual identity that survives without the same software, font and private convention.

Tangut and jianzipu differ from ordinary modern Chinese text, but they reveal the same modelling boundary. An inventory of atomic encoded units cannot straightforwardly represent every graph produced or encountered in a spatially compositional tradition.

A grapheme-based digital layer

The alternative proposed here is a grapheme-based encoding and rendering layer. A character or notational graph would be represented through identifiable components, their allographs, structural relations and graphotactic constraints. The composed result could then be exchanged as a structured textual object rather than as an image or a privately assigned glyph.

Such a system would need more than the current IDS operators. It would require:

  • a standard inventory of functional graphemes and their positional allographs;
  • expressive spatial relations, including overlap and script-specific configurations;
  • canonical composition rules where interchange requires equivalence;
  • a distinction among character identity, graph instance and glyph styling;
  • rendering and input systems that preserve the structured representation.

This would not require Unicode to be discarded. Existing code points remain necessary for ordinary text and backwards compatibility. A compositional layer could supplement them, map composed structures to encoded characters when available and retain a usable structure when no atomic code point exists.

Component-based input methods already show that Chinese users can select characters through form rather than Latin-letter pronunciation. Their present output remains limited to the encoded repertoire. A truly compositional method would build and exchange lawful graphs, including rare historical forms, without waiting for every complete combination to receive an atomic code point.

What follows from the phonosemantic model

Chinese writing is not a collection of pictures, an inventory of opaque word signs or an arrested stage on a path toward the alphabet. It is a mature phonosemantic system based on finite graphemes, fuzzy sound families, semantic fields and two-dimensional graphotactics.

The model explains why the script can remain structurally stable while pronunciations change, why character families are available to learners and why an atomic digital inventory never fully captures its productive capacity. It also produces claims that can be tested and refined: the size and function of the grapheme inventory, the cognitive effects of spatial packaging, the regularity of phonetic series across varieties and the requirements of interoperable composition.

The square is not the absence of structure. It is where the structure has been compressed.

References

  1. Kulmala, Toni. The Chinese Writing System: A Phonosemantic Script. Unpublished doctoral research manuscript. Cited as the theoretical framework of this article.
  2. Qiu Xigui. Chinese Writing. Translated by Gilbert L. Mattos and Jerry Norman. Society for the Study of Early China and Institute of East Asian Studies, University of California, Berkeley, 2000. Used for palaeographic and documentary comparison, not as the adjudicating framework.
  3. Xu Shen. Shuowen jiezi, postface, c. 100 CE.
  4. Nguyen, Thi Phuong, Jie Zhang, Hong Li, Xinchun Wu and Yahua Cheng. “Teaching Semantic Radicals Facilitates Inferring New Character Meaning in Sentence Reading for Nonnative Chinese Speakers.” Frontiers in Psychology 8 (2017): 1846.
  5. Shen, Helen H. and Chuanren Ke. “Radical Awareness and Word Acquisition Among Nonnative Learners of Chinese.” The Modern Language Journal 91, no. 1 (2007): 97–111.
  6. The Unicode Consortium. “Chapter 18: East Asia.” The Unicode Standard, Version 17.0, 2025. Cited as technical documentation of the current model.
  7. The Unicode Consortium. “Private-Use Characters, Noncharacters and Sentinels FAQ.” Accessed 18 August 2026.
  8. West, Andrew and Viacheslav Zaytsev. “Tangut Character Additions and Glyph Corrections.” Unicode Technical Note 42, version 2, 2019.
  9. West, Andrew. “Proposal to encode one newly-identified Tangut ideograph.” ISO/IEC JTC 1/SC 2/WG 2 N5314, 2025.
  10. West, Andrew. “Feedback on Proposals to Encode New Ideographic Description Characters.” L2/22-136, 2022.
  11. Chen, Zhuang. “Proposal to Encode Jianzi Musical Notation.” ISO/IEC JTC 1/SC 2/WG 2 N5041, 2019.
  12. Unicode CJK and Unihan Working Group. “Recommendations for UTC Meeting 186.” L2/26-009, 2026.
  13. CHISE Project. “CHaracter Information Service Environment.” Accessed 18 August 2026. Cited as an example of component-aware research infrastructure.