Inferred grammar

Four layers, each derived from the layer above. Glyph statistics and affixes are counted directly; categories and rules are fitted models and are labelled as such.

layer 1 · observed

Glyph inventory and recurrent groups

Character n-grams ranked by pointwise mutual information: sequences that co-occur far more than their component glyphs would predict. These are the candidates for single underlying signs.

Glyph frequency

o
13.3%
e
10.5%
h
9.3%
y
9.2%
a
7.5%
c
6.9%
d
6.8%
i
6.1%
k
5.7%
l
5.5%
r
3.9%
s
3.9%
t
3.6%
n
3.2%

Bound glyph groups

aiin11.1bdaii8.7biin7.5bqoke8.3bqoka8.4bkain8.7bhedy7.4bched7.5bkaii8.1baii6.3bpche8.5bokai7.3bqok6.2bshed7.3bdain7.8beedy6.5bopch7.9bokee6.4bche5.3btaii7.6bain5.9bchol6.5bkeed6.5bchey6.2bedy5.1botai6.8b

Hover for counts. High-PMI groups such as these behave like units: they resist being split by the morphology model and cluster at fixed positions inside words.

layer 2 · observed

Morphology

Affixes kept only if they attach to at least 12 distinct stems that are themselves attested as words. 96% of word types receive a non-empty prefix or suffix.

Prefixes

formstemstokens
o-6895,836
q-4804,864
ch-4693,638
qo-4104,540
d-3572,839
y-3541,364
che-2651,965
sh-2491,753
l-2641,145
ot-2161,674
ol-233898
s-220900
cho-211766
t-211740
ok-1791,800
k-172868

Suffixes

formstemstokens
-y321014,554
-n10535,931
-dy6035,661
-ey3363,487
-ar3691,580
-r3263,482
-aiin3441,937
-l3053,869
-ol3112,311
-s317796
-al2911,293
-or2851,460
-in2444,608
-edy2333,729
-o224883
-ain2011,120

Most productive stems

ai60ch146ke81ee121ed49ka52eo54da44ol71ta35sh90ckh62cth62al67ar54te45od61kai21or46ko41kch57pch65ok51che71to33he5ot35tch43ho6do25dai19dy11qo8s1ek45kee46

layer 3 · inferred

Word classes

Seven clusters over context, position and affix features. The role names are descriptions of measured behaviour — they are hypotheses about grammatical function, not identifications of parts of speech in a known language.

C1 · modifier-like

inferred
types
107
tokens
12,020
line-initial
8%
ending
-y

frequent members

daiinolchedyshedycheyqokeeyqokeedydarqokainsheyqokedyqokaiindaldy
evidence trail
  • observeddistributional counts

    12,020 tokens across 107 word types; 7.8% line-initial, 9.2% line-final, mean relative position 0.51.

  • observedmorphological signature

    dominant ending "-y" covers 27.5% of class tokens; mean word length 4.85 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

C2 · predicate-like (verb / attribute)

inferred
types
341
tokens
3,246
line-initial
6%
ending
-dy

frequent members

ytaiinoteodyokeodychedaiinqokeodyqokcheyopcheyokamotchdyokchedyorainqokamqotchedyqotchdy
evidence trail
  • observeddistributional counts

    3,246 tokens across 341 word types; 6.3% line-initial, 22.1% line-final, mean relative position 0.58.

  • observedmorphological signature

    dominant ending "-dy" covers 22.8% of class tokens; mean word length 5.53 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

C3 · opener / topic marker

inferred
types
139
tokens
1,325
line-initial
59%
ending
-ol

frequent members

dshedytchedyyteedydchedydchorpolycheeydcholsoiintchorydaiindoiindcheyokeeol
evidence trail
  • observeddistributional counts

    1,325 tokens across 139 word types; 59.3% line-initial, 4.8% line-final, mean relative position 0.22.

  • observedmorphological signature

    dominant ending "-ol" covers 17.4% of class tokens; mean word length 5.61 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

C4 · function word / connective

inferred
types
62
tokens
3,757
line-initial
10%
ending
-

frequent members

aiinoraralsotardaircheorainamrsarairo
evidence trail
  • observeddistributional counts

    3,757 tokens across 62 word types; 9.8% line-initial, 11.1% line-final, mean relative position 0.51.

  • observedmorphological signature

    dominant ending "-∅" covers 40.0% of class tokens; mean word length 3.12 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

C5 · modifier-like

inferred
types
67
tokens
919
line-initial
17%
ending
-r

frequent members

otorokorykarsharqotorsairkorytarodaraiirtordaiirokeorqor
evidence trail
  • observeddistributional counts

    919 tokens across 67 word types; 16.8% line-initial, 7.2% line-final, mean relative position 0.46.

  • observedmorphological signature

    dominant ending "-r" covers 100.0% of class tokens; mean word length 4.34 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

C6 · closer / terminal marker

inferred
types
127
tokens
1,684
line-initial
9%
ending
-y

frequent members

olyotchyshdyodyckhychotyokchyshcthyalychecthydalykchydchyary
evidence trail
  • observeddistributional counts

    1,684 tokens across 127 word types; 8.6% line-initial, 24.1% line-final, mean relative position 0.59.

  • observedmorphological signature

    dominant ending "-y" covers 100.0% of class tokens; mean word length 4.74 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

C7 · modifier-like

inferred
types
109
tokens
4,160
line-initial
11%
ending
-y

frequent members

cholchorsholychyshocthyshyshordamokyqotyotolokol
evidence trail
  • observeddistributional counts

    4,160 tokens across 109 word types; 10.8% line-initial, 9.6% line-final, mean relative position 0.47.

  • observedmorphological signature

    dominant ending "-y" covers 18.9% of class tokens; mean word length 4.48 glyphs.

  • inferredcategory assignment

    k-means (k=7, cosine) over left/right context vectors built from the 40 most frequent forms, plus positional and affix features. The role label follows from the profile above; no natural language is assumed.

layer 4 · inferred

Candidate syntax rules

Every class-to-class transition compared against what independent placement would predict. z is the standardised deviation; ⟦line edge⟧ stands for the start or end of a written line.

Strongly preferred sequences

ruleobservedexpectedz
⟦line edge⟧C3786163+48.7
C4C41,352753+21.8
C1C17,7316402+16.6
C2⟦line edge⟧1,001618+15.4
C7C71,096708+14.6
C4⟦line edge⟧1,077694+14.5
C6⟦line edge⟧406208+13.8
C5C4309183+9.3
⟦line edge⟧C5246169+5.9
C3C7237172+5.0
⟦line edge⟧C7799673+4.8
C5C57245+4.1

Avoided sequences (grammatical constraints)

⟦line edge⟧⟦line edge⟧0640-25.3
C4C11,5452196-13.9
C1C3234517-12.4
C1C41,6812196-11.0
C4C348177-9.7
C1C71,7572129-8.1
C3⟦line edge⟧64163-7.8
C2C383158-6.0

Under-represented transitions are as informative as preferred ones: a system that forbids certain orderings is behaving like a grammar rather than like a word-salad generator.