Beinecke MS 408 · computational study
A grammar inferred from the manuscript, not imposed on it.
This tool induces word boundaries, morphology, grammatical categories and syntax rules purely from the transcription, correlates vocabulary with what the illustrations depict, and emits multiple candidate readings with confidence scores. It optimises one thing: how well the inferred grammar predicts folios it has never seen.
Observed
Counted directly in the transcription. Reproducible, no modelling.
22 distinct glyphs, 190,876 glyph tokens, 16 recurrent endings.
Inferred
Output of a statistical model fitted to the observations and scored on held-out pages.
7 word classes, 10.93 bits/word on unseen folios.
Speculative
A semantic guess. Carries no evidential weight and is never used to score models.
Every English gloss in this app. Never used to score a model.
pipeline
How a reading is built
Each stage consumes only the output of the stage above it, so any gloss can be unwound back to raw counts.
1. Glyph inventory
observedCount glyphs and detect recurrent glyph groups by pointwise mutual information.
2. Morphology
observedDiscover 16 prefixes and 16 suffixes that recur across at least 12 stems; segment 96% of word types.
3. Word classes
inferredCluster words by their left/right context, line position and affixes into 7 distributional categories.
4. Syntax rules
inferredScore every class-to-class transition against an independence model; keep the significant ones.
5. Illustration linkage
observedMeasure word × illustration-type mutual information (1.017 bits total) to locate semantic domains.
6. Candidate readings
speculativeFill the inferred slots with domain-appropriate placeholder glosses, and report several competing hypotheses.
corpus
Illustration sections
Pages carry the illustration category recorded in the transcription archive. These labels are inputs, not conclusions.