Beinecke MS 408 · computational study

A grammar inferred from the manuscript, not imposed on it.

This tool induces word boundaries, morphology, grammatical categories and syntax rules purely from the transcription, correlates vocabulary with what the illustrations depict, and emits multiple candidate readings with confidence scores. It optimises one thing: how well the inferred grammar predicts folios it has never seen.

36,895
word tokens
5,191 lines
8,433
word types
71% occur once
225
folio sides
labelled by illustration type
11%
perplexity cut vs. no grammar
on held-out folios
observed

Observed

Counted directly in the transcription. Reproducible, no modelling.

22 distinct glyphs, 190,876 glyph tokens, 16 recurrent endings.

inferred

Inferred

Output of a statistical model fitted to the observations and scored on held-out pages.

7 word classes, 10.93 bits/word on unseen folios.

speculative

Speculative

A semantic guess. Carries no evidential weight and is never used to score models.

Every English gloss in this app. Never used to score a model.

pipeline

How a reading is built

Each stage consumes only the output of the stage above it, so any gloss can be unwound back to raw counts.

  1. 1. Glyph inventory

    observed

    Count glyphs and detect recurrent glyph groups by pointwise mutual information.

  2. 2. Morphology

    observed

    Discover 16 prefixes and 16 suffixes that recur across at least 12 stems; segment 96% of word types.

  3. 3. Word classes

    inferred

    Cluster words by their left/right context, line position and affixes into 7 distributional categories.

  4. 4. Syntax rules

    inferred

    Score every class-to-class transition against an independence model; keep the significant ones.

  5. 5. Illustration linkage

    observed

    Measure word × illustration-type mutual information (1.017 bits total) to locate semantic domains.

  6. 6. Candidate readings

    speculative

    Fill the inferred slots with domain-appropriate placeholder glosses, and report several competing hypotheses.

corpus

Illustration sections

Pages carry the illustration category recorded in the transcription archive. These labels are inputs, not conclusions.

herbal
10,703
stars
10,660
biological
6,862
pharmaceutical
2,476
cosmological
2,460
text-only
1,612
zodiac
1,292
astronomical
830