「‍」 Lingenic

Markov Chain Text Analysis

(⤓.md ◇.md); γ ≜ [2026-07-17T120407.600, 2026-07-17T135416.643] ∧ |γ| = 3

Markov Chain Text Analysis

Origin. A. A. Markov, lecture to the physical-mathematical faculty of the Imperial Academy of Sciences, St Petersburg, 23 January 1913 (Julian): "An Example of Statistical Investigation of the Text Eugene Onegin Concerning the Connection of Samples in Chains." The target was Nekrasov's claim that independence is necessary for the law of large numbers; the text was the instrument, not the subject.

Mechanism. A sequence of dependent trials still obeys limit laws if the dependence reaches back only one step. Prose supplies a long, cheap, indisputably dependent sequence to demonstrate this on: whether a letter is a vowel is not independent of whether the last one was, and the deviation from independence is measurable rather than argued. Reducing a text to two states throws away everything a philologist cares about, which is what makes the residual dependence attributable to structure rather than to meaning.

Procedure. Fix a corpus and a binary classification of its units. Markov took the first 20,000 letters of Eugene Onegin — the whole first chapter and sixteen stanzas of the second — dropped the hard and soft signs, and classified each letter as vowel or consonant. Write the sequence in 10×10 squares of 100 and count by rows and columns, which turns a serial count into two independent tallies that check each other. Estimate p, the unconditional probability of a vowel; then p₁ and p₀, the probabilities of a vowel after a vowel and after a consonant; then the four transition counts p₁,₁, p₁,₀, p₀,₁, p₀,₀. Compare the conditional estimates against p. Markov found 8,638 vowels against 11,362 consonants, p ≈ 0.43, and conditional probabilities diverging from it sharply enough to refute independence — the chain alternates.

Applies to. Any long sequence with a defensible finite state classification: text, transcribed speech, musical or metrical sequence, symbolic time series. It is the ancestor of every n-gram model, of Shannon's letter-prediction experiments, and of Soviet mathematical linguistics.

Limitations. The classification carries the result: two states see alternation and nothing else, and Markov's own analysis is, as he intended, linguistically superficial — it addresses neither meter nor rhyme nor meaning, and treats the poem as a stream of letters. First-order dependence is an assumption, not a finding; longer memory requires re-tallying against a state space that grows exponentially. The hand procedure is exact and does not scale: Markov counted 20,000 letters personally, and the 10×10 checking device exists because he had no other error control.

© 2026 Lingenic LLC