Not Deciphering The Voynich Manuscript
The Piltdown man was the most successful artifactual hoax of all time. The relevant skull was a chimeric melange of orangutan and human bone fragments from several individuals “discovered” in England in 1912, and meant to represent a transitional form between apes and humans. Given the intense interest in evolutionary biology, it took almost fifty years to discover the Piltdown fabrication. The study of human evolution was derailed. Adding insult to injury, a further sixty years was required to conclusively demonstrate the identity of the culpable fraudster. If you must look up the identity of this scientific blasphemer, go ahead, but we won’t give his name space here.
Students of fraud and fabrication always combat headwinds. In the winter. Uphill both ways.
The evidentiary standards for such claims of fabrication grow higher yet given the passage of centuries and an unknown creator of the subject of interest. And so we arrive at the Voynich manuscript.
It’s a book some 240 pages long, proved by carbon dating techniques to have originated sometime in the 15th century, in the style of the late Italian middle ages. No known author or point of origin. It contains fanciful drawings and is written in an unknown script or language. Today, the physical artifact resides at a Yale library. A small literature devoted to the study of this object has originated, its mystique derived from the uncracked code or untranslated language that make up the book.
My personal view is that this manuscript is a completely meaningless fabrication, the Piltdown man of vellum text artifacts. An elaborate, extended medieval doodle, though it’s impossible to divine the scribe’s intention exactly. This post will marshal and present some of the evidence for such an interpretation, though admittedly, it doesn’t rise to the modern evidentiary standards of fraud. Like I said, uphill, both ways.
Background
A lot of the information about the physical artifact itself comes from René Zandbergen’s wonderfully done website. If you want to go down this rabbit hole, I’d start there. As I mentioned before, there’s also a scholarly literature on Voynich and it’s collected a lot of interest across the internet, for a variety of different reasons. Zandbergen even gives the following piece of helpful advice.
The Voynich manuscript was captured on vellum, a type of parchment made from calfskin. In the process of obtaining data and information for this post, I naturally gathered all the photographed raw page-level images. The book starts out as a collection of unknown, fanciful plants, the so-called “Herbal Section”.
At some point the subject matter switches to what might be described as almost gnostic.
Note the visual motifs repeated on each page. Some are human or animal motifs; in the earlier pages, the plants’ leaves and appendages themselves also repeat, though it looks a little better given that plants actually do have many leaves that look quite similar. Keep an eye on the theme of repetition because we’ll come back to this idea. A central question here is does this art look more like my favorite medieval tapestry, created at about the same time as Voynich
or does it look more like an art class doodle which was not created at that time?
I think you can probably guess my answer to this question. Maybe my comparison of artwork in a vellum book to a giant tapestry is unfair, but you can also look at illuminated manuscripts from the 15th century (books of hours are good examples) to see comparators drawn on a similar medium and format, and with a real language.
Since I’m not an art historian, and don’t even play one on TV, the bulk of this post will focus on the written Voynich elements, not the illustrated ones. They’re much more amenable to statistical analysis. As you can see, the latter pages of the book tend to be full pages of text written in the same mysterious script you see above.
These are not Latin letters in fancy medieval handwriting. They don’t resemble any known written language or writing system. The Voynich text corpus is truly unique. The text features repeated but modified word motifs. Words depend on the words directly preceding them, which suggests a formulaic generative process, not a meaningful excerpt of natural language.
There have been several independent transcription attempts to digitally encode the Voynich manuscript. As far as I can tell, they are all high quality and quite similar to each other which serves our purposes well here, though for this analysis, I used the Zandbergen-Landini transcription.
Statistical Analysis
Given this uniqueness, the hypothesis space about what the written Voynich script means is complicated. Is it an unknown natural language, a known natural language in an unknown script, cipher-text of a known natural language, a shorthand language known to only a select few (maybe just one person), or is it alien language guidebook abandoned on Earth hundreds of years ago, a la the History Channel? I certainly don’t mind entertaining disreputable ideas here. Some of these hypotheses are falsifiable though, and the process here is to systematically evaluate how likely each of these ideas is based on different statistical strategies used to evaluate the Voynich text.
The original work here actually started a year or so ago. After I got started, I did consult some of the computational literature that has accreted about Voynich, and some of what I’m doing here is simply replicating the results of others from the digitally transliterated text, which looks like this in my IDE.
All of the computational work here is my own, even though I’m not necessarily breaking new ground. So first things first. Is the Voynich manuscript comprised of a natural language?
Conditional Entropy
This was the cornerstone statistic of Lindemann and Bowern, 2021, which compares glyph-level conditional entropy across hundreds of real languages to the supposed Voynichese language.
Conditional entropy measures how much unpredictability remains concerning a new glyph once you know the glyph before it. Lindemann and Bowern do a more thorough job than I do going over hundreds of different comparator languages but you can see the money shot in my figure here.
All the real written languages here, even made up ones like Esperanto, fall within a pretty small entropy band here. The lower the conditional entropy figure is, the more previous glyphs predict about the next ones. Voynichese is a massive outlier here, and looks much more like a literal Markov chain than a real language.
Remember, this is a log base 2 scale. The Voynichese 1.3 bit deficit here compared to Latin might not sound like much, but given a preceding glyph, Voynichese behaves as if choosing among roughly 4 equally-likely continuations where Latin has about 10 such choices. Our figure here reproduces Lindemann and Bowern’s Figure 11.
These authors are Voynich believers and think the document presents an authentic, meaningful written text. On the idea of fabrication, they state
The major unsolved question of the Voynich text is whether it represents meaningful language. It could be a medieval hoax that is designed to look like an esoteric alchemical text, in which no sequences of letters correspond to any meaningful words or concepts (see Rugg 2004; Timm and Schinner 2020, amongst others). If so, the creators did an incredible job of imitating the patterns of an authentic language text considering what was known about the structure of language at the time. In so doing, they must have modeled their fake language after a real language that they were familiar with, imbuing it with familiar, language-like patterns.
I don’t know how this statement squares with Figure 12. Their paper is entitled Character Entropy in Modern and Historical Texts: Comparison Metrics for an Undeciphered Manuscript, and so character entropy is their central statistic whose calculation plainly contradicts the statement above. A footnote adds a bit more context and even intrigue.
Timm and Schinner 2021 make this point especially forcefully, even going as far as accusing us of a lack of scientific rigor for not finding their arguments convincing. While the aim of this paper is to make language comparisons across natural and constructed languages, rather than to make arguments about any particular theory, we reiterate a point we have made several times elsewhere (including in the review article that they criticize): that Voynichese appears unnatural only below the word level. At the level of page and paragraph, Voynichese is comparable to natural language and structured text.
Their results show that Voynichese is unusual at the glyph level. The interpretative dispute is whether its regularities distinguish a meaningful language excerpt from plausible text-generation procedures that people like Timm have proposed. So I think they recognize this issue, but prefer to rely on other higher level structural statistics. How informative are those statistics?
Zipf’s Law
This is now a word-level analysis, as Lindemann and Bowern prefer. Zipf’s law states that in natural languages, a word’s rank order is inversely proportional to its frequency in a corpus. So, in a linguistic corpus, a word second ranked in order of frequency is used more often than the third rank ordered word, etc. No one understands precisely why natural languages obey this law, but they do, and Zipf-like laws are a workhorse for some areas of computational linguistics and cryptography. Zipf-like properties are necessary conditions for a corpus to form part of a natural language, but as we will see, they’re not sufficient to establish that.
Unlike with the conditional entropy calculations, Voynichese sits squarely in the natural language band. Right along with other notable classical authors like Dante, Goethe, and an Order 1 Markov process. The point is that Zipf-ness is not strong evidence of this being a natural language.
To my surprise, I don’t think anyone in the Voynich literature has explicitly reported this statistic for the manuscript. Reddy and Knight, 2011 state that it “follows Zipf’s law”, but don’t report the parameters. It’s 1.854, as reported above. Maybe everyone thinks this is so trivial that it’s not worth mentioning, but you heard it here first. I suppose this is an example of what Lindemann and his colleagues mean by “imitating the patterns of an authentic language text”.
Local Dependence
Torsten Timm is the guy who’s done the most work studying the Voynich’s local word dependence. It’s long been known that Voynich seems to have similarly spelled words close together, and the lines and paragraphs look more systematically integrated than you might expect from a natural language. The swashbuckling linguist and cryptographer Captain Prescott H. Currier first noticed this in 1976, and on the structural Voynich issues he raises in his monograph, gives us this gem.
These [structural] findings are definite enough, I think, to warrant much further study by anyone who is going to be involved in seriously attacking the text of the Voynich manuscript. I have no interpretations of them, by the way; I have no solutions. All I know is that they are significant — and damn significant. Anyone who attempts to work on the text without considering these, ignores them at his own peril.
So Currier’s not sure what to make of the book’s text, but Timm builds upon these ideas in 2016 and makes an appearance in the footnote quoted above. He thinks that Voynich is a hoax or fabrication.
In real books, words and close variations of those words can get repeated within paragraphs and within pages. This happens because there’s thematic correlation: pages and paragraphs close together in a book are talking about similar things, and similar things are often described using similar words, or the same words are repeated. But Timm states this isn’t what’s happening here, and provides some pretty good tests proving it. He says
In natural languages a word normally (cf. poems) is used because of its meaning and not because it is similar to a previously written one. The result that the words are arranged such that they co-occur with similar ones is therefore not compatible with a linguistic system. An English text with similar features would mainly consist of words similar to the words “the”, “and” and “to”. Additionally, a word “the” would co-occur with words like “khe”, “phe”, “fhe”, “tha”, “tho”, “thy”, “thee” and “theee”
His contention is that the level of word lookalikes and co-occurrence in Voynichese exceeds what can be explained by thematic correlation or by linguistic inflection if it’s actually meaningful text. My own analysis concurs. Over the pages of the manuscript, I looked how many times you see adjacent word lookalikes. In English, these would look like single edit distance pairs like “book took”, “tick tock”, or “wow how”.
A page whose vocabulary is concentrated in a small number of words produces such pairs by chance. The green series, the same tokens reordered, is the baseline the blue must be read against. Words further apart from each other are not measured here, but we do that in the next analysis. Against a Latin baseline you see this enormous enrichment in adjacent lookalikes, and the effect persists after you control for the vocabulary on a page; Timm correctly points out how strange this is. It’s much better accounted for by a scribe semi-systematically modifying new words using past words as a template. This is supported by some of the inter-line level analysis he does.
So you don’t just see motifs preserved but slightly changed in adjacent Voynich words. You see it vertically between lines as well. It’s as if the scribe looks around what’s recently been written, copies some kind of base template, and makes a modification. I’d remind you that this is entirely of a piece with the conditional entropy calculations above and some of the repetitive, formulaic illustrations.
GRU
The above analysis only looks at adjacent words. I’ve been on a bit of a machine learning applied to old texts kick lately, and wanted to try my hand here. There have not been many ML-themed methods applied to the Voynich problem over the past few years, at least in a scholarly setting. I found this guy on Reddit who tried something similar to what I’ll do here, but with a somewhat different strategy and intent.
GRU is a Gated Recurrent Unit. This neural network architecture is what preceded the modern transformer architecture that make up all the modern LLMs. Like the transformer, this model is trained on text to predict the next token in a string of language. The conditional entropy statistic we examined before only knows about the token immediately preceding what you’re trying to predict. Our GRU can look at a hundred or more glyphs preceding what we’re trying to predict and model long term dependencies. GRUs were replaced by transformers because they can’t be trained at the massive, parallel scale that enables modern AI systems, but the Voynich manuscript is not a massive dataset, and what was old is now new.
I trained this neural network on a random subset of the Voynich manuscript’s pages, along with several control models on corpora of comparable size but with different underlying properties or an entirely different language. The idea is to test out the strength of the local copying effects described before: if you mask a previous Voynich lookalike word, how much worse does the trained model do predicting the downstream lookalike compared to when you do it with a real language? Functionally, this means you make a fake token like [mask], randomly include it during training so the model is accustomed to seeing it, and then if you see “paw” followed by “caw” later in the sentence, “paw” gets replaced with the three letter fake word [mask][mask][mask]. If there are these local copying effects in the manuscript, the Voynich model is going to do worse than the trained control models over many such predictive rollouts, and the effect can be quantified. We get this, the central finding of this post.
So even though the Voynichese GRU model is trained to predictive quality comparable to a GRU model underlying Timm’s preferred copy-method that he asserts explains the generation of the Voynich manuscript or a Markov process, masking a lookalike word rather than a randomly chosen word degrades the model performance much more in Voynichese than it does in a Markov process or a real language like Latin. In other words, the trained model relies on the earlier lookalike in downstream prediction. Timm’s method produces the same effect to a slightly larger degree than Voynichese, but keep in mind it’s entirely synthetic and built to do that.
Hold on there a sec, RBA! Haven’t you heard of inflected languages? Latin is one of them. Nouns and adjectives change from a base form depending on what they’re doing in a sentence. Puer is the Latin word for boy when the boy is a sentence’s subject but it’s puerum when the boy is the object. Those sure look like similar words with a small change and Latin definitely is a real language.
However, the idea that this is morphology doesn’t hold water even under a fairly broad set of potential morphological templates.
The lookalike degradation in Voynichese is highest when the lookalike word pairs differ in the glyphs at the beginning and end, which is unlike real languages like Latin or Esperanto. I was surprised to find this convoluted construction has actually been proposed as the primary inflection mechanism in Voynichese. Reddy and Knight, 2011 state
Several hypotheses about VMS word structure have been proposed. Tiltman (1967) proposed a template consisting of roots and suffixes. Stolfi (2005) breaks down the morphology into ‘prefix-midfix-suffix’, where the letters in the midfixes are more or less disjoint from the letters in the suffixes and prefixes. Stolfi later modified this to a ‘core-mantel-crust’ model, where words are composed of three nested layers.
As with the conditional entropy statistics, I couldn’t find any evidence that any meaningful written language uses this tripartite inflection structure at anything like the Voynichese rate, but as you can see from the figure, this makes sense if Timm is right and this is a copy and edit scheme a knavish medieval prankster is deploying.
Discussion
So if the Voynich manuscript’s text doesn’t look like a natural or even constructed language, what’s left?
On this point, Lindemann’s claim that a 15th century prankster could not possibly have known about Zipf’s law or other statistical requirements for a hoax cuts both ways. This could be a cipher-text of a known or unknown language. It could even be a sophisticated verbose cipher in which a single plaintext character gets carried into multiple cipher-text glyphs, but I don’t think the vocabulary math works out there.
But just as no one had discovered Zipf’s law, a lot less was known about cryptography in the 15th century than is known now. This shrinks the viability of these parts of the hypothesis space quite a bit.
The conditional entropy calculations for the Voynich manuscript we and other investigators have done exclude simple substitution ciphers. What’s interesting here is that the 15th century milieu from which Voynich originated was sort of a turning point for cryptographic technology. People understood that frequency analysis made substitution ciphers very vulnerable to a determined code-breaker and the innovations at the time demonstrate that. Leon Battista Alberti invented polyalphabetic ciphers in 1467. Those actually increase conditional entropy and so I’m skeptical that Alberti’s theoretical breakthrough can rescue the cryptographic Voynich hypothesis but it is notable nonetheless. Mary Queen of Scots used a nomenclator cipher (also increases conditional entropy!) in 1586, over a century after Alberti, had it broken by Thomas Phelippes, and it cost her her head. The Voynich document clearly is not substitution enciphered text from the language we looked at here, or any known language for that matter. Figure 12 is striking, the most important figure in this post.
There are other options. What about a compressed or shorthand language meant for a small enclave of people? What about the most extreme example, a shorthand language for a single person? Wouldn’t that be unbreakable with no secondary reference text? I have some other work addressing these possibilities I might present in a part two of this post, but you’ll have to wait and see.
There’s also a taboo in action that’s worth mentioning here, and I don’t mind breaking it. There is this common fascination with the ancient, the mysterious, the occult, and the Voynich manuscript is the mayor of that space. The zodiac drawings, the secret language, potential hidden ancient wisdom might have a very strong hold on people, even academics at Yale. This is why people watch the History Channel and why Joe Rogan can’t quite ever shut up about how it was aliens or some ancient advanced technology that allowed the pyramids to be built. Even though there were scientists skeptical about Piltdown Man from the very beginning, people wanted Piltdown Man to fill in for missing evolutionary evidence, and so the hoax went on. The conditional entropy calculations make Voynich a massive information theoretic outlier, and even sophisticated people are willing to give it the benefit of the doubt.
Coda
It’s 2026. I’d like to point out that I’m making myself particularly vulnerable calling the Voynich a fabrication or meaningless piece of doodling. Just the past few weeks, the most powerful LLMs have made cryptographic breakthroughs unusually likely. During that time, Simon Klee cracked a hitherto unbroken papal cipher to read a letter that hadn’t been read since the 16th century. Literally the day after that was published, Astra solved an unbroken German cipher from WWI. I promised a part two, but I might look like an idiot really soon.