Using LLMs to trace alchemical knowledge and decode 17th century letters

Using LLMs to trace alchemical knowledge and decode 17th century letters 图片 1

I’ve written previously about the pitfalls and use cases for AI in augmenting historical research, but things have changed significantly since 2024-25. Occasioned by the dueling releases of GPT-6 Sol and Opus 5.5 this week, I thought I’d share some early results with using these models not just to perform “research assistant” type functions like transcribing documents, but to try to actually solve existing historical problems.

The TLDR is that pairing historians working in collaborative groups with the current frontier models would, in my view, produce numerous advances in historical knowledge and interpretation. My guess is that many of these could end up being quite meaningful. This was not the case as recently as last year. I think AI labs, historical researchers, and funding agencies should start actively pursuing these collaborations.

Finding traction

As we’ve seen with the field of mathematics, these models do best when they have a set of problems that LLMs invariably tend to describe as “tractable.” In other words:

• Have experts in the field already identified a group of problems that need solving?

• Is the data needed to answer these problems fully digitized and accessible?

• Do the problems lend themselves to the “spiky” capabilities of frontier AI models — namely multilingual reasoning, advanced math, and/or ability to conduct autonomous research through large datasets or across disciplinary subfields?

• Are they amenable to solutions that involve writing bespoke code?

• Most importantly: can a potential solution be clearly proven or disproven? (This last one, it seems to me, is a key part of why reasoning models have run rampant in mathematics but not in humanistic fields).

The above factors mean that the types of historical “open problems” which frontier AI can reasonably be expected to help with are fairly constrained:

  1. Anything involving cryptography and codebreaking (For instance, see Astra decrypting a 1941 German army communication and a WWI German radio cipher, or the work that Daniel Bourdeau has been doing here, or my own attempt to use GPT-6 Astra to figure out what is going on with the Elizabethan occultist John Dee’s coded magical book, Liber Loagaeth).
  2. Tracing texts across translations and adaptations. As an example of this, I was able to use GPT-6 Astra to determine the identity of a passage that Isaac Newton had freely translated into Latin from a French alchemical text, an identification that seems to have not previously been made.
  3. Drawing links between existing findings that are reported only in discrete or niche subfields, or are not yet integrated into scholarship.

This last one might end up being the most impactful new method that these tools open up for historical researchers. For instance, if you read the writeup of Astra breaking a July 10, 1941 Enigma message that had resisted decipherment, it turns out that the key breakthrough was not anything to do with the codebreaking itself, but with noticing the full range of information that was available. Historical cryptological researcher Frode Weierud writes:

We are still analysing the GPT–6 Astra logs to see exactly how it executed the break. And we are discovering amazing details. In July 2026, I made the following announcement on the webpage with the 1941 Message List:Note: In July 2026, research in the German Bundesarchiv revealed several
collections of radio messages, both enciphered and in cleartext. One of
these message collections was from SS-Totenkopf Division’s logistics
command, Nachschubführer. Many of these messages were sent to the Ib
(Quartiermeister) radio station and are identical to those in this list.
Others are new, but most likely related. These new messages are added to
the 1941 Message List in bold, with the indicator NF (Nachschubführer)
after the message number, indicating that these message numbers belong
to the NF numbering. All NF messages are outgoing; hence, the message
numbers are in blue.
It appears that GPT–6 Astra discovered this note about the collections of radio messages at the German Bundesarchiv.

What’s fascinating about this note is that even the leading human experts don’t entirely understand what GPT-6 Astra did as it gathered together these bits of information and used them to find a solution. Weierud writes:

The file references GPT–6 Astra mentions, RS 3–3/20a and RS 3–3/63b, are correct, but they are not available on the Crypto Cellar Research website. GPT–6 Astra mentions a private collection, but it is not clear what this is, whether it has succeeded in accessing the Bundesarchiv’s digitised collections or whether it has found these files elsewhere.

Shades of the Hugging Face incident here: these models are maniacally determined when giving a problem they deem tractable. They will push their search for potential solutions as far as they possibly can, often in ways that human experts find difficult to trace.

What can be done now

I mentioned above that I tried to using GPT-6 Astra to “solve” John Dee’s coded manuscript, Liber Loagaeth. Dee is one of my favorite historical figures ever, and if you haven’t heard of him, I recommend his Wikipedia page — his story is endlessly fascinating and weird. Among other things, Dee is thought to have influenced both Shakespeare’s depiction of the wizardly Prospero in The Tempest and Christopher Marlowe’s portrayal of the devil-bargaining Faust in Doctor Faustus.

One of the weirdest parts of a very weird life was Dee’s work with the “scryer” Edward Kelley to transcribe what he called a “book of mystery” which was written in the “angelicall language” (Dee believed that Kelley was, in effect, a prophet who was receiving new works of divine revelation written in code). You can read a full transcription of this book here.

Astra’s verdict, which I think makes sense given that Kelley was pretty clearly a charlatan, is that the supposedly coded book is not in code at all: it is almost entirely nonsense syllables. It created a report of its findings here.

However, the model’s analysis did yield a few interesting things. For instance, it was able to cross-check its mathematical analysis of how often characters repeat in the text to the evidence from John Dee’s diary. It concluded that Kelley started getting increasingly lazy after a specific date and began repeating himself more:

Astra was also able to determine that one passage of this apparent gibberish actually did encode meaning: a reference to Bornogo, one of the angelic beings in what we might call the “John Dee cinematic universe” of invented mythology.

Is this a meaningful breakthrough in John Dee studies? No. And it’s worth acknowledging that even a genuine breakthrough in a niche historical subfield like this is far from an equivalent to solving Navier-Stokes.

But - this sort of thing is, I think, a genuine sign that expert historical knowledge combined with frontier models and a lot of compute can yield unexpected results.

Three quick case studies

I initially threw Astra and Opus 5.5 at the challenge of finding more WW2 and WW1 era encrypted messages to solve, but the low hanging fruit here seems to have been plucked — they came up empty (although it was fascinating seeing how they trolled through lists of German troop rosters to find plausible names to check).

Darwin’s monkey tails

I started getting better results when I moved into my own wheelhouse as a specialist in the history of science and medicine. As I write, GPT-6 is currently working through the writings of Charles Darwin and searching his references to where he gathered information relating to natural selection; the idea is to find undiscovered links in the chain of knowledge between Darwin and his informants. Interestingly, this was an idea that GPT-6 suggested on its own. However, it is actually a good match with my professional intuition about what would constitute a worthy research project (somewhere on the spectrum between a research paper and a PhD dissertation, in terms of potential payoff) using this material. In the past, AI models struck me as lacking this ability to independently conceive of worthwhile historical research projects at this scale — they were more useful for, say, making data visualizations.

Here is an example of the model’s reasoning traces as it contemplates whether to continue to research a reference to a kangaroo larynx in one of Darwin’s notebooks!

This one is currently in progress and hasn’t yielded anything worth mentioning yet as a decisive result, but I think it’s a good example of how the very patient, collaborative work of historical researchers and archivists — namely the team behind the wonderful Darwin Correspondence Project — can serve as a foundation for emerging research methods. It’s certainly the case that humans can, and have, traced the references to named figures in Darwin’s notes and letters, but the multilingual nature of language models makes me suspect that they will be able to find new links here, especially in extremely large corpora of sources that are beyond the ability of any one human to read in full.

Tracing the transmission of alchemical knowledge

Another great candidate: the papers of Samuel Hartlib, the self-described “intelligencer” who was an influential early member of the Royal Society and a key node in the network of early modern science. These are fully digitized, they are drawn from sources in several languages, and they span a wide range of academic fields and intellectual niches. All of which means they are unusually tractable for a frontier model.

Opus 5.5 set to work downloading over 5,000 primary source files from Hartlib’s archive, then created sub-agents to troll through Google Books and other archive sites to cross check the unidentified sources of Hartlib’s information across different languages. The goal was to find moments when Hartlib had received important scientific information from an anonymous or unidentified source, and then discover that identity.

Opus is actually still working through this as I write, but a preliminary report is written up here. The top findings are, in my view, real and meaningful. Not earth-shattering by any means, but the sort of thing I could imagine spending a week of research on.

Did you catch the bit about the anagram? This is where the reasoning/math ability of these models becomes relevant: Opus 5.5 noticed that both Newton and Hartlib used different anagrams/codes for the key ingredient, Hungarian vitriol. This sort of coded language is common in early modern alchemy, but I certainly never would have noticed it. Opus explains:

Newton’s is a true anagram. “Vltimorui” uses exactly the letters of vitriolum (v-i-t-r-i-o-l-u-m), rearranged. The Newton edition’s editors identify it that way.Hartlib’s is closer to backwards writing, and even that is imperfect. Reverse each word of Miloirtiua riciragnun letter by letter and you get: Miloirtiua → auitriolim, close to uitriolum (= vitriolum)riciragnun → nungaricir, close to ungaricum

This felt like a stretch to me, but it further clarified things by sharing the specific marginal annotation that had been written to clarify this for 17th century readers as well:

So what Opus identified here was not just the anagram for a key alchemical ingredient, but more importantly, the parallel between both Newton and Hartlib employing anagrams for it. This, along with the same quantities being described by both, and other matches across the texts, seems to me to be very compelling evidence that Hartlib’s manuscript was the one Newton drew upon.

As far as I can tell, this actually is a new finding, and given Newton’s historical significance, it may be one that would merit publication, especially if it can be fleshed out with other findings along the same lines.

Deciphering early modern coded letters

A final case study: literally while I was writing this post, Opus 5.5 partially deciphered two 16th century Spanish letters written in the secret code of Emperor Charles V:

The catch? Both had already been deciphered! One had been decrypted back at the time of authorship, in the 1530s, with the plain text written in a set of pages that followed the coded ones. The second, after some digging through Google Books, turned out to have been deciphered in 1916.

This was a good example of the importance of expertise and “desk research,” since (being a complete amateur when it comes to historical cryptography) I could easily have wasted several more hours duplicating the work of a careful scholar well over a hundred years ago. At the same time, it was also a great test case for determining that Opus 5.5 really is capable of doing this sort of work, since it was able to verify its own interpretations as correct once it found the “gold standard” plain text from 1916. Below is a chart Opus made showing this, and a complete website it created with a writeup of that work:

It’s worth mentioning again here that Daniel Bourdeau has an amazing website collecting open problems for historical cryptography and documenting his attempts to use these same models to solve them. It’s a great guide for this sort of thing.

What next?

The obvious next step is not people like me using up their personal Codex and Claude Code allowances each week poking around in this haphazard way. It’s a systematic effort based on collaborative research and sharing of information between historians, archivists and other researchers, and I think it’s time for the major AI labs and foundations to start funding and assisting this work.

Why? So much of what frontier models can currently accomplish is because they have access to publicly accessible primary sources. They can make breakthroughs in, say, cracking an Enigma cipher because enormous volunteer effort has gone toward making these documents transcribed and available online, and because so much collaboration has happened between humans to establish what questions should be asked, what the problems are.

For now, the results for historical research, archives, and related fields (like archaeology) are going to be much more scattershot and limited than what we’ve seen in math. That’s partly a matter of what these models find tractable, and it’s true that mathematical proofs are just fundamentally different from how historical knowledge is amassed. But I think three key interventions would move the needle toward real breakthroughs in the field of history:

  1. Collaborate across libraries and archives to digitize unavailable historical manuscripts and make them freely accessible online. Repeatedly, in my testing, the bottleneck turns out to be access to archival documents. These are often digitized but are not available unless you have privileged access. Relaxing these restrictions would go a long way, but it’s even more important to remember that the vast majority of premodern historical manuscripts remain undigitized. This is a very solvable problem that just needs institutional will and funding.
  2. Providing historians with free API access/compute. I might be wrong, but I don’t think anyone actually knows what happens when a medium to large amount of compute (on the order of hundreds or thousands of agents) is thrown at active historical problems.
  3. Historians can band together to identify “millennium problems” just as mathematicians have. I should clarify here that the major debates in historical scholarship have nothing really to do with “solving problems” or “disproving theorems” — again, history is just fundamentally different from math or physics in this way. The things that historians get passionate about, and devote our careers to, are often issues of interpretation and subjective analysis that have no single “solution” at all. But — there also are actual mysteries that could be solvable if sufficient attention and resources were devoted to them. John Dee’s Liber Loagaeth is one: does it encode more meaningful information than the snippet the AI was able to spot? Quite possibly - we just don’t know right now. The famous Voynich manuscript may be another, although I personally believe it likely has no semantic information at all (my theory is that it’s the product of an early modern person suffering from graphomania). And then there’s Linear A, and all the still-encrypted historical primary sources, and on and on…

I’m intrigued enough by all this that I am planning on emailing historian friends and colleagues to create an informal survey of which “open problems” in history they think would lend themselves best to this sort of approach. The list would then be made publicly available as a list on a website. Please get in touch if you’d like to be involved in this:

Email me

Clearly, there will be more advances in historical code-breaking from these models. But what interests me is what additional forms of historical knowledge that general set of skills can uncover. In other words, the problem space around actual cryptography.

Personally, I suspect that issues relating to provenance, quotation (including previously undetected cases of historical plagiarism!) and influence across languages and genres are going to be where frontier models end up being most useful.

But this is where pooling the expertise of historians and archivists, and getting direct input from AI researchers, is most helpful. There are so many offshoots of historical knowledge that lead in niche directions that it’s impossible for one person to actually know what questions to ask.

As an example, GPT-6 Pro has spent the past several hours churning through a 17th century Sanskrit astronomical text (the Karaṇakesarī of an astronomer named Bhāskara) trying to reconstruct the algorithms Bhāskara used to model solar eclipses.

Is this actually historically useful? I have absolutely no idea.

And that’s exactly why I find these tools interesting, despite all the legitimate societal concerns and existential anxieties they have introduced into our lives.

AI, if used for writing or as a replacement for original thought, surely encourages damaging cognitive offloading. But when used to expand research questions beyond the horizon of what any single person can know, they do something else, something I for one find mind-expanding and curiosity-inducing. I think it’s worth seeing where it leads.

Weekly links

• I was honored to receive one of 80 Cosmos Institute grants announced earlier this month. I’ll be working with Nathan Davies, a PhD student at Oxford, on Humanity’s First Exam, a corpus of historical sources and questions relating to human autonomy and the relationship between humans and machines that we’ll be using to benchmark how various AI models reason about this topic. In particular we’re interested in finding the areas where they fail to encompass the breadth of the various documented human viewpoints on these issues (i.e. the topics where all AI models converge on a median answer, but humans demonstrate way more variance - I think this “epistemological flattening” is increasingly important to document as humans become increasingly reliant on asking LLMs how to think about our own history, consciousness, and experience). (Github for the prototype)

• Gotta love premodern children’s books: “We then home in on man’s lifecycle: the baby, saved from the eagle, sets out to become rich; by panel four he is a prosperous gentleman — but, of course, death comes for us all. “O MAN !” the last panel exclaims. “Now see thou art but dust…” (Public Domain Review)

Res Obscura is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

I would love to hear from people in the comments about which unsolved “historical mysteries” or other historical questions you think would be “tractable” for frontier models. Also eager to hear any results you might have gotten from doing so.

Leave a comment

As with mathematical research, seems is an important qualifier here - I tried searching around for it in the secondary scholarship but it’s entirely possible that this link has already been made in published or unpublished work I didn’t find.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论