AI can crack undeciphered ancient scripts at extraordinary speed, but meaning remains out of reach
AI is compressing decades of manual linguistic cross-referencing into minutes for undeciphered writing systems like Linear A and Etruscan, yet researchers warn that statistical pattern-matching hits a hard ceiling: a model can achieve fluency in a dead language without ever knowing what the words actually mean. A June claim that Linear A belongs to the Semitic language family remains under expert review, illustrating why AI-assisted decipherment still depends on human insight and a comparative anchor that computing power alone cannot manufacture.

AI Can Read Linear A. It Still Can't Tell You What the Words Mean.
Using AI-built programming scripts, a researcher reportedly assigned values to 40 signs and compiled a 408-word lexicon, arguing that Linear A is an extinct member of the Semitic language group, which includes Hebrew and Aramaic. The claim is still under expert review 1.
The statistical pattern-matching that let AI cross-reference thousands of ancient characters in minutes is the same paradigm powering every large language model in production today. It carries the same structural blind spot: a model can achieve fluency in a language without knowing what a single word means.
The researcher used AI scripts to test a sound-pattern across the surviving Linear A corpus, work that would have taken months by hand. The AI did not originate the hypothesis. It was the research assistant, rapidly checking a human idea against thousands of characters 1.
AI excels at large-scale pattern testing: checking whether a guess about one sign holds up across an entire archive in minutes, spotting repeated sequences a human eye would miss, and predicting likely characters to restore damaged inscriptions. The same techniques already worked on Ugaritic, an extinct Semitic language spoken during the late Bronze Age in what is now Syria, because the language family was known going in 1.
Linear A has no such anchor. Linguists call it a language isolate because it has no confirmed link to any known language, living or dead, and there is no bilingual text like the Rosetta Stone. Etruscan, a language of pre-Roman Italy, is only slightly better off: a partial vocabulary survives from short funerary inscriptions, but its grammar and deeper meaning still elude linguists 1.
Statistical pattern-matching cannot manufacture meaning from nothing. It needs an anchor to tell it which patterns are meaningful and which are coincidence, and no amount of computing power invents one from scratch 1.
Given enough data, a model could theoretically reach the point where it recognizes repeating structures well enough to repurpose Linear A phrases in new combinations, approximating something like limited conversation. But it could not hand you a translation. Fluency and meaning are not the same thing. The model learns which signs follow which, which words cluster together, without ever knowing what any of them refer to 1.
Linear A's entire surviving corpus is roughly 7,500 characters, short enough to fit on a single screen. With that little data, almost any hypothesis can find scattered matches to support it, which is why decipherment claims depend on independent expert scrutiny and peer review rather than statistical confidence scores 1.
The Linear A Semitic claim may turn out to be correct. It may turn out to be a compelling coincidence that a statistical model made look like a breakthrough. Either way, the output is polished, internally consistent, and impossible to verify against ground truth without expertise the machine does not possess. That is the same risk profile compliance teams and product leaders face whenever a model generates legal analysis, medical summaries, or technical documentation: confident, fluent output that may or may not correspond to reality, with no native speaker to check against.
AI is not a dead end for ancient-language research. It compresses years of manual cross-referencing into minutes and lets more people attempt problems than institutional resources ever allowed 1. But it does not remove the two ingredients decipherment has always required: a comparative anchor, and rigorous human review to separate a real discovery from an appealing pattern match.
The Linear A claim is now in the hands of reviewers. Watch how they reach their verdict, not how confident the model sounded. If the claim survives because independent experts find corroborating evidence in the archaeological record, that validates AI-assisted human insight. If it falls apart because the statistical matches were coincidental, that reveals the cost of confusing fluency with understanding. Either outcome will be instructive, and neither will be settled by a confidence score.
References
Cite this story
ProvenBrief (2026). "AI can crack undeciphered ancient scripts at extraordinary speed, but meaning remains out of reach." ProvenBrief. https://provenbrief.com/story/ai-can-crack-undeciphered-ancient-scripts-at-extraordinary-speed-but-meaning-rem
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.