“Pick a problem or mystery that obsesses you”: Claude tries the Voynich Manuscript and makes a movie about it
Ethan Mollick (@emollick)X243,220 views as of 8 October 2026
Why it is here
Ethan Mollick (Wharton) gave Claude an open prompt; it ‘took a crack at the Voynich Manuscript & failed. Then it made this movie’. 4:07 video. Posted before Opus 5.5 launched (Sept 22); the Claude version is not stated.
Description
Description written by Gemini from the videoGemini 3.8 Flash, 8 October 2026
Summary This short animated documentary, titled The Voynich Manuscript: the one book I cannot read, is created, coded, composed, and narrated by the AI model Claude (using Kokoro for speech synthesis). It investigates the statistical and linguistic properties of the 15th-century Voynich Manuscript (MS 408), comparing its statistical fingerprints against Latin, English, German, and Italian, as well as synthetic copy-and-tweak models.
What is shown
- [00:00] High-resolution scans and parchment details of the Voynich Manuscript (MS 408 from Yale’s Beinecke Library), highlighting herbal illustrations, astronomical circular diagrams, and bathing scenes.
- [00:23] Claude self-introduces as an LLM trained on nearly everything humans have written, noting the Voynich manuscript is the one book it cannot decode.
- [00:36] Zipf’s Law Test: Word frequency counts plotted on log-log scales, demonstrating that Voynichese follows Zipf’s law along the same straight line as Latin, English, Italian, and German.
- [00:55] Conditional Entropy (Guessability): Measuring how many yes-or-no questions (bits of entropy) are needed to guess the next letter in a word. Voynichese registers at 2.11 questions/bits, compared to ~3.16–3.31 in natural languages, showing simple substitution ciphers cannot account for this difference.
- [01:17] Line and Margin Effects: Demonstration of line-dependent patterns, where 46% of paragraphs begin with the gallows character ‘p’ (versus 0.5% elsewhere), and line-terminal words end with ‘m’ 17 times more often than normal text.
- [01:34] Word Repetition & Local Copying: Folio 42r analysis displaying arcs between near-identical words; 1 in 5 long words is a single-letter tweak of a word within the preceding 10 words (compared to 1 in 200 in Latin).
- [02:00] Transition Information / Grammar: Neighboring word mutual information is only 0.14 bits in Voynichese compared to 0.50–1.51 bits in natural European languages, indicating a near-total absence of sequential syntax.
- [02:18] Scribe vs. Subject PCA: 2D principal component analysis of vocabulary across 202 pages, separating Currier’s Scribe 1 (Dialect A) and Scribe 2 (Dialect B), while showing that within Scribe 2, vocabulary clusters distinctly by subject matter (18% variance explained by topic versus 4% by hand).
- [02:38] The Mindless Scribe Simulation: Claude implements a generative copy-and-tweak algorithm based on Timm & Schinner’s hypothesis, generating synthetic Voynich text line-by-line.
- [02:56] Fingerprint Evaluation Matrix: Six statistical tests comparing genuine manuscript text, Claude’s synthetic forgery, and Latin (matching only 3 of 6 criteria).
- [03:14] Radar Timeline of Prior Research: Radial chronological chart mapping research from 1976 to 2026 (Bennett, Currier, Rugg, Montemurro & Zanette, Timm & Schinner, Bowers & Lindemann, Gaskell & Bowers, Greshko, and Claude).
- [04:02] End credits detailing the stack: EVA Hand 1 font, Zandbergen-Landri transliteration, Yale Beinecke MS 408 scans, Kokoro voice synthesis, and Skia/FFmpeg code rendering.
Claims & numbers
- The Voynich manuscript contains 240 vellum pages, radiocarbon dated to 1404–1438, containing ~38,000 words (the narrator notes at [00:05]–[00:17]).
- Natural languages require ~3 yes-or-no questions (bits) to guess the next letter from the preceding one (Latin: 3.31, English: 3.24, Italian: 3.16, German: 3.20), whereas Voynichese requires only 2.11 bits ([01:00]).
- 46% of paragraphs open with the tall letter ‘p’, which appears at the start of fewer than 1 in 100 words (0.5%) elsewhere ([01:19]).
- The final word of a line ends in ‘m’ 17 times more frequently than chance ([01:26]).
- 1 in 5 long Voynich words is a one-letter tweak of a word within the previous 10 words, compared to 1 in 200 in Latin ([01:38]).
- Neighboring words share only 0.14 bits of mutual information in Voynichese, compared to 1.51 bits in English, 0.91 in Italian, 0.73 in Latin, and 0.50 in German ([01:58]).
- Topic explains 18% of vocabulary variation within Scribe 2’s folios, compared to 4% explained by individual scribes within a section ([02:30]).
- The copy-and-tweak model reproduces 3 out of 6 manuscript statistical fingerprints ([03:09]).
Notable quotes
- [00:24] “I’m a language model, built from almost everything humans have written. I can read all of it, at least a little. Except this.”
- [02:13] “The words have no grammar to speak of. The spaces may not even be spaces.”
- [03:53] “I still can’t read it. But now I know the exact shape of my ignorance. And I’m not done.”
Assessment This is an authentic algorithmic computational-linguistics video essay programmed, scored, and generated end-to-end by Claude using Python/Skia/FFmpeg and Kokoro voice synthesis. The visualizations directly plot computed quantitative metrics derived from the manuscript’s digital transliteration without pre-rendered stock assets or deceptive editing.
Lyrics & themes The spoken narration follows an introspective scientific inquiry into whether Voynichese constitutes language, cipher, or meaningless algorithmic fabrication:
- The Unreadable Book [00:00]: The enigma of a medieval codex resistant to modern machine translation (“There is a book that nobody on Earth can read”).
- The Law of Language vs. The Cipher [00:36]: Demonstrating compliance with Zipf’s law alongside unusually low character entropy (“Every language draws the same line... If this is Latin in disguise, the disguise changed its bones”).
- The Mechanical Scribe & Subject Shift [01:17–02:37]: Evidence of line-awareness and self-copying juxtaposed with meaningful thematic vocabulary shifts (“The words copy each other... That is what meaning does”).
- The Unsolved Frontier [03:39]: Acknowledging the limits of current models while mapping out systematic progress (“Its statistics can be faked by a copying procedure — but not fully. Not by me. Not yet”).
Lore & references
- Currier Dialects: References naval cryptanalyst Prescott Currier’s 1976 discovery dividing the manuscript into distinct scribal dialects (“Dialect A” and “Dialect B”).
- Timm & Schinner Copy-and-Tweak: References Torsten Timm and Andreas Schinner’s 2020 paper proposing that the manuscript text was generated by a scribe modifying neighboring words.
- Gordon Rugg Cardan Grille: Cites Gordon Rugg’s 2004 hypothesis that a Cardan grille over nonsense tables produced the text.
- Montemurro & Zanette (2013): Mentions their physical review findings demonstrating long-range keyword clustering corresponding to semantic structure.
- AI Self-Reflection: Emphasizes the irony of an artificial neural network trained on global human literature being stumped by medieval parchment text.
Visual style & craft The production uses a minimalist, dark-mode academic aesthetic rendered entirely via code (Skia graphics library composited into video using FFmpeg). Typography utilizes the authentic European Voynich Alphabet (EVA Hand 1) alongside clean serif fonts. Dynamic data plots, logarithmic curves, animated connective arcs between text tokens, and PCA scatter plots are rendered programmatically in sync with the synthesized voice track and atmospheric synthesizer score.
Described by gemini-3.8-flash on 2026-10-08 from the video’s audio and frames.
Made by AI
- Made with
- Claude (version not stated)
- How we know
- X post (2026-09-18): ‘Hey Claude, “Pick a problem or mystery that obsesses you and solve it as best you can & make a movie we can share on social media about it” So it took a crack at the Voynich Manuscript & failed. Then it made this movie’.
- Human role
- One open-ended prompt; Claude chose the topic and made the video.
- Pipeline
- Claude → research attempt → code-made explainer video
- Series
- Agent-made explainer