A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models
Summary
This paper presents evidence supporting the Pastiche Hypothesis for the Voynich Manuscript by applying large language models to analyze its structure and generative grammar.
View Cached Full Text
Cached at: 09/21/26, 08:59 AM
# A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models Source: [https://arxiv.org/abs/2609.20835](https://arxiv.org/abs/2609.20835) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
The Probabilistic Structure of Large Language Models
The paper provides a unified probabilistic framework for large language models, describing them through probability measures, training via maximum-likelihood estimation, and text generation as stochastic simulation, with insights into phenomena like hallucination and the role of diffusion models.
Evidence of Layered Positional and Directional Constraints in the Voynich Manuscript: Implications for Cipher-Like Structure
ArXiv preprint quantifies layered RTL and LTR constraints in the Voynich Manuscript, showing 97 % of cross-boundary mutual information lies in specific grapheme transitions and that simple generative models cannot simultaneously reproduce all observed structural signatures.
Large Language Models of Babel
The article reflects on the history of text generation, drawing parallels between modern LLMs like GPT-4 and earlier concepts from Jorge Luis Borges and Claude Shannon. It explores how Shannon's probabilistic experiments and Borges' 'Library of Babel' metaphor help clarify fundamental questions about the nature of generated text and data structure.
A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
This paper introduces a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books to generate synthetic parallel corpora for fine-tuning machine translation models, achieving ChrF++ gains of up to +8.8 on three low-resource languages.
Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research
This paper systematically evaluates the applications of large language models in low-resource language research, analyzing opportunities and challenges across linguistic variation, historical documentation, cultural expressions, and literary analysis. The study emphasizes interdisciplinary collaboration and customized model development to preserve linguistic and cultural heritage while addressing issues of data accessibility, model adaptability, and cultural sensitivity.