伏尼契手稿的生成语法与拼凑假说:基于大型语言模型的证据
摘要
本文通过应用大型语言模型分析伏尼契手稿的结构和生成语法,提供了支持拼凑假说的证据。
arXiv:2609.20835v1 Announce Type: new
Abstract: Background: The Voynich Manuscript is a fifteenth-century codex written in an unknown script whose content remains undeciphered. Previous studies suggest that its statistical properties resemble those of natural languages, while its illustrations - primarily plants - recall medieval herbals.
Methods: We present a multidisciplinary analysis combining probabilistic modeling, phonetic decomposition, rare-event detection, and multimodal image analysis, based on a newly transliterated corpus. Word- and letter-level distributions are modeled using position-dependent probabilistic grammars, while phonetic patterns are compared across Indo-European, Semitic, and Asian languages. Image-text alignment methods based on large language models are applied to identify potential botanical correspondences.
Results: The results indicate that Voynich symbols behave as letters rather than syllabic units, while word-length distributions resemble syllabic structures. Phonetic analyses show closer alignment with consonant-heavy languages such as Hebrew or Arabic than with Indo-European languages. Probabilistic modeling reproduces Zipf-like distributions and reveals extremely low probabilities for repeated initial-letter sequences, indicating a structured imitation of natural language. Image analysis suggests strong correspondences between Voynich plant illustrations and those found in Pseudo-Apuleius herbals from the Mediterranean tradition, consistent with an imitation of medieval medicinal books.
Perspectives: These findings support the hypothesis that the Voynich Manuscript follows a structured generative system combining linguistic regularities and herbal knowledge, and demonstrate the value of integrating probabilistic and AI-assisted approaches in the analysis of historical manuscripts.
查看缓存全文
缓存时间: 2026/09/21 08:59
# A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models Source: [https://arxiv.org/abs/2609.20835](https://arxiv.org/abs/2609.20835) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
相似文章
大型语言模型的概率结构
该论文为大型语言模型提出了一个统一的概率框架,通过概率测度描述模型,通过最大似然估计进行训练,并将文本生成视为随机模拟过程,同时对幻觉现象和扩散模型的作用提供了深入见解。
伏尼契手稿中层叠位置与方向约束的证据:对类密码结构的启示
ArXiv 预印本量化了伏尼契手稿中多层 RTL 与 LTR 约束,显示 97% 的跨边界互信息集中于特定字形转换,且简单生成模型无法同时复现所有观测到的结构特征。
# 巴别塔的大语言模型
本文反思了文本生成的历史,在现代大语言模型(如 GPT-4)与豪尔赫·路易斯·博尔赫斯和克劳德·香农的早期概念之间建立了联系。文章探讨了香农的概率实验以及博尔赫斯“巴别图书馆”的隐喻,如何有助于阐明关于生成文本本质和数据结构的根本问题。
基于语法书的低资源机器翻译合成数据生成的析因研究
本文介绍了一个流程,该流程使用大型语言模型从语法书中提取语法规则、示例句子和词汇表,以生成合成平行语料库,用于微调机器翻译模型,在三种低资源语言上实现了高达+8.8的ChrF++增益。
大语言模型在低资源语言人文学科研究中的机遇与挑战
本文系统评估了大语言模型在低资源语言研究中的应用,分析了在语言变异、历史文献、文化表达和文学分析等方面的机遇与挑战。研究强调了跨学科合作和定制化模型开发,以保护语言和文化遗产,同时解决数据可获取性、模型适应性和文化敏感性问题。