Evidence of Layered Positional and Directional Constraints in the Voynich Manuscript: Implications for Cipher-Like Structure

arXiv cs.CL Papers

Summary

ArXiv preprint quantifies layered RTL and LTR constraints in the Voynich Manuscript, showing 97 % of cross-boundary mutual information lies in specific grapheme transitions and that simple generative models cannot simultaneously reproduce all observed structural signatures.

arXiv:2604.19762v1 Announce Type: new Abstract: The Voynich Manuscript (VMS) exhibits a script of uncertain origin whose grapheme sequences have resisted linguistic analysis. We present a systematic analysis of its grapheme sequences, revealing two complementary structural layers: a character-level right-to-left optimization in word-internal sequences and a left-to-right dependency at word boundaries, a directional dissociation not observed in any of our four comparison languages (English, French, Hebrew, Arabic). We further evaluate two classes of structured generator against a four-signature joint criterion: a parametric slot-based generator and a Cardan grille implementing Rugg's (2004) gibberish hypothesis. Across their full tested parameter spaces, neither class reproduces all four signatures simultaneously. While these results do not rule out generator classes we have not tested, they provide the first quantitative benchmarks against which any future generative or cryptanalytic model of the VMS can be evaluated, and they suggest that the VMS exhibits cipher-like structural constraints that are difficult to reproduce from simple positional or frequency-based mechanisms alone.
Original Article
View Cached Full Text

Cached at: 04/23/26, 10:02 AM

# Evidence of Layered Positional and Directional Constraints in the Voynich Manuscript: Implications for Cipher-Like Structure
Source: [https://arxiv.org/html/2604.19762](https://arxiv.org/html/2604.19762)
###### Abstract

The Voynich Manuscript \(VMS\) exhibits a script of uncertain origin whose grapheme sequences have resisted linguistic analysis\. We present a systematic analysis of its grapheme sequences, revealing two complementary structural layers: a character\-level right\-to\-left optimization in word\-internal sequences and a left\-to\-right dependency at word boundaries, a directional dissociation not observed in any of our four comparison languages \(English, French, Hebrew, Arabic\)\. Mutual information decomposition shows that 97% of cross\-boundary MI resides in specific within\-class grapheme transitions rather than in class\-level labels, and a Markov simulation demonstrates that the directional dissociation is reproducible from surface word\-level statistics without reference to positional classes or generative mechanisms\. Positional analysis identifies Zipfian boundary distributions and an 80\.6% end\-class→\\tostart\-class transition rate, both quantitatively distinct from the comparison languages\.

We further evaluate two classes of structured generator against a four\-signature joint criterion: a parametric slot\-based generator and a Cardan grille implementing Rugg’s \(2004\) gibberish hypothesis\. Across their full tested parameter spaces, neither class reproduces all four signatures simultaneously\. The slot\-based generator cannot achieve high cross\-boundary MI and Zipfian boundary distributions together; the grille cannot achieve the E→\\toS rate and Zipfian shape together\. These are structurally distinct failure modes arising from mechanistically different generators\. While these results do not rule out generator classes we have not tested, they provide the first quantitative benchmarks against which any future generative or cryptanalytic model of the VMS can be evaluated, and they suggest that the VMS exhibits cipher\-like structural constraints that are difficult to reproduce from simple positional or frequency\-based mechanisms alone\.

Keywords:Voynich Manuscript; directionality; positional structure; n\-gram perplexity; mutual information decomposition; Cardan grille; generative model evaluation

## 1Introduction

The Voynich Manuscript \(MS 408, Yale University Beinecke Library; hereafter VMS\) remains one of the most studied undeciphered texts in history\. Three competing hypotheses dominate scholarly discourse: that the VMS is meaningless gibberish\(Rugg,[2004](https://arxiv.org/html/2604.19762#bib.bib9); Timm and Schinner,[2020](https://arxiv.org/html/2604.19762#bib.bib12); Gaskell and Bowern,[2022](https://arxiv.org/html/2604.19762#bib.bib5)\), that it represents a natural or constructed language\(Bowern and Lindemann,[2021](https://arxiv.org/html/2604.19762#bib.bib2)\), or that it is a ciphertext of a known language such as Latin or Italian\(D’Imperio,[1978](https://arxiv.org/html/2604.19762#bib.bib3); Greshko,[2025](https://arxiv.org/html/2604.19762#bib.bib6)\)\.

A fundamental yet underexplored question concerns the*directionality*of the Voynich script\. While the manuscript was almost certainly*written*left\-to\-right \(LTR\), based on ink directionality and stroke analysis, whether the underlying text should be*read*LTR or right\-to\-left \(RTL\) has received little quantitative treatment\. In prior work\(Parisel,[2025](https://arxiv.org/html/2604.19762#bib.bib8)\), we introduced a language\-agnostic n\-gram perplexity asymmetry method that revealed consistent RTL optimization in the VMS character stream\.

That same work also uncovered an anomaly at word boundaries\.Ashraf and Sinha \([2018](https://arxiv.org/html/2604.19762#bib.bib1)\)showed that natural languages universally exhibit asymmetric grapheme distributions at word boundaries, a property exploitable for directionality detection using Gini and Shannon entropy measures\. However, when we applied these boundary metrics to the VMS, we found that they were unsuitable: the distribution of word\-initial and word\-final graphemes in the VMS follows a Zipfian curve, in stark contrast to the plateau\-shaped distributions observed in the comparison languages\. This distributional difference \(which we call the*Zipfian boundary effect*\) means that conventional boundary\-based directionality metrics cannot be meaningfully benchmarked against the comparison languages for the VMS\.

Motivated by these observations, we developed two new analytical methods that work*across*word boundaries rather than*at*them, circumventing the Zipfian boundary effect while probing the positional properties of VMS words\. A Markov simulation establishes that the directional dissociation is reproducible from surface word\-level statistics alone, without reference to positional classes or any generative hypothesis\. This is a clarifying result, not a deflating one: it rules out the directional dissociation as an independent diagnostic and redirects analytical attention to the positional properties that are*not*reproducible by surface resampling\. It is those properties, and their joint behaviour under generative testing, that form the main contribution of this paper\.

With the directional dissociation set aside as a surface\-level consequence of VMS word structure, we ask what structural properties of the VMS cannot be similarly explained\. The answer is a set of four positional signatures: boundary concentration, bilateral positional extremity, cross\-boundary mutual information, and Zipfian boundary distributions\. We test whether these can be jointly reproduced by two classes of structured generator, a parametric slot\-based generator and a Cardan grille generator\(Rugg,[2004](https://arxiv.org/html/2604.19762#bib.bib9)\), each tested across its full parameter space\. Neither achieves all four simultaneously; the failure modes are mechanistically distinct and robust across parameter variation, though they do not rule out generator classes we have not tested\.

We emphasise at the outset that our comparison set comprises only four languages\. This is sufficient to establish that the VMS differs from*these four languages*on the metrics reported, but is far too small to support claims about natural language in general\. Throughout this paper, statements about the VMS being “different” or “exceeding” baselines refer exclusively to the four\-language comparison set, not to natural language as a class\.

Figure[1](https://arxiv.org/html/2604.19762#S1.F1)illustrates the two\-layer directional profile\.

PREFIXsuffixPREFIXsuffixPREFIXsuffixPREFIXsuffixWord sequence: LTR\-optimizedCharacter stream: RTL\-optimizedMIMIMIFigure 1:The two\-layer directional profile of the Voynich Manuscript\. Words are composed internally of prefix\-class and suffix\-class graphemes\. The character stream is RTL\-optimized \(red arrows\), while word\-to\-word transitions are LTR\-optimized \(blue arrow\)\. Cross\-boundary mutual information \(green, dashed\) flows between suffix graphemes of one word and prefix graphemes of the next\. A word\-level Markov simulation \(Section[4\.5](https://arxiv.org/html/2604.19762#S4.SS5)\) establishes that this two\-layer pattern is a surface\-level consequence of VMS word\-internal structure and local transition probabilities, clarifying that the directional dissociation itself is not diagnostic\. The analysis therefore focuses on the positional signatures that are not reducible to surface statistics\.
## 2Background

### 2\.1Perplexity\-based directionality

InParisel \([2025](https://arxiv.org/html/2604.19762#bib.bib8)\), we defined a directional asymmetry measure

Δ=XLTR−XRTL\\Delta=X\_\{\\text\{LTR\}\}\-X\_\{\\text\{RTL\}\}\(1\)whereXXis the average cross\-entropy per token under annn\-gram model\. PositiveΔ\\Deltaindicates RTL optimization \(lower RTL perplexity\); negativeΔ\\Deltaindicates LTR optimization\. Applied to the VMS \(RF1b\-e EVA transcription\), we found consistently positiveΔ\\Deltawith tight bootstrap confidence intervals acrossn=2n=2ton=4n=4\.

### 2\.2The Zipfian boundary anomaly

Ashraf and Sinha \([2018](https://arxiv.org/html/2604.19762#bib.bib1)\)demonstrated that natural languages universally exhibit asymmetric grapheme distributions at word boundaries, with word\-final positions being more constrained than word\-initial ones\.Winstead \([2024](https://arxiv.org/html/2604.19762#bib.bib13)\)extended this finding across a large multilingual corpus, combining Gini index and Shannon entropy into a directional score\. However, these boundary metrics assume that word\-initial and word\-final grapheme frequency distributions follow the plateau\-shaped curves characteristic of natural languages\.

We showed that the VMS violates this assumption: its boundary grapheme distributions follow a Zipfian \(power\-law\) curve rather than a plateau\. This makes Gini and entropy measures unreliable for benchmarking VMS directionality against the comparison languages, as the metrics are sensitive to distributional shape, not just asymmetry\. This finding motivated the development of the cross\-boundary methods presented in this paper\.

### 2\.3Positional slot structure

Several researchers have proposed that VMS words are built from ordered positional slots\.Zattera \([2022](https://arxiv.org/html/2604.19762#bib.bib15)\)proposed a 12\-slot positional structure for Voynich words\.Greshko \([2025](https://arxiv.org/html/2604.19762#bib.bib6)\)adapted this into a two\-class system of type\-1 \(prefix\) and type\-2 \(suffix\) affixes with disjoint glyph sets as part of the Naibbe cipher model\. Our analysis does not depend on any specific slot model or cipher hypothesis; we use the concept of positional structure purely as a descriptive framework for characterising the observed distributional properties of VMS graphemes\.

## 3Methods

### 3\.1Data

We use five corpora:

- •Voynich:RF1b\-e EVA transcription\(Zandbergen,[2025](https://arxiv.org/html/2604.19762#bib.bib14)\), 37,016 words across 5,820 sentences\.
- •English \(LTR\):Melville’s*Moby Dick*\(Melville,[1851](https://arxiv.org/html/2604.19762#bib.bib7)\)\.
- •French \(LTR\):Dumas’*Le Comte de Monte Cristo*\(Dumas,[1844](https://arxiv.org/html/2604.19762#bib.bib4)\)\.
- •Hebrew \(RTL\):SVLM Hebrew Wikipedia Corpus\(SVLM,[2024](https://arxiv.org/html/2604.19762#bib.bib10)\)\.
- •Arabic \(RTL\):The Big Arabic Corpus\(The Arabic Big Corpus,[2024](https://arxiv.org/html/2604.19762#bib.bib11)\)\.

These four languages were chosen to include both LTR and RTL scripts and both European and Semitic language families\. This is a convenience sample, not a typologically representative one; it excludes agglutinative, polysynthetic, and tonal languages, among others\. All comparative statements in this paper are limited to these four languages\.

EVA graphemes are tokenized using a longest\-match greedy algorithm against the STA1 grapheme inventory\. Natural language corpora are tokenized at the character level\. For Hebrew and Arabic, which are stored in logical \(LTR memory\) order, we apply a visual transformation \(reversing word order within sentences\) before cross\-boundary analysis so that the “forward” direction corresponds to actual reading order\.

### 3\.2Cross\-boundary directionality test

For each pair of consecutive words\(wi,wi\+1\)\(w\_\{i\},w\_\{i\+1\}\)in a sentence, we extract the*boundary transition*: a pair consisting of the lastnngraphemes ofwiw\_\{i\}\(the*condition*\) and the firstnngraphemes ofwi\+1w\_\{i\+1\}\(the*target*\)\. We compute the conditional entropy

H​\(target∣condition\)=−∑c,tP​\(c,t\)​log2⁡P​\(t∣c\)H\(\\text\{target\}\\mid\\text\{condition\}\)=\-\\sum\_\{c,t\}P\(c,t\)\\log\_\{2\}P\(t\\mid c\)\(2\)and the mutual information

MI​\(condition;target\)=H​\(target\)−H​\(target∣condition\)\\text\{MI\}\(\\text\{condition\};\\text\{target\}\)=H\(\\text\{target\}\)\-H\(\\text\{target\}\\mid\\text\{condition\}\)\(3\)for both forward \(original word order\) and backward \(reversed word order\) directions\. Word\-internal grapheme order is never modified; only the sequence of words varies between conditions\.

The directional asymmetry is

ΔCB=Hfwd−Hbwd\\Delta\_\{\\text\{CB\}\}=H\_\{\\text\{fwd\}\}\-H\_\{\\text\{bwd\}\}\(4\)where positive values indicate RTL optimization and negative values indicate LTR optimization\.Weighting note\.ΔCB\\Delta\_\{\\text\{CB\}\}is computed from corpus\-wide frequency counts and is therefore implicitly weighted by the number of boundary transitions per sentence\. Confidence intervals are computed via paired bootstrap resampling over sentences \(B=500B=500,α=0\.05\\alpha=0\.05\), which resamples sentences with equal weight regardless of length\. When these disagree in sign, the n\-gram order is treated as inconclusive\. A shuffle control \(random permutation of word order within sentences\) verifies that detected signals reflect genuine sequential structure\.

### 3\.3Positional diagnostic

We classify each grapheme as*start\-preferring*,*end\-preferring*, or*ambiguous*based on a 2:1 ratio threshold of word\-initial versus word\-final occurrence counts\. The*polarization index*is the fraction of graphemes that are clearly start\- or end\-preferring\.

We then decompose cross\-boundary mutual information into two components:

MItotal=MIclass\+MIwithin\\text\{MI\}\_\{\\text\{total\}\}=\\text\{MI\}\_\{\\text\{class\}\}\+\\text\{MI\}\_\{\\text\{within\}\}\(5\)whereMIclass\\text\{MI\}\_\{\\text\{class\}\}measures information carried by the positional class labels alone, andMIwithin\\text\{MI\}\_\{\\text\{within\}\}measures the residual information carried by specific grapheme identities within their positional class\.

We apply word\-order shuffling to both components to distinguish structural \(order\-independent\) from sequential \(order\-dependent\) contributions\.

### 3\.4Markov simulation

To test whether the observed directional combination requires explanation beyond surface word\-level statistics, we train a word\-level Markov chain on the VMS and generate synthetic corpora:

1. 1\.A word\-level Markov chain of orderkk\(tested atk=1k=1andk=2k=2\) is trained on VMS word sequences, with<BOS\>and<EOS\>tokens marking sentence boundaries\.
2. 2\.Synthetic sentences are generated by sampling from the chain, with sentence lengths drawn from the real VMS length distribution\. Only words that pass EVA tokenization are retained\.
3. 3\.Both diagnostics \(character\-levelΔchar\\Delta\_\{\\text\{char\}\}and cross\-boundaryΔCB\\Delta\_\{\\text\{CB\}\}\) are computed on each synthetic corpus\.
4. 4\.The procedure is repeated for 10 independent runs at each Markov order\.

The simulation preserves VMS word\-internal structure \(because it samples real VMS words\) and approximate word\-to\-word transition probabilities, but has no knowledge of grapheme\-level positional classes or any generative model\. If the directional combination is reproducible under these conditions, it is a consequence of surface statistics rather than evidence for a specific generative mechanism\.

### 3\.5Generative model testing

To test whether the four positional signatures can be jointly reproduced by structured generators, we implement two distinct generator classes and evaluate each against all four signatures simultaneously\.

##### Four\-signature evaluation\.

For a generated corpus to “pass” the joint profile, it must simultaneously satisfy: \(Sig1\) E→\\toS rate in the range 70–95%; \(Sig2\) bilateral positional extremity present in\>\>50% of runs; \(Sig3\) cross\-boundary MI\>\>0\.10 bits; and \(Sig4\) Zipfian boundary distribution \(power\-law fitR2\>0\.85R^\{2\}\>0\.85, coefficient of variation\>\>0\.8\)\. These thresholds are calibrated to the VMS observed values and represent the minimum bar for a plausible replication\. All metrics are computed using the same pipeline as the main analysis\. Each configuration is run 20 times with different random seeds; results are reported as bootstrapped means with 95% confidence intervals\. Cohen’sddrelative to the VMS observed value measures quantitative proximity, not just threshold passage\.

##### Slot\-based generator\.

A parametric generator builds synthetic corpora from ordered positional slots, mirroring the structural hypothesis that VMS words are composed from prefix\-class and suffix\-class graphemes\. The generator constructs a lexicon of approximately 1,000 cipher words by sampling from prefix and suffix slot pools, assigns Zipfian frequency weights, and samples word sequences via a sparse Markov chain with tunable boundary\-pair preferences\. Key parameters include: bridge zone size \(graphemes shared between prefix and suffix terminal slots, controlling E→\\toS\); Markov sparsity \(top\-kksuccessors per source word, controlling MI\); and Zipf exponent \(concentration of within\-slot frequency, controlling boundary distribution shape\)\. We test 12 ablation conditions \(including single\-pool, random\-order, overlapping\-pool, and near\-miss variants\) and five sensitivity sweeps \(S1 pool overlap, S2 slot count, S3 Zipf exponent, S4 vocabulary size, S5 boundary pair strength, S6 bridge zone width, S7 Markov top\-kk\)\.

##### Cardan grille generator\.

A second generator implements the Cardan grille model proposed byRugg \([2004](https://arxiv.org/html/2604.19762#bib.bib9)\)as an alternative to linguistic hypotheses\. A grille table of dimensions\(rows×cols\)\(\\text\{rows\}\\times\\text\{cols\}\)is filled with graphemes drawn from column\-specific pools; a grille \(a binary mask ofnholesn\_\{\\text\{holes\}\}column positions\) is applied to a selected row, and the non\-blank graphemes at hole positions form one word\. We implement four word\-generation modes:random\(Rugg’s original model: each word uses an independently drawn grille position\),shift\(grille advances one column after each word\),rotate\(grille cycles through all unique rotations\), andsequential\(words are generated by reading successive rows of a table whose column distributions are learned from a real corpus rather than pre\-specified pools\)\.

We distinguish explicitly between configurations that impose pre\-separated prefix/suffix pools \(*circular on Sig1*\) and*honest*configurations where no positional pool structure is imposed\. Honest variants include: a uniform\-pool table with large row count \(establishing the column\-skew\-free baseline\), a blank\-gradient table \(testing whether density variation alone induces E→\\toS\), learned\-column\-distribution tables built from English text and random grapheme sequences \(testing whether real corpus structure transmits through the grille\), and row\-sequential traversal of a 201,723\-word English corpus at four jump probabilitiesp∈\{0\.00,0\.05,0\.10,0\.30\}p\\in\\\{0\.00,0\.05,0\.10,0\.30\\\}\(testing whether sequential scribe behaviour generates MI\)\. Five sensitivity sweeps cover blank probability, hole count, table columns, column skew, and table rows\.

## 4Results

### 4\.1Character\-level perplexity asymmetry

Table[1](https://arxiv.org/html/2604.19762#S4.T1)presents the character\-stream perplexity results from our prior work\.

Table 1:Character\-stream perplexity asymmetryΔ=XLTR−XRTL\\Delta=X\_\{\\text\{LTR\}\}\-X\_\{\\text\{RTL\}\}with Laplace smoothing and 95% bootstrap confidence intervals\. PositiveΔ\\Delta: RTL\-optimized; negativeΔ\\Delta: LTR\-optimized\.The VMS shows consistent RTL optimization \(Δ\>0\\Delta\>0\) at all n\-gram orders, while English and French show consistent LTR optimization \(Δ<0\\Delta<0\)\. This establishes that VMS word\-internal grapheme sequences are more predictable when read right\-to\-left\.

### 4\.2Cross\-boundary directionality

Table[2](https://arxiv.org/html/2604.19762#S4.T2)presents cross\-boundary results at boundarynn\-gram sizes of 1 and 2\.

Table 2:Cross\-boundary directionality: conditional entropyH​\(target∣condition\)H\(\\text\{target\}\\mid\\text\{condition\}\), mutual information, and directional asymmetry\.ΔCB=Hfwd−Hbwd\\Delta\_\{\\text\{CB\}\}=H\_\{\\text\{fwd\}\}\-H\_\{\\text\{bwd\}\}; negative values indicate LTR optimization, positive values RTL\. Where the aggregate and bootstrap CI disagree in sign, the result is marked inconclusive \(†\)\.†Voynichn=2n=2: aggregateΔCB\\Delta\_\{\\text\{CB\}\}\(token\-weighted\) is negative \(LTR\) while the bootstrap CI \(sentence\-weighted\) is entirely positive \(RTL\)\. Thisnn\-gram order is treated as inconclusive\.

The VMS shows LTR optimization at word boundaries atn=1n=1\(ΔCB=−0\.243\\Delta\_\{\\text\{CB\}\}=\-0\.243\), with forward mutual information \(0\.230\) substantially exceeding backward \(0\.051\)\. Shuffle controls yieldΔshuf≈0\.005\\Delta\_\{\\text\{shuf\}\}\\approx 0\.005, confirming the signal reflects genuine sequential structure\. All directional claims for the VMS rest onn=1n=1; then=2n=2result is inconclusive\.

### 4\.3Directional profile summary

Table[3](https://arxiv.org/html/2604.19762#S4.T3)summarises the directional profiles across corpora\.

Table 3:Directional profiles across corpora\. The*Layer Agreement Index*\(LAI\) issign​\(Δchar\)×sign​\(ΔCB\)\\text\{sign\}\(\\Delta\_\{\\text\{char\}\}\)\\times\\text\{sign\}\(\\Delta\_\{\\text\{CB\}\}\): positive indicates agreement between layers, negative indicates opposite\-direction optimization\. Within this four\-language comparison set, only the VMS shows opposite\-direction optimization; however, the Markov simulation \(Section[4\.5](https://arxiv.org/html/2604.19762#S4.SS5)\) demonstrates that this pattern is reproducible from surface word\-level statistics\.∗Hebrew shows positiveΔchar\\Delta\_\{\\text\{char\}\}atn=2n=2but negative atn=3n=3andn=4n=4, precluding a clean directional assignment\.

The VMS is the only text in this comparison set where the character\-level and word\-boundary directional signals point in opposite directions\. As shown in Section[4\.5](https://arxiv.org/html/2604.19762#S4.SS5), this combination is reproducible from surface word\-level statistics and does not constitute evidence for any specific generative mechanism\.

### 4\.4Positional diagnostic

#### 4\.4\.1Positional polarization

Table[4](https://arxiv.org/html/2604.19762#S4.T4)shows grapheme positional classification across corpora\.

Table 4:Grapheme positional classification \(2:1 ratio threshold\)\.*Polarization*: fraction of graphemes classified as clearly start\- or end\-preferring\.*End→\\toStart %*: fraction of cross\-boundary transitions where an end\-class grapheme is followed by a start\-class grapheme\.All four comparison languages show some degree of grapheme positional polarization \(0\.694–0\.860\), reflecting phonotactic and orthographic constraints at word boundaries\(Ashraf and Sinha,[2018](https://arxiv.org/html/2604.19762#bib.bib1)\)\. The VMS cross\-boundary end→\\tostart transition rate of 80\.6% is substantially higher than any of the four comparison languages \(19\.8%–35\.5%\)\. Whether this gap would narrow for languages with richer prefix\-suffix morphology \(e\.g\., Turkish, Finnish, Swahili\) or templatic structure is unknown\.

The extreme positional ratios of individual Voynich graphemes further illustrate this pattern\. The graphemeqappears 5,285 times word\-initially but only 2 times word\-finally \(ratio 1,762:1\), whileiinappears 0 times initially and 3,958 times finally\. Counting graphemes with positional ratios exceeding 100:1:

- •Voynich:at least 5 graphemes across both classes \(iinat 3,958:1,qat 1,762:1,inat 541:1,irat 473:1,chat 327:1\)\.
- •English:2 graphemes \(iat 166:1,Tat 102:1\)\.
- •French:2 graphemes \(xat 1,631:1,zat 188:1\)\.
- •Hebrew:4 graphemes, all orthographic final letter forms:*mem sofit*\(1,239:1\),*nun sofit*\(1,143:1\),*kaf sofit*\(299:1\),*pe sofit*\(212:1\)\.
- •Arabic:4 graphemes, all position\-dependent letter variants:*alif maqsura*\(2,595:1\),*ta marbuta*\(2,475:1\),*hamza*\(1,245:1\),*alif madda*\(437:1\)\.

In the four comparison languages, extreme positional ratios are concentrated in a small number of graphemes with specific orthographic explanations: final letter forms \(Hebrew*sofit*\), position\-dependent variants \(Arabic\), or rare letters \(xin French\)\. In Hebrew and Arabic, all extreme\-ratio graphemes are end\-preferring\. The VMS differs in that its extreme\-ratio graphemes span*both*positional classes \(qandchare almost exclusively word\-initial, whileiin,in, andirare almost exclusively word\-final\), with no known orthographic rule to explain the pattern\. We call this*bilateral positional extremity*\. Whether this property is absent in all natural languages or merely absent in these four cannot be determined from the present data\.

#### 4\.4\.2Mutual information decomposition

Table[5](https://arxiv.org/html/2604.19762#S4.T5)presents the MI decomposition for all corpora\.

Table 5:Mutual information decomposition at word boundaries\. MItotal\{\}\_\{\\text\{total\}\}: total cross\-boundary MI at the grapheme level\. MIclass\{\}\_\{\\text\{class\}\}: MI explained by positional class labels alone\. MIwithin\{\}\_\{\\text\{within\}\}: residual MI from specific grapheme identities within class\. Subscript “shuf” indicates values after word\-order shuffling\.Several observations emerge from the MI decomposition\. These are descriptive properties of the corpora tested; we do not claim they generalise beyond this comparison set\.

##### \(i\) Class\-level MI is negligible everywhere\.

Across all corpora, the positional class labels account for 0\.2%–4\.1% of total MI\. For the VMS, the 80\.6% end→\\tostart transition rate is a rigid pattern, but its very consistency renders it uninformative in the information\-theoretic sense: it is essentially a near\-constant that carries almost no predictive power\.

##### \(ii\) Within\-class MI dominates and depends on word order\.

In the VMS, 97% of cross\-boundary MI resides in specific grapheme\-to\-grapheme transitions within the positional classes\. This within\-class MI is largely destroyed by word\-order shuffling: VMS MI drops from 0\.223 to 0\.049, a reduction of 78%\. The pattern is qualitatively consistent across all corpora tested\.

##### \(iii\) The VMS has the highest total MI in this comparison set\.

At 0\.230 bits, VMS cross\-boundary MI exceeds all four comparison languages \(0\.054–0\.173\)\. This indicates stronger word\-to\-word sequential dependencies at the grapheme level than any of the four comparison languages, though the comparison set is too small to determine whether this is unusual for natural language in general\.

##### \(iv\) The VMS retains the most MI after shuffling\.

After word\-order shuffling, the VMS retains 21% of its original MI \(0\.049 out of 0\.230\), compared to 14% for English, 8\.5% for French, 7\.6% for Hebrew, and 9\.8% for Arabic\. This residual represents order\-independent cross\-boundary predictability, that is, baseline transition regularity that does not depend on which word follows which\.

##### \(v\) The VMS has the largest forward\-to\-backward MI ratio\.

The ratio MIfwd\{\}\_\{\\text\{fwd\}\}/MIbwd\{\}\_\{\\text\{bwd\}\}atn=1n=1is 4\.5:1 for the VMS, compared to 1\.6:1 for French, 1\.9:1 for Arabic, and 2\.3:1 for Hebrew \(Table[2](https://arxiv.org/html/2604.19762#S4.T2)\)\. The forward direction is more predictable relative to the backward direction in the VMS than in any of the four comparison languages\.

### 4\.5Markov simulation

Table[6](https://arxiv.org/html/2604.19762#S4.T6)presents the results of the Markov simulation described in Section[3\.4](https://arxiv.org/html/2604.19762#S3.SS4)\.

Table 6:Markov simulation results \(means±\\pmstandard deviations over 10 independent runs\)\.Δchar\\Delta\_\{\\text\{char\}\}: character\-stream perplexity asymmetry atn=2n=2\(positive = RTL\)\.ΔCB\\Delta\_\{\\text\{CB\}\}: cross\-boundary asymmetry atn=1n=1\(negative = LTR\)\. Dissociation: number of runs whereΔchar\>0\\Delta\_\{\\text\{char\}\}\>0andΔCB<0\\Delta\_\{\\text\{CB\}\}<0\.ConfigurationMeanΔchar\\Delta\_\{\\text\{char\}\}MeanΔCB\\Delta\_\{\\text\{CB\}\}DissociationOrder\-1 Markov\+0\.0057±0\.0004\+0\.0057\\pm 0\.0004−0\.2601±0\.0070\-0\.2601\\pm 0\.007010/10Order\-2 Markov\+0\.0048±0\.0006\+0\.0048\\pm 0\.0006−0\.2646±0\.0071\-0\.2646\\pm 0\.007110/10*Real VMS**\+0\.0099\+0\.0099**−0\.2431\-0\.2431**—*The Markov chain reproduces the opposite\-direction combination \(positiveΔchar\\Delta\_\{\\text\{char\}\}, negativeΔCB\\Delta\_\{\\text\{CB\}\}\) in10 out of 10 runsat both order 1 and order 2\.

##### Why the character\-level RTL signal survives\.

The Markov chain samples real VMS words, each of which carries its internal grapheme structure intact\. The RTL signal is a*word\-internal*property: any process that reuses VMS words will reproduce it, regardless of whether it has knowledge of positional classes or cipher tables\.

##### Why the cross\-boundary LTR signal survives\.

Even an order\-1 Markov chain captures sufficient word\-to\-word transition structure to produce forward predictability at word boundaries\.

##### Consequence\.

The opposite\-direction combination isnot diagnostic of any specific generative mechanism\. It is a consequence of two properties already present in the VMS: \(1\) word\-internal grapheme sequences that are more predictable right\-to\-left, and \(2\) word\-to\-word transition probabilities that are more predictable in forward order\. Any process that preserves these surface statistics will reproduce the combination\.

This result does not bear on the other properties reported in this paper \(positional polarization, boundary concentration, Zipfian distributions, MI decomposition\), which characterise VMS word structure itself\.

### 4\.6Generative model testing

#### 4\.6\.1Joint profile criterion

We evaluate each generator against the four positional signatures jointly, requiring simultaneous satisfaction of all four thresholds\. Table[7](https://arxiv.org/html/2604.19762#S4.T7)summarises the results\. No generator configuration achieves 4/4 across either class\.

Table 7:Interpretation matrix for generative model testing\. Each row is one generator configuration\. Sig1: E→\\toS rate 70–95%; Sig2: bilateral extremity in\>\>50% of runs; Sig3: MI\>\>0\.10 bits; Sig4: Zipfian boundary shape\. ✓ = VMS\-like,∼\\sim= marginal,×\\times= absent\. Joint = number of signatures passed simultaneously\.ConfigurationSig1 \(E→\\toS\)Sig2 \(Bilat\)Sig3 \(MI\)Sig4 \(Zipf\)JointSlot\-based generatorBaseline \(bridge=1,kk=5,α\\alpha=1\.2\)✓✓✓×\\times3/4Best near\-miss \(bridge=1,kk=5,α\\alpha=2\.0\)✓✓×\\times✓3/4Single pool \(ablation D\)×\\times×\\times∼\\sim×\\times0/4Random word order \(ablation C\)×\\times✓×\\times∼\\sim1/4Agglutinative mimic×\\times✓×\\times✓2/4Templatic mimic×\\times✓∼\\sim✓2/4Cardan grille generator \(honest configurations\)Uniform pool \+ random×\\times×\\times∼\\sim✓1/4Blank\-gradient \+ random×\\times×\\times∼\\sim✓1/4Learned\-English columns \+ random×\\times×\\times∼\\sim×\\times0/4Learned\-random columns \+ random×\\times×\\times∼\\sim×\\times0/4Row\-sequential English \(pp=0\.00\)×\\times×\\times∼\\sim×\\times0/4Row\-sequential English \(pp=0\.05\)×\\times×\\times∼\\sim×\\times0/4Cardan grille generator \(circular configurations, for comparison\)SPLIT \+ RANDOM \(Rugg core\)×\\times✓×\\times∼\\sim1/4SPLIT \+ SHIFT×\\times✓✓✓3/4SPLIT \+ ROTATE×\\times✓✓✓3/4Naibbe cipherNaibbe over Latin source×\\times✓×\\times×\\times1/4VMS \(observed\)✓✓✓✓4/4
#### 4\.6\.2Slot\-based generator: the Sig3/Sig4 tension

Across 12 ablation conditions and seven sensitivity sweeps \(totalling over 2,000 individual runs\), the best slot\-based configuration achieves 3/4\. The failing signature alternates predictably with parameters: configurations optimised for MI \(Sig3\) produce Intermediate rather than Zipfian boundary distributions \(Sig4 fails\), while configurations that achieve Zipfian shape do so through high Zipf exponent, which simultaneously collapses MI\. This trade\-off is mechanistic: high Zipf exponent concentrates frequency mass on a few graphemes, producing the power\-law boundary distributions characteristic of Sig4, but the same concentration reduces the effective diversity of cross\-boundary transitions, destroying the sequential dependencies that generate Sig3\. The slot\-based generator cannot simultaneously satisfy Sigs 3 and 4 regardless of bridge zone width, vocabulary size, Markov sparsity, or boundary pair strength\.

#### 4\.6\.3Cardan grille: the Sig1/Sig4 tension

Across 15 honest configurations and five sensitivity sweeps, no grille variant achieves all four signatures\. The sensitivity sweep over column skew \(Table[8](https://arxiv.org/html/2604.19762#S4.T8)\) reveals the governing trade\-off directly\.

Table 8:Cardan grille sensitivity to column skew \(α\\alpha\): E→\\toS rate and boundary distribution shape as a function of the Zipfian concentration parameter applied to column grapheme distributions\. SPLIT specialisation,randommode, 20 runs of 37,000 words each\. Higher skew produces Zipfian boundary distributions \(Sig4\) but destroys E→\\toS \(Sig1\)\. No skew value satisfies both simultaneously\.Atα=0\\alpha=0, E→\\toS reaches 73% but shape is Intermediate; atα≥1\.5\\alpha\\geq 1\.5, shape is Zipfian but E→\\toS collapses below 14%\. The mechanism is the same as in the slot\-based generator but operating on a different parameter: high column skew concentrates each column’s distribution onto one dominant grapheme, producing Zipfian boundary frequencies \(Sig4\) through a pure frequency\-concentration effect\. But that same dominant grapheme now appears at roughly equal rates at word starts and ends \(because it dominates both the first and last columns of any word generated from this table\), collapsing the positional asymmetry that drives E→\\toS \(Sig1\)\. No intermediate skew value achieves both\.

The honest row\-sequential configurations confirm that sequential traversal of a structured source does not rescue MI\. With a 201,723\-word English corpus as source and pure sequential traversal \(pjump=0\.00p\_\{\\text\{jump\}\}=0\.00\), the net sequential signal \(MIorig\{\}\_\{\\text\{orig\}\}−\-MIshuf\{\}\_\{\\text\{shuf\}\}\) is approximately 0\.002 bits, two orders of magnitude below the VMS target of 0\.181 bits\. The random\-source null yields an identical net signal \(0\.002 bits\), confirming that with a corpus large enough to avoid row repetition, sequential traversal contributes no measurable MI regardless of source structure\.

#### 4\.6\.4Complementary failure modes

The two generator classes fail on complementary signatures\. The slot\-based generator achieves E→\\toS \(Sig1\) and bilateral extremity \(Sig2\) by construction from its pool structure, but cannot achieve Zipfian shape \(Sig4\) without destroying MI \(Sig3\)\. The Cardan grille, in its honest form, achieves neither E→\\toS \(Sig1\) nor Zipfian shape \(Sig4\) simultaneously; in its circular form \(SPLIT \+ SHIFT or ROTATE\), it achieves Sig3 and Sig4 but destroys Sig1\. The VMS simultaneously achieves all four\. This complementary failure pattern, two mechanistically distinct generators failing on the same joint profile from different directions, provides stronger evidence for the distinctiveness of the VMS profile than either generator class alone\.

#### 4\.6\.5Naibbe cipher

We additionally test the Naibbe cipher proposed by Greshko as a cipher\-based generator\. Naibbe encrypts Latin or Italian plaintext by respacing it into unigram \(single\-letter\) and bigram \(two\-letter\) tokens, then substituting each token with a glyph string drawn from a card\-weighted lookup table with disjoint prefix and suffix pools for bigram tokens\. The cipher was designed to produce VMS\-like output and is evaluated here as the most structurally motivated cipher hypothesis currently in the literature\.

We implement the full encoder faithfully, including the UNAMBIGUOUS mode that rejects bigram combinations whose glyph strings are ambiguous under re\-tokenization\. Two source texts are tested: Pliny the Elder’s*Naturalis Historia*\(Books 16–17\), the same Latin corpus used by Greshko; and Melville’s*Moby Dick*as an English comparison\. Twenty independent runs of approximately 37,000 output tokens each are evaluated through the same four\-signature pipeline\.

Results are consistent across both sources\. Over Pliny Latin: E→\\toS = 48\.7% \[48\.7, 48\.8\], MI = 0\.040 \[0\.040, 0\.041\], bilateral extremity in 100% of runs, shape Intermediate\. Over Moby Dick English: E→\\toS = 52\.6% \[52\.5, 52\.6\], MI = 0\.053 \[0\.052, 0\.053\], bilateral extremity in 100% of runs, shape Intermediate\. Joint score: 1/4 for both sources, passing only Sig2\. The Latin values are marginally lower on all three failing metrics, consistent with Latin’s somewhat flatter letter\-frequency distribution producing slightly weaker positional concentration at the EVA tokenization level\. The result is source\-independent\.

Both failures are structurally explained by two properties of the cipher\. First, the∼\\sim47% unigram tokens \(drawn from a pool that overlaps with both prefix and suffix boundary graphemes at the EVA tokenization level\) systematically dilute the E→\\toS rate; the theoretical ceiling is well below the VMS’s 80\.6%\. Second, each card draw is statistically independent, so no sequential dependency exists between the end of one cipher word and the start of the next, producing near\-zero cross\-boundary MI regardless of source language\. Naibbe scores 1/4, the same as Rugg’s original Cardan grille \(G0\), and for partially overlapping structural reasons\.

This finding is not a criticism of the Naibbe cipher hypothesis, which was not designed to reproduce these specific metrics\. It is a constraint: any cipher\-based model of the VMS must account for the joint profile, and the Naibbe architecture as currently specified does not satisfy it\.

## 5Discussion

### 5\.1The Markov result as analytical pivot

The VMS exhibits opposite\-direction optimization: RTL at the character\-stream level and LTR at the word\-boundary level\. This combination is not observed in any of the four comparison languages\. The Markov simulation establishes that this dissociation is a surface\-level consequence of two properties already present in the VMS: word\-internal grapheme sequences that are more predictable right\-to\-left, and word\-to\-word transition probabilities that are more predictable in the forward direction\. Any process that preserves these surface statistics will reproduce the combination\.

This result has a clarifying, not a deflating, role in the paper’s argument\. By showing that the directional dissociation reduces to surface statistics, it rules it out as an independent diagnostic and concentrates analytical attention on the positional properties that do not reduce in the same way\. In natural languages, both directional levels emerge from the same underlying phonotactic and morphological system, so they tend to agree on directionality\. The VMS dissociation reflects the fact that its word\-internal and word\-sequencing structures carry different directional fingerprints—a property of VMS word structure that is itself informative, even if it is not uniquely diagnostic\.

### 5\.2What the positional properties describe

The positional properties that are not reproducible by Markov resampling, and that differ quantitatively from all four comparison languages, are:

##### Boundary concentration\.

The 80\.6% end→\\tostart transition rate exceeds the four comparison languages by a factor of 2\.3–4\.1×\\times\. This means VMS word boundaries enforce a rigid alternation between positional grapheme classes that is quantitatively different from these four languages\. The gap could narrow for morphologically richer languages not yet tested\.

##### Bilateral positional extremity\.

Extreme positional ratios \(\>\>100:1\) span both start\-preferring and end\-preferring grapheme classes, unlike the four comparison languages where such ratios cluster in one direction with specific orthographic explanations\. This bilateral pattern is consistent with a system where word\-initial and word\-final positions draw from largely non\-overlapping grapheme pools, but we cannot exclude the possibility that some untested natural languages exhibit similar properties\.

##### Zipfian boundary distributions\.

VMS word\-boundary grapheme distributions follow a Zipfian curve rather than the plateau shape observed in the four comparison languages\(Parisel,[2025](https://arxiv.org/html/2604.19762#bib.bib8)\)\. This is a fundamental property of VMS word structure\.

##### High cross\-boundary MI with high structural residual\.

The VMS has the highest total cross\-boundary MI \(0\.230 bits\) and the highest MI retention after word\-order shuffling \(21%\) among the five corpora tested\. The high retention indicates that a substantial fraction of cross\-boundary predictability is structural \(order\-independent\)\.

These properties are observations, not explanations\. They describe quantitative differences between the VMS and four specific natural languages\. They constrain the space of plausible text\-generation models in the sense that any proposed model must reproduce them, but they do not by themselves identify the generative process or exclude any hypothesis\.

### 5\.3Constraints and non\-constraints on VMS hypotheses

We can state clearly what these findings do and do not constrain:

##### What they constrain\.

Any proposed text\-generation model for the VMS must reproduce: \(a\) the 80\.6% end→\\tostart boundary concentration; \(b\) bilateral positional extremity across the grapheme inventory; \(c\) Zipfian word\-boundary distributions; \(d\) the MI decomposition profile \(negligible class\-level MI, high within\-class MI destroyed by shuffling\); and \(e\) word\-internal grapheme sequences that are more predictable right\-to\-left\. The generative model testing reported here extends this list with an empirical constraint: \(f\) two specific generator classes, a parametric slot\-based generator and a Cardan grille, each fail to achieve all four signatures simultaneously across their full tested parameter spaces\.

##### What they do not constrain\.

These findings do not discriminate between the natural language, gibberish, and ciphertext hypotheses\. The positional properties are observations that could in principle arise from a natural language with extreme positional morphology, from a structured generative process we have not yet tested, or from a cipher with positional slot structure\. The results narrow the space of plausible generators by identifying a structural incompatibility within each tested class, but do not close it: they demonstrate that generators in which frequency concentration is the sole mechanism for both positional and distributional signatures cannot satisfy the joint profile\. Generators where E→\\toS and Zipfian shape arise from genuinely independent mechanisms remain untested\.

### 5\.4Limitations

1. 1\.Small comparison set\.Four languages is insufficient to characterise natural language in general\. The absence of agglutinative languages \(Turkish, Finnish\), polysynthetic languages, tonal languages, and scripts with complex templatic morphology is a critical gap\. All comparative claims are limited to these four languages and should not be generalised\.
2. 2\.Generator class coverage\.The slot\-based generator, Cardan grille, and Naibbe cipher cover three important structural hypotheses but do not exhaust the space of possible generators\. Generators with independently tunable positional pool structure and frequency concentration, where E→\\toS and Zipfian shape are controlled by separate mechanisms rather than both depending on the same frequency distribution, may achieve the joint profile\. Such generators would, however, require explicit architectural separation of these properties, which itself constitutes a structural claim about the VMS that can be evaluated\.
3. 3\.Circularity in grille SPLIT configurations\.Grille configurations with pre\-separated prefix/suffix column pools are partially circular with respect to Sig1, since the pool separation is imposed rather than emergent\. Honest configurations \(uniform pool, blank\-gradient, learned\-column\) avoid this but score lower\. We report both and label them clearly; interpretations should rest on honest configurations where possible\.
4. 4\.Transcription dependence\.The cross\-boundary test depends on the accuracy of word segmentation in the EVA transcription\. Ambiguous word boundaries in the VMS could affect results\.
5. 5\.Coarse positional classification\.The MI decomposition uses a three\-class partition \(start/end/ambiguous\) derived from a 2:1 ratio threshold\. A finer\-grained analysis using the full slot positions ofZattera \([2022](https://arxiv.org/html/2604.19762#bib.bib15)\)could reveal additional structure\.
6. 6\.Single transcription system\.All results are based on the EVA transcription\. Different transliteration systems could yield different grapheme inventories and therefore different positional statistics\.

## 6Conclusion

We have characterised the Voynich Manuscript’s directional and positional properties at two levels, benchmarked against four natural languages, and tested whether the resulting four\-signature joint profile can be reproduced by two classes of structured generator\. The VMS exhibits the following quantitative properties that differ from all four comparison languages:

1. 1\.Opposite\-direction optimization \(surface\-level\)\.The VMS character stream is RTL\-optimized while its word\-boundary transitions are LTR\-optimized\. None of the four comparison languages shows this combination\. A word\-level Markov simulation reproduces this pattern in 10/10 runs, establishing that it is a surface\-level consequence of VMS word structure and local transition statistics\. This result clarifies the paper’s analytical focus: the directional dissociation is real but reducible, and the properties that are*not*reducible to surface statistics are items 2–6 below\.
2. 2\.High boundary concentration\.VMS word boundaries enforce an 80\.6% end\-class→\\tostart\-class transition rate, compared to 19\.8%–35\.5% in the four comparison languages\.
3. 3\.Bilateral positional extremity\.Extreme positional ratios \(\>\>100:1\) span both start\-preferring and end\-preferring grapheme classes, unlike the comparison languages where such ratios cluster in one direction with specific orthographic explanations\.
4. 4\.Zipfian boundary distributions\.Word\-boundary grapheme frequencies follow a power\-law curve rather than the plateau shape observed in the comparison languages\.
5. 5\.High structural MI at boundaries\.The VMS has the highest total cross\-boundary MI \(0\.230 bits\) and the highest MI retention after word\-order shuffling \(21%\) among the five corpora tested\.
6. 6\.Negligible class\-level MI\.Despite the rigid 80\.6% boundary alternation, positional class labels account for only 3% of cross\-boundary MI; the remaining 97% resides in specific grapheme\-to\-grapheme transitions that depend on word order\.

The generative model testing adds a new result: across the full parameter spaces of a slot\-based generator \(12 ablations, 7 sensitivity sweeps,\>\>2,000 runs\), a Cardan grille generator \(15 configurations including honest non\-circular variants, 5 sensitivity sweeps\), and the Naibbe cipher \(20 runs each over Pliny’s*Naturalis Historia*in Latin and Melville’s*Moby Dick*in English\), no configuration achieves all four signatures simultaneously\. The slot\-based generator and Cardan grille fail on complementary signatures through mechanistically distinct processes: in the slot generator, cross\-boundary MI and Zipfian boundary distributions require opposite frequency\-concentration conditions; in the grille, E→\\toS rate and Zipfian shape require opposite column\-skew values\. The Naibbe cipher, despite being designed to produce VMS\-like output, scores 1/4, failing on E→\\toS, MI, and Zipfian shape, for structural reasons intrinsic to its card\-independent encryption mechanism\. The VMS holds all four simultaneously; no tested generator achieves this combination\.

These results are descriptive, not explanatory\. They constrain what any proposed generative model must reproduce, and they demonstrate that the simplest instances of two prominent generative hypotheses, structured slot\-based composition and Cardan grille, do not meet this constraint across their tested parameter spaces\. Whether more complex variants of these generators, typologically diverse natural languages, or different generative approaches can reproduce the joint profile is an open question and the necessary next step\.

## References

- Ashraf and Sinha \(2018\)Ashraf, M\. I\. and Sinha, S\. \(2018\)\.The handedness of language: Directional symmetry breaking of sign usage in words\.*PLoS ONE*, 13\(1\):e0190735\.
- Bowern and Lindemann \(2021\)Bowern, C\. L\. and Lindemann, L\. \(2021\)\.The Linguistics of the Voynich Manuscript\.*Annual Review of Linguistics*, 7\(1\):285–308\.
- D’Imperio \(1978\)D’Imperio, M\. \(1978\)\.*The Voynich Manuscript: An Elegant Enigma*\.Fort Meade: National Security Agency\.
- Dumas \(1844\)Dumas, A\. \(1844\)\.*Le Comte de Monte Cristo*\.Project Gutenberg,[https://www\.gutenberg\.org/ebooks/17989](https://www.gutenberg.org/ebooks/17989)\.
- Gaskell and Bowern \(2022\)Gaskell, D\. E\. and Bowern, C\. L\. \(2022\)\.Gibberish after all? Voynichese is statistically similar to human\-produced samples of meaningless text\.In*Proc\. International Conference on the Voynich Manuscript 2022*, University of Malta\.
- Greshko \(2025\)Greshko, M\. A\. \(2025\)\.The Naibbe cipher: a substitution cipher that encrypts Latin and Italian as Voynich Manuscript\-like ciphertext\.*Cryptologia*\.[https://doi\.org/10\.1080/01611194\.2025\.2566408](https://doi.org/10.1080/01611194.2025.2566408)\.
- Melville \(1851\)Melville, H\. \(1851\)\.*Moby Dick*\.Project Gutenberg,[https://www\.gutenberg\.org/ebooks/2701](https://www.gutenberg.org/ebooks/2701)\.
- Parisel \(2025\)Parisel, C\. \(2025\)\.Directionality of the Voynich Script\.arXiv:2509\.10573v4\.
- Rugg \(2004\)Rugg, G\. \(2004\)\.An Elegant Hoax? A possible solution to the Voynich Manuscript\.*Cryptologia*, 28\(1\):31–46\.
- SVLM \(2024\)SVLM Hebrew Wikipedia Corpus \(2024\)\.[https://github\.com/NLPH/SVLM\-Hebrew\-Wikipedia\-Corpus/](https://github.com/NLPH/SVLM-Hebrew-Wikipedia-Corpus/)\.
- The Arabic Big Corpus \(2024\)The Arabic Big Corpus \(2024\)\.[https://github\.com/mohataher/arabic\_big\_corpus/](https://github.com/mohataher/arabic_big_corpus/)\.
- Timm and Schinner \(2020\)Timm, T\. and Schinner, A\. \(2020\)\.A possible generating algorithm of the Voynich manuscript\.*Cryptologia*, 44\(1\):1–19\.
- Winstead \(2024\)Winstead, J\. \(2024\)\.Writing Direction Detection\.[https://jhnwnstd\.github\.io/projects/writing\-direction/](https://jhnwnstd.github.io/projects/writing-direction/)\.
- Zandbergen \(2025\)Zandbergen, R\. \(2025\)\.Text Analysis – Transliteration of the Text\.*The Voynich Manuscript*\.[https://www\.voynich\.nu/transcr\.html](https://www.voynich.nu/transcr.html)\.
- Zattera \(2022\)Zattera, M\. \(2022\)\.A new transliteration alphabet brings new evidence of word structure and multiple ‘languages’ in the Voynich Manuscript\.In*Proc\. International Conference on the Voynich Manuscript 2022*, University of Malta\.

Similar Articles

How far are we from solving The Voynich Manuscript?

Reddit r/ArtificialInteligence

The article explores the potential of using latest AI models to attempt deciphering the Voynich Manuscript, a centuries-old unsolved code, amid recent announcements of other historical puzzles being solved.

Leveraging Morphology for Historical Script Metrological Analysis

Hugging Face Daily Papers

This paper presents a transformer-based architecture with prototype learning that enables scalable paleographic measurements from historical documents using only line-level transcriptions, demonstrating effectiveness on a 160-page codex with minimal training data.