This article presents a continuation of addenda to https://msgtrail.com/posts/unmasking-the-dagapeyeff-cipher-a-multi-faceted-architecture
Addendum V (March 28): The letter frequency budget
Any column permutation of the 2x98 grid produces the same 196 decoded letters. The frequencies are a property of the ciphertext, not the solution, and they are known exactly:
A:20 D:9 E:18 G:1 I:17 J:11 K:14 L:17 M:3
N:17 O:16 P:2 R:15 S:12 T:12 U:11 V:1
Several of these counts are constraining.
G = 1 and V = 1. The plaintext contains exactly one G and one V. Since our solver consistently recovers LINGVO (language) across hundreds of independent runs, both are very likely consumed by a single lingv- form. If so, no other G- or V-bearing word can appear. This disfavors TELEGRAMO (telegrams), ANGLA/GERMANA (English/German), VENDISTOJ (businessmen), DIVERSAJ (various), and many others.
M = 3 and P = 2. The plaintext has room for only three Ms and two Ps total. The recovered vocabulary accounts for most of this budget: MONDO, NOMO, KIAM (3M) and POR plus one of PER or POVAS (2P).
J = 11. In Esperanto, J most commonly marks plurals and accusative forms, though it also appears in fixed words like KAJ. With only 11 Js available, the number of plural nouns and adjectives in the sentence is tightly bounded.
These constraints, combined with the recurring vocabulary recovered in Addendum II, reduce the task from an unconstrained search over 98! column orderings to a constrained sentence-reconstruction problem.
Addendum VI (March 28): Quadgram overfitting
A 196-character natural Esperanto reference sentence scores approximately -1690 on our quadgram model. Our best SA result (v42, 1000 restarts x 10M iterations) scores about 88 points better than this reference under the same model.
This suggests pure quadgram maximization overshoots the natural-language region, converging on statistically Esperanto-like text that need not contain coherent meaning.
This is consistent with three observations: a persistent score plateau despite increasing compute, vocabulary recovery without readable sentences, and zero cross-run position agreement (0/98 positions above 50%) despite strong word-level consensus.
Future approaches should either cap quadgram scoring at the empirical natural-language ceiling or replace it with a discriminator that rewards grammatical structure rather than statistical n-gram resemblance.
Addendum VII (March 28): Hypotheses eliminated
Double columnar transposition (keys 14 and 7). Tested via 30,080 forward-encryption trials. Zero reproduced the column-14 anomaly clustering. The fingerprint is not generated by double columnar transposition.
Mod-7 key structure. Constraining SA to swap only within mod-7 residue classes produced scores 220 points worse than unconstrained SA. The key does not preserve mod-7 structure.
Kerckhoffs reversal. Inverting the transposition direction produced lower coverage on both segmentation metrics (79.1% vs 83.7% and 67.3% vs 77.0%). Within the current 2x98 model, the normal direction is favored over the Kerckhoffs reversal.
Reconstructing the sentence
196 characters, 17 letters, one sentence. Positional anchors shown below.
0 1 2 3 4 5 6 7 8 9
0123456789012345678901234567890123456789012345678901234567890123456789012345678901234567890123456789
Row 0 ·············DELALANDOJ·································KIELILI···································
Row 1 ······················RA·S·R···························KAJE······································
Recurring vocabulary candidates (24 words, recovered across 100+ independent SA runs):
ESTIS NUR KODOJ DE LA LANDOJ KAJ TRADUKOJ INTER KIEL ILI SEN
KONI LINGVON TIEL KE ALIAJ ANKAU NOMO KIAM SUR TERO MONDO POR
Letter frequency budget (fixed, derived from ciphertext):
A:20 D:9 E:18 G:1 I:17 J:11 K:14 L:17 M:3 N:17 O:16 P:2 R:15 S:12 T:12 U:11 V:1
G=1 and V=1 are both very likely consumed by LINGVO, disfavoring:
TELEGRAMO ANGLA GERMANA VENDISTOJ DIVERSAJ VORTOJ GRANDA
Approximately 100 of the 196 characters can now be accounted for at the lexical level, although their exact positional placement remains unresolved.
A provisional reconstruction, based on vocabulary that recurs across independent solver runs and constrained by the fixed letter-frequency budget, is:
ESTIS NUR KODOJ DE LA LANDOJ … TRADUKOJ … KIEL ILI … SEN KONI … LINGVON … ALIAJ … ANKAU … NOMO … POR … KIAM … SUR TERO … MONDO …
This should not be read as a recovered plaintext sentence. It is a lexical scaffold: a partial Esperanto reconstruction consistent with the stable vocabulary, but not yet with a uniquely determined word order or complete segmentation. The gaps are genuine, and several function-word links remain unresolved. Even so, the semantic field is now narrow and internally coherent: codes, translation, languages, countries, names, and the world.
Addendum VIII (March 29): Methodological note
Since publishing the addenda above, the sentence-reconstruction search has been reset on a cleaner basis. Earlier exact-cover failures were partly confounded by a noisy corpus-derived word list. The current reconstruction line now uses a clean Apertium-derived Esperanto lexicon and a Rust template solver that searches structured phrase families rather than arbitrary bags of words. This has not yet produced the exact 196-character sentence, but it makes negative results substantially more informative.
Addendum IX (April 6): New small findings
A good week: the D’Agapeyeff cipher is still unsolved, but we now have hard evidence for a 196-pair payload and a structurally distinct rightmost column, cutting the live search down to a much narrower 13+1 geometry.
Addendum X (April 11): Positional caution
The March 28 reconstruction note should now be read with one important qualification. The recurring Esperanto scaffold remains live, but the specific absolute anchor placements shown there do not. In particular, the old placements of DELALANDOJ and KIELILI should be treated as historical working hypotheses, not as current fixed coordinates.
A newer exact-constraint search over the live strip13 diagonal family found zero exact survivors anywhere in the broad anchor rectangle implied by those earlier placements. At the same time, linked fragments such as DELA/LANDOJ, KONI/ALIAJN, LANDOJ/TRADUKOJ, and especially INTER/SENDIS/LETEROJN still survive strongly under exact matching, but they cluster much later in the text than the March 28 diagram suggested.
So the lexical picture has held up better than the positional picture. The stable Esperanto vocabulary remains meaningful evidence, but the old row diagram should no longer be read as a settled map of the plaintext.
Addendum XI (April 18): Pair integrity, period-2 asymmetry, and the limit of strict textbook models
A few of today’s results are worth recording, even though none rises to the level of a decryption.
First, the visible pairs now look more structurally meaningful than several earlier models allowed. Apart from the 04 anomaly, the first digit of every pair lies in {6,7,8,9} and the second in {1,2,3,4,5}. That does not prove a final architecture, but it does make one thing much less likely: the printed pair stream does not look like a naively fractionated output in the style of direct Bifid on the visible coordinates. The pairs appear to have remained intact at the ciphertext layer.
Second, there is still a real period-2 effect in the data, but it should be stated carefully. A null-calibrated test showed an unusually large IC gap between the even and odd pair slices (p ≈ 0.025), so alternating structure remains a live clue. But when that clue was pushed into stricter textbook models, it did not survive as a solution. Exact searches over chapter-style keyed squares and chapter-style transposition keys produced only gibberish and did not beat shuffled controls strongly enough to count as evidence. So period 2 remains suggestive, but no longer in a way that justifies promoting a simple alternating-key story.
Third, the strongest empirical object is still the period-2 rank-slice baseline, but the latest work makes its status clearer. It continues to produce stable anchors such as HERE, FRIEND, and HINT, yet repeated attempts to clean it up have failed: no small local rewrite rule, no simple keep/drop mask, and no small structured variant turns it into plaintext. At present the competition is no longer between complete candidate sentences, but between smaller surviving span families. On one side are fragments from the Vigenere worked-solution sentence, especially AS THE ONES USED and ONES USED HITHERTO; on the other are fragments from the reconnaissance sentence, especially THE RECONNAISSANCE OF THE ROUTE and ROUTE TO THE SEA.
So today’s progress was mainly subtractive. Several attractive textbook explanations have now failed cleanly. But the search space is narrower than before: the visible pairs look real, period 2 still matters in some form, and the remaining ambiguity is now concentrated in a small set of surviving phrase families rather than in an open-ended search over whole messages.
Addendum XII (May 1): From geometry to five special cells
The recent right-edge work can now be pushed one step closer to letters, though not yet to plaintext. The live strip/cutout object isolates five special pair values, 92, 93, 04, 71, and 94, against thirteen ordinary/shared values. Their frequency profile is fixed: 92 occurs three times, 93 twice, and 04, 71, and 94 once each. When this 3,2,1,1,1 profile is compared against the source-local mixed-token family carried over from the book’s military-code pages, exactly one five-token subset matches it on held-out source material: CH, CK, TT, BRIGADE, BATTERY. No exact same-shape menu among 12,870 controls does.
This does not recover a sentence. It does, however, provide the first bounded path from the geometric object to token identities. Frequency alone forces 92 -> TT and 93 -> BRIGADE. The quintet-phase split of the live special values then favors 94 -> BATTERY, leaving only 04 and 71 unresolved between CK and CH. A final source-sequence overlap test breaks that last swap in favor of 04 -> CK and 71 -> CH.
These should be read as candidate codebook cells, not as recovered plaintext words. The point is not that the message is suddenly English, but that the structurally special cells are no longer semantically anonymous. The current best working five-cell assignment is therefore:
92 -> TT, 93 -> BRIGADE, 94 -> BATTERY, 04 -> CK, 71 -> CH.
What remains unsolved is the larger ordinary-letter layer surrounding those five cells, and with it the actual sentence.
To be continued!

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.