Skip to content

Chapter IX: Simple Substitution — Fundamentals (2)

Text size

At (e), the first line of the cryptogram is repeated, as it would appear after the making of this exchange. The beginning of the message can almost be read: The first word appears to be _notify_, furnishing two new substitutes. Three more can be furnished in the suggested sequence _to go ahead with_. And here, the word _with_ would be tried in any case, because it is a common word, and because the frequency of the letter _P_ is suggestive of _w_. Arrived at this point, we begin to notice patterns: _postpone_, _council_, _account_, _matter_, and so on; so that the rest of the solution is largely a matter of filling in framework. In the given example, it would also be noticed that _F_ and _N_ have resulted from reciprocal encipherment; this may not be the case with other letters, but it presents a possibility which is always well-worth investigating.

Figure 68

Digram Count for the Longer Cryptograms

(First-Letters)

A B C D E F G H I J K L M N O P Q R S T U V W X Y Z
A 1 1 1 1 2 1 3 1 A 11
B B
C C
D 1 1 1 1 1 3 3 1 2 4 3 D 21
E 2 E 2
F 2 2 1 3 F 8 *
S G 1 1 1 1 1 G 5
e H 1 1 1 1 H 4 *
c I I
o J 1 2 1 1 2 2 J 9
n K 1 1 1 10 1 1 1 K 16
d L 1 1 L 2
{. M 1 1 1 1 2 M 6
L N 2 1 1 1 N 5 *
e O 1 2 1 3 1 O 8
t P 1 1 P 2
t Q 1 2 1 Q 4
e R 3 3 2 1 1 1 4 4 1 3 1 2 2 R 28 *
r S 2 1 3 4 1 1 3 2 1 S 18
s T 2 2 1 1 1 1 6 1 T 15
U 1 2 1 1 1 3 U 9
V 1 1 1 4 6 1 1 1 1 2 2 1 2 V 24
W 2 1 2 2 1 W 8
X 2 1 1 2 2 1 X 9
Y 2 6 2 Y 10
Z 1 2 1 1 2 1 1 Z 9
A B C D E F G H I J K L M N O P Q R S T U V W X Y Z
9 5 4 27

Most Frequent Digrams

RK, 10 KV, 6 WD, 4 SR, 4
VT, 6 DY, 6 RR, 4 KS, 4
HV, 4

* * *

_The Digram-Solution Method_. — This method, representing another of our many debts to M. E. Ohaver, may be used either in conjunction with the vowel method, or independently, as the fundamental method of attack. For a satisfactory demonstration, however, we need more material, and Fig. 67 shows our cryptogram again, together with a suspected reply. Thus we have a length of 235 letters, so that the preparation of contact-notations, which we found sufficient in the preceding case, becomes here an irksome task.

For these longer cryptograms, it is usually best to put all of our data into the form of a _digram-count_, as indicated in Fig. 68. This is most easily done as follows: Using a sheet of cross-section paper, mark off the limits of a 26 x 26 square; write the normal alphabet across the top, so that each of its letters will govern a column; and write it again along one side, so that each letter will govern a row. For added convenience, these two alphabets may be repeated, as they are shown in the figure. Now, remembering that each letter in the text is the first letter of a digram (except the two which are finals), our two texts, with their total of 235 letters, are to provide a count on 233 digrams. Taking letters one by one, just as they come in the cryptograms, find each letter in the upper alphabet; find, in the side alphabet, the letter which immediately follows it in the cryptogram, and count this digram by placing a tally-mark in the cell at which the column and row governed by these two letters are found to intersect. In the figure, the tally-marks have been replaced with numbers showing their totals. It will be noted that the process described is identically the method which would have been used by Meaker in preparing the digram chart; and, just as in the case of the digram chart, the counting of the digrams has automatically counted the single letters. To obtain their frequencies, we may total either the columns or the rows, taking the larger figure in those few discrepancies caused by initial or final letters. With the chart understood, the digram-method of solution can be shown in a nutshell.

An inspection of this chart enables us to find quickly that the leading digrams are those listed: _RK_, _VT_, _KV_, _DY_, etc. These, almost certainly, are the substitutes for digrams ranking high on the normal list, and many others, having a frequency of 3, are very likely indeed to be substitutes for digrams from that same high-frequency class. Our text, of course, is still short, even with 235 letters, and we do not invariably find, in texts of this length, that the ranking digram (in this case _RK_, frequency 10) is the substitute for _th_, though the chances are, at all times, that it is. And should it prove here that _RK_ does not represent _th_, then we may be quite sure that _th_ is represented in one of the digrams _VT_, _KV_, _DY_, having the next frequency, 6. With the single exception of _RR_, each digram of the nine which are listed below the chart can be checked against three other digrams: Its own reversal; the doubling of its first letter; and the doubling of its second letter. In addition, it may be checked through the individual frequencies of its two component letters. These points of comparison, made for each of the nine leading digrams, have been tabulated in Fig. 69, so that the discussion may be easily followed.

Examining _RK_, assumed to represent _th_: Its reversal, _KR_, has not appeared on the chart, which is satisfactory for a digram of no greater frequency than its supposed original, _ht_. The doubling of its first letter, _RR_, has appeared four times, which is satisfactory for _tt_, one of the leading doubles of our language. The doubling of its second letter, _KK_, has not appeared, which is eminently satisfactory for a digram as rare as _hh_. Its first letter, _R_, has a frequency of 28, the highest in the cryptogram, which is not at all unusual in the case of _t_; and its second letter, _K_, has a frequency of 16, a little high for _h_, but not unsatisfactory. Thus, we find nothing, so far, to contradict the supposition that the digram _RK_ is the substitute for _th_. But if _K_ represents _h_, it should be possible to find digrams beginning with _K_ which will check equally well as the substitutes for _ke_ and _ka_. We do, in fact, find _KV_ and _KS_. But which is which? Examination of Fig. 69 shows that one of these, _KS_, has a reversal, _SK_, frequency 1; but this is not informative, since it would be equally expected of _eh_ or _ah_. Further examination shows that _V_ has been doubled, which is far more characteristic of _e_ than of _a_. Also, the individual frequency of _V_, 24, is the second highest in the cryptogram, and more likely to be that of _e_ than that of _a_. Thus we may assume that _KV_ represents _he_ and that _KS_ represents _ha_. This automatically identifies the digram _SR_ as _at_. As to _VT_, this, apparently, involves the only reversal of any prominence in the cryptogram. Its first letter has already been identified as _e_, and the outstanding reversal of the language is _er_-_re_. This is not so certain as in the preceding cases, but the frequency of _T_ is satisfactory as that of _r_.

Thus we have identified the letters _t_, _h_, _e_, _a_, _r_, which is as far as the tabulation has been carried. Having the substitute for _h_, we may now bring in the vowel-solution method through examination of digrams _KD_, _KJ_, _KT_, _KZ_; or continue with the digram-solution method by looking over the field for some of the other _h_-digrams: _sh_, _ch_, _wh_, _ph_, _gh_, and so on. The first of these should be easily identified by the frequency of _s_, and, in addition to the regular three check-digrams, we might check this against a possible _st_, another of our leading English digrams. With the process explained, we need not go further; the substitution of letters _t_, _h_, _e_, _a_, _r_, _s_, will surely break any simple substitution cryptogram. Possibly, enough has not been said as to the use of the trigram list, the consideration of common affixes, common short words, and so on; but these are all points which the student can best develop for himself.

Figure 69

Digram Doubled Letter Letter Frequency Supposed
Identity
Original Reversed 1st 2d 1st 2d

R K 10 K R... R R 4 K K... R 28 K 16 t h
V T 6 T V 2 V V 1 T T... V 24 T 15 e r
K V 6 V K... K K... V V 1 K 16 V 24 h e
D Y 6 Y D... D D... Y Y... D 2l Y 10 ?
W D 4 D W 1 W W 1 D D... W 8 D 2l ?
S R 4 R S... S S... R R 4 S 18 R 28 a t
K S 4 S K 1 K K... S S... K 16 S 18 h a
R R 4 .... .... .... R 28 t t
H V 4 V H... H H... V V 1 H 5 V 24 ? (-e)

Another point, however, must not be overlooked: _the long repeated sequences_ _HVXXU_, _ZDYFZJX_, _DRKVT_, _GSRRVT_. Repeated sequences of these lengths will usually come from _repeated whole words_, making it possible, to some extent, to attack the cryptogram by word-division methods. It is, in fact, the repetition of sequences, these and many others, which, in the beginning, has led us to assume that both cryptograms are using the same key. As to the recovery of this key, we need not wait until solution is complete. Even in simple substitution, it is well, during the identification of substitutes, to have before us a sort of skeleton key, in which the plaintext alphabet has been written out in normal order, so that the substitutes, as fast as their identities are discovered, can be placed below their originals.

Thus, having identified as many as twelve letters in our present cryptogram, this skeleton key, or framework, might begin to assume the appearance which is indicated in the upper tabulation of Fig. 70. Here, we are able to note a reciprocal encipherment between _A_ and _S_, _F_ and _N_, _R_ and _T_, and _U_ and _Y_, suggesting that the whole encipherment may have been reciprocal; if so, we have the identities of four additional substitutes: _O_, _I_, _H_, _E_, representing _d_, _j_, _k_, _v_, respectively. If they are present in the cryptogram, these four substitutions may be tried; but with or without their presence in the cryptogram, they can be added to the skeleton key, as in the lower tabulation of the figure. Notice that when this has been done, the cipher alphabet is beginning to show alphabetical sequences (reversed). We find _H I J K_, and, just before this, _D F_, which is an alphabetical sequence if the letter _E_ has been taken out for use in a key-word. Between _DF_ and _HIJK_ of the cipher alphabet, we need only _G_ to fill out the sequence; therefore either _l_ or _m_ must belong to the key-word; comparing this with what is found at the other end of the sequence, we find that either _L_ or _M_ would be the substitute for _g_. Between _NO_, we find _V_, evidently misplaced; and, following _O_ and preceding _S_, we find two positions which may be occupied by two of the letters _PQR_, of which _R_ has already been placed (under _t_). That is, where the encipherer has used a key-word-mixed alphabet without troubling to carry it through a transposition process of any kind, we are often able to build it up again, and make it help us in the solution. This is especially true if he has used reciprocal encipherment; with the substitutes which may actually be found in our foregoing cryptograms, a little rearrangement is all that is needed in order to discover exactly what the original key was. When the cipher alphabet has been carried through a transposition block, it is not so easy to recover during the actual process of solution; afterward, however, it is not usually difficult to treat it by one of the transposition processes, just as if it were a transposition cryptogram of 26 letters. In the examples which follow, the key-word-mixed alphabets were used as they stood, though we believe that none of the encipherment was reciprocal. In one case, however, the plaintext and cipher alphabets were both mixed, according to different key-words, so that the recovery of this key may prove troublesome.

Figure 70

Supposing 12 substitutes to have been identified:

Plaintext alphabet: a b c d e f g h i j k l m n o p q r s t u v w x y z
CIPHER ALPHABET: S V N K J F D T A R Y U

Assuming reciprocal substitution:

Plaintext alphabet: a b c d e f g h i j k l m n o p q r s t u v w x y z
CIPHER ALPHABET: S O V N K J I H F D T A R Y E U
Q?P? * L? G? C?B?* * T?
M?

45. By PICCOLA.

S C Y J T O P N R M J T U E A W S R O R O A E P Q R J C R O A R M P H Q K J Q S R S J H A X P F K E A Q R M Y S R P Q P M P S E C A H G A W S R O P E E E S H A Q O P V S H I R O A Q P F A E A H I R O P H N P Q R J H T F U A M C J M R Y R O M A A W A E E B T Q R W M S R A S R J H A I M J T K U A E J W P H J R O A M P H N Q A A W O P R Y J T Q A A L.

46. By PICCOLA. (Plaintext and Cipher Alphabets have each a key-word).

J C W E H S N D F S B N J I V T E A G V D H O C Q Q I Q F R P H F K Q E A R F Q A R F A H F Q E J C B N J N H B E O C B N L N O V H B L F Q J B N A B L F V H C A J I V B N W N S T B L E A G V A J N S R F W N S Y R V S S C A E H V A Q F C J E A G J N A W N S O V B V C Q Y D C S P H E H O C S P E A G B E O N A F R L C A G N E A K C S O N S H A C B E F A C Q X.

47. By PALOMITA. (No key-word).

B O Y B A N K I L L A P K R I Y A P Y Y U P B L Y E R P B P L G Y G M H L A B O Y K J A K L P Y L H H J A C R P O R C Q U Y N B H L A B O Y G N A Z N Y L H B O Y K N A N P R B R W O J C B R C Q D N P K.

48. By PICCOLA. (Of these two, one has normal word divisions; the other has not).

W T E I C H E P P C A E P T J W P O Y D Q P R M E L U E I N D E P Q T C Q D Q D P C P D H K G E P U O P Q D Q U Q D J I C. I S Y E Q T C P V E M Y R E W M E K E C Q E S P E L U E I N E ? P D Q H U P Q C G P J T C V ! E O E E I Q M C I, P K J P E S X Q T E Q M C I P K J P D Q D J I D P U P U C G G Y J R Q T E V E M Y P D H K G E P Q F D I S.

49. By PICCOLA.

P B K L A B E I C D J D B I L Y P K L D O I X L Y I P K V Y A L ? A G F Y A M I L K L Y I K I D C A G G L D O I X V D J R K L Y I C P B R P B N X D Q A J I ? Q K J I S P B R K L Y A L A B R M X Q F P L F P E O L D I B R V Y P E Y O B D V X D Q P C E G I A J F I J C I E L G X P K S I A B P B N P L K. A J P K L D E J A L A B B D L P K L Y P K B D !

Comments

Log in to leave a comment.

Elementary cryptanalysisChapter IX: Simple Substitution — Fundamentals (2)

0%13 min left in chapter