Shakespeare’s plays occupy eight of the top ten places in the all-word bigram comparison with Neville’s writing, although they make up only 36 of the 236 plays tested. A bigram is a two-word sequence, such as “I will” or “write to”. Counting thousands of these combinations tests whether Neville’s habitual language resembles Shakespeare’s plays more closely than other contemporary drama. The three methods below compare all-word bigrams, pairs formed after filtering for frequent words, and samples of similar length.
What was compared
The dramatic texts come from Pervez Rizvi’s database of 527 early modern plays. The comparison here is limited to 236 dramatic texts dated 1590–1615: the 36 plays of the First Folio and 200 texts by other writers. Pericles, Edward III and The Two Noble Kinsmen are excluded from both groups. Collaborative First Folio plays, including Henry VIII, count once as complete plays in the group statistics. The Shakespeare section of Henry VIII is also shown separately in the tables, using the database’s division between Shakespeare and Fletcher.1
On Neville’s side, the study used 165 English letter records and three longer documents, comprising 137,551 words after processing.1 The question is how each play resembles Neville’s collected writing, not which play resembles a single selected letter.
The comparison works with base forms of words, so that forms such as writes and writing can be counted under write. It therefore examines word use rather than the appearance of particular historical spellings. A pair is never made by joining the last word of one Neville document to the first word of another.
What a two-word comparison measures
A bigram is a sequence of two words. In the illustrative phrase “I will write to you,” the consecutive pairs are “I will,” “will write,” “write to” and “to you.” A passage contributes many such overlapping pairs. Counting them records not just which words a writer uses, but which words repeatedly occur together.
Each text receives a frequency profile: some pairs occur often, others occasionally, and many not at all. The comparison asks how closely a play’s balance of those frequencies resembles Neville’s. Cosine similarity compares the relative pattern of pair frequencies, rather than raw totals. This removes the direct effect of multiplying every count by the same amount, but longer texts can still differ because they contain more varied and more stable samples of language. Its score runs from zero for no shared pairs towards one for increasingly similar profiles; it is not a percentage of identical wording.
This differs from selecting striking verbal parallels. A rare expression can be illuminating when read in context, but these tests also count ordinary combinations. A high score can arise from many small similarities in usage without any one phrase being remarkable. The three methods vary the vocabulary and the amount of text compared, allowing us to see which results depend on those choices.
How to read the percentages
Each play first receives a similarity score. The group comparison then pairs every Shakespeare play with every play by another writer and asks which has the higher score. With 36 Shakespeare plays and 200 others, there are 7,200 pairings. In the all-word test, the Shakespeare play scores higher in 5,894, or 81.9 per cent, of them.
Thus 81.9 per cent describes how often a Shakespeare play outranks another play in its similarity to Neville. It does not mean that 81.9 per cent of their words match, or that there is an 81.9 per cent probability of common authorship. A result of 50 per cent would mean that the two groups win equally often. The graph shows that Shakespeare’s advantage remains substantial under all three methods, although it becomes smaller when the comparison is restricted to samples of common words.

Each table contains the twenty highest-ranked complete plays plus Henry VIII (Shakespeare section), with complete plays numbered 1–20 and the supplementary section left unnumbered at its position by similarity score. Bold titles identify First Folio plays and the supplementary Shakespeare section. The section is not counted as an additional play in the percentages, group means or random-group tests.
1. All-word bigrams
The first method uses all eligible word pairs, without limiting the vocabulary to common words. It can therefore reflect both habitual phrasing and subject matter: correspondence about government, for example, may share vocabulary with a history play. It asks which plays most closely resemble the overall pattern of Neville’s paired words.
The Winter’s Tale ranks first among complete plays, Henry VIII second and Cymbeline third. Shakespeare plays occupy twelve of the top twenty complete-play places, although they make up only about 15.3 per cent of the 236-text comparison. The 81.9 per cent pairing result shows that the advantage extends beyond the plays at the very top.
| Rank | Play |
|---|---|
| 1 | The Winter’s Tale |
| 2 | Henry VIII (full play) |
| 3 | Cymbeline |
| 4 | All’s Well That Ends Well |
| 5 | Cynthia’s Revels |
| 6 | Henry V |
| 7 | Henry IV, Part 2 |
| Henry VIII (Shakespeare section) | |
| 8 | Measure for Measure |
| 9 | The Royal King and the Loyal Subject |
| 10 | Coriolanus |
| 11 | Every Man Out of His Humour |
| 12 | The Tragedy of Charles Duke of Byron |
| 13 | Love’s Labor’s Lost |
| 14 | Hamlet |
| 15 | Henry IV, Part 1 |
| 16 | King Lear |
| 17 | 1 Edward the Fourth |
| 18 | 2 Edward the Fourth |
| 19 | Volpone |
| 20 | Every Man in His Humour |
2. Frequent-word bigrams
The second method concentrates on the 200 most frequent base words in the dramatic collection’s plays dated 1580–1620. The list is selected from the plays, rather than from expressions chosen because they favour Neville. It includes small grammatical words such as and, of and to, but also common words with substantial meaning, including king, love and honour.
Words outside that list are removed before the pairs are formed. This matters: two words counted as a pair may originally have had other words between them. The method tests sequences in the remaining common-word pattern, not necessarily phrases that a reader would find printed side by side. It reduces the contribution of unusual vocabulary while retaining much of the ordinary language through which writers construct their sentences.
For a simple illustration, take the invented sentence “I humbly ask you.” The study’s 200-word list includes I and you, but neither humbly nor ask:
| Stage | Words or pairs |
|---|---|
| Invented example | I humbly ask you |
| Ordinary adjacent pairs | I humbly; humbly ask; ask you |
| After the 200-word filter | I you (humbly and ask removed) |
| Pair now counted | I + you |
The filtered pair “I you” records the sequence left after removal; it does not claim that those words stood together in the original sentence. The same operation is applied to Neville and the plays.
Henry VIII ranks first among complete plays and Henry V second. Shakespeare plays again occupy twelve of the top twenty complete-play places and score higher in 78.1 per cent of the 7,200 pairings. Their advantage therefore survives when much of the less frequent vocabulary is removed. The result does not depend entirely on unusual shared expressions.
| Rank | Play |
|---|---|
| 1 | Henry VIII (full play) |
| Henry VIII (Shakespeare section) | |
| 2 | Henry V |
| 3 | Cymbeline |
| 4 | The Winter’s Tale |
| 5 | The Tragedy of Charles Duke of Byron |
| 6 | Henry IV, Part 2 |
| 7 | Cynthia’s Revels |
| 8 | The Conspiracy of Charles Duke of Byron |
| 9 | Coriolanus |
| 10 | The Family of Love |
| 11 | Henry VI, Part 2 |
| 12 | Sejanus His Fall |
| 13 | Henry VI, Part 1 |
| 14 | Macbeth |
| 15 | Love’s Labor’s Lost |
| 16 | The Hector of Germany, or The Palsgrave, Prime Elector |
| 17 | King John |
| 18 | Richard II |
| 19 | Philotas |
| 20 | The Widow’s Tears |
3. Comparing passages of similar length
The third method uses the same common-word list but compares passages rather than treating each full text as one sample. A long collection of letters offers a more extensive picture of usage than a short play. Sampling reduces that imbalance and asks whether the affinity remains when the amount of language being compared is more nearly alike.
The study takes fifty samples of approximately 10,000 retained words from Neville’s pooled writing and uses their average pattern as the comparison profile. For each sufficiently long play, it takes fifty passages of the same size and averages their similarity scores. Plays with between 5,000 and 10,000 retained words contribute their whole text once; shorter plays are excluded. The sample size refers to words remaining after the vocabulary filter, not words on the original page. Neville’s samples contain slightly fewer than 10,000 actual words because document-boundary markers occupy some sampled positions. The passages can overlap; they are repeated samples of the same works, not new independent texts.
This leaves 221 plays: 36 First Folio plays and 185 others. Henry VIII ranks first among complete plays, The Winter’s Tale second and Henry V fifth. The separately scored Shakespeare section of Henry VIII comes above the full play in this method. Shakespeare takes eleven of the top twenty complete-play places and scores higher in 73.9 per cent of the 6,660 cross-group pairings. Sampling reduces the advantage but does not remove it.
| Rank | Play |
|---|---|
| Henry VIII (Shakespeare section) | |
| 1 | Henry VIII (full play) |
| 2 | The Winter’s Tale |
| 3 | The Tragedy of Charles Duke of Byron |
| 4 | The Conspiracy of Charles Duke of Byron |
| 5 | Henry V |
| 6 | Macbeth |
| 7 | The Hector of Germany, or The Palsgrave, Prime Elector |
| 8 | Cymbeline |
| 9 | Philotas |
| 10 | The Family of Love |
| 11 | Richard II |
| 12 | All’s Well That Ends Well |
| 13 | Coriolanus |
| 14 | The Atheist’s Tragedy, or The Honest Man’s Revenge |
| 15 | Henry IV, Part 2 |
| 16 | King John |
| 17 | Henry VI, Part 1 |
| 18 | Sejanus His Fall |
| 19 | The Rape of Lucrece |
| 20 | 1 The Troublesome Reign of King John |
What the pattern shows
The second graph displays every eligible complete-text score. Filled dots represent the First Folio group and open dots the comparison group; a square identifies Henry VIII (full play) and a diamond its Shakespeare section. The section is supplementary and is excluded from the group means. Shakespeare scores lie further to the right on average in every panel, but the groups overlap. Chapman’s two Byron plays, for example, rank prominently in the common-word comparisons.

A further test asks how unusual the Shakespeare group’s high average position is. For each method, the computer selected 20,000 random groups of 36 plays from that comparison’s pool. None had an average rank as high as the actual Shakespeare group. Allowing for the finite number of random draws gives a p value of about 0.00005; an adjustment for conducting three tests gives about 0.00015, or 0.015 per cent.1
These figures make the concentration statistically significant under this unadjusted rank-concentration test. “Unadjusted” here means that date, length and other textual differences are not controlled; the adjustment for three tests is a separate matter. Chapter 13 reports a different test that adjusts for date and length. They answer a specific question: is the Shakespeare group more highly ranked than groups formed by randomly choosing plays? They do not measure the probability that Neville wrote Shakespeare. Nor does this test hold composition date, play length and dramatic genre constant, or allow for the fact that several plays share an author. Those features can influence language and help explain why a group resembles Neville.
The three methods also share the same source material and much of the same vocabulary. Their agreement is useful evidence that the affinity survives different ways of making the comparison, but it is not three independent discoveries whose probabilities can be multiplied. The finding is a sustained resemblance in patterns of two-word usage. Establishing how distinctive that resemblance is to Neville requires comparisons with other writers’ prose as well as with other dramatists’ plays.
The Confession, Case and Advice: a closer look at Hamlet
A separate comparison asks what happens when the focus narrows from Neville’s collected writing to particular documents. It tests the Confession alone, then the Confession together with the Case, and finally those two defences together with the later Advice. Here the comparison pool contains 331 plays dated 1590–1630. Its wider date range means its ranks should be read within that pool, rather than as positions in the 236-play tables above.2
Using the frequencies of all two-word sequences, Hamlet ranks first in all three comparisons. Adding the Case and Advice therefore preserves Hamlet’s leading position while broadening the Neville material under examination. Across these three document combinations, Shakespeare plays win about 77, 80 and 80 per cent of pairings with other plays. With all three documents combined, ten Shakespeare plays appear in the leading twenty.
This result depends on the kind of resemblance measured. When the test counts only distinct pairs found in no more than five plays, Hamlet no longer ranks first. Its strength here is in the overall frequency pattern, including ordinary language, rather than an exceptional total of rare verbal matches. The specific passages examined in the Hamlet chapter provide the context that a numerical ranking alone cannot supply.
Comparisons with other letter-writers help define the limit of the result. In tests using equal-sized samples of correspondence, John Chamberlain’s letters also showed substantial affinity with Shakespeare: their typical cross-group pairing result was about 78.5 per cent, against 77.0 per cent for Neville’s letter samples. Neville performed better on a separate measure emphasizing Shakespeare’s concentration towards the top of the ranking. Neither result makes general Shakespeare affinity unique to Neville; these letter samples also do not constitute a matched test of the Confession, Case and Advice individually.2
The chapter’s positive finding is therefore substantial but specific. Shakespeare plays are concentrated near the top when compared with Neville’s writing by the frequency of two-word sequences; several leading plays recur across the three methods, and Hamlet leads the focused document comparisons. These are suggestive patterns of linguistic affinity, to be considered alongside the documentary and verbal evidence, rather than definitive authorship identifications.
Notes and sources
Rerun of 3 October 2026: 165 English letter records and three longer documents (the Confession, Case and Advice), totalling 168 analytical records and 137,551 alphabetic lemma tokens. One duplicate witness of the 4 June 1606 letter was removed; enclosed and quoted material within English-labelled records remains. French records were excluded. The earlier descriptive census in Chapter 11 uses different counting rules. The principal comparison uses 36 First Folio plays and 200 other dramatic texts dated 1590–1615; Pericles, Edward III and The Two Noble Kinsmen are excluded from scoring and from the 1580–1620 pool used to select the 200-word vocabulary. The twenty-one-entry tables comprise the top twenty complete plays plus Henry VIII’s Shakespeare section; only complete texts enter group statistics. Database divisions identify 9,994 Shakespeare and 14,718 Fletcher tokens in Henry VIII; after frequent-word filtering these become 6,833 and 10,401. The section scores concatenate their assigned passages, following the original implementation. Similarity is cosine similarity of lowercased lemma-pair counts, with no pairs crossing Neville document boundaries. The windowed method averages fifty 10,000-position samples for longer texts; texts with 5,000–10,000 retained words are scored whole. Neville’s profile averages fifty such samples, with boundary markers occupying some positions. Thus the section and full play are not strictly equal-length samples. A separate boundary-aware check of all possible windows at equal lengths of 3,000, 5,000 and 6,500 tokens places the Shakespeare section above the combined play and the Fletcher section under all three profiles. These overlapping windows are not independent observations. Each random-group test draws 20,000 groups of 36 complete texts: zero equalled or bettered the observed mean rank, giving add-one p = 1/20,001 and Holm-adjusted p = .000150 across three tests. The test assumes exchangeable labels and does not adjust for date, length, genre or shared authorship. Text sources and the database are documented by Pervez Rizvi, “Collocations and N-grams”. These analyses are this book’s research, not Rizvi’s results; study data and scripts have not yet been publicly deposited. The database’s dates and text records are used as supplied, without independently redating the plays. It includes masques, entertainments and additions as well as full-length plays.↩︎
The focused document comparison uses sentence- and document-bounded base-word pairs in 331 plays dated 1590–1630, with 37 in its Shakespeare group. Full bigram-frequency cosine puts Hamlet first for Confession, Confession + Case, and Confession + Case + Advice; their cross-group winning fractions are .768340, .796286 and .802078. The last combination places ten Shakespeare plays in the top twenty. Counting distinct pairs occurring in at most five candidate plays gives Hamlet tied positions 20–37, 9–19 and 35–46 respectively. Equal-size correspondence-sample controls give median bigram winning fractions of .770 for Neville and .785 for Chamberlain; average precision, a measure emphasizing the concentration of Shakespeare plays near the top, is .310 and .275 respectively. These are controls for letter samples, not matched controls for each document combination. The play texts come from Rizvi’s collection, cited in note 1. The focused pool requires at least 4,000 dialogue words and an assigned date within 1590–1630. This separate earlier study includes Pericles in its 37-play Shakespeare group and retains its original comparison pool; it has not been rerun under the principal study’s new First Folio-only selection. Group statistics were recalculated from the preserved per-play scores on 28 September 2026 to correct its earlier classification; the scores and Hamlet’s first-place rankings are unchanged. The comparison correspondents were processed through a shared lemma map, whereas the primary Neville samples retained their stored lemmas; normalization was therefore not identical across these controls. Its effect on small differences between correspondents remains to be tested under consistent processing.↩︎