Henry NevilleReading edition · 0.94.335
Book PDF
Chapter 13

Testing the Affinity

A study of 180 language features selected from Neville’s correspondence found a consistent advantage for the Shakespeare group across 236 plays. A Shakespeare-group play outscored a comparison play in about seven out of ten pairings. After allowing for date and length, the estimated difference was 0.44 standard deviations. It remained positive when each play was removed in turn, and when each feature was removed in turn: no single play or individually scored feature accounted for the result.1

The study extends the comparison beyond parallels noticed by a reader familiar with Shakespeare. It selected the language from correspondence before scoring the plays. The result is evidence of linguistic affinity; the central question is how much that affinity reflects authorship, shared professional language or dramatic subject matter. The principal statistical test fell short of the preset significance threshold, as explained below.

Choosing the language first

The source comparison used 168 English records from the current Neville corpus and 1,412 letters from the Parsed Corpus of Early English Correspondence, second edition. Comparing correspondence with correspondence helped distinguish Neville’s usage from ordinary letter-writing convention: a feature had to stand out against other correspondents before it could enter the test.2

The features included single words, consecutive phrases, and short sequences allowing intervening words. They were divided among diplomatic and epistolary language, legal and institutional language, and other lexical patterns. Each division contributed equally, as did the five feature classes within it. This prevented one abundant kind of expression from dominating the score. The selection excluded Neville’s name and known transcription artifacts.

The list includes ordinary words and patterns, not just striking phrases. Actual selected features include the single word honor, the consecutive pair your honor, the consecutive phrase the duke of, and the ordered sequence and … so … your … honor. In the last example, one to three words in total may intervene while the order stays the same. These are normalized word forms used by the program, not quotations of Neville’s spelling. Related forms such as honour and honor are treated alike.3

A phrase can therefore be common in English and still qualify if Neville uses it relatively often compared with the control letters. The question is whether the whole selected set is distributed differently across the plays. Its interest lies in that collective pattern, not in treating every occurrence of king or your honor as an individual authorship clue.

The selection program used correspondence without loading drama or play labels. It fixed a list of 180 features and their weights before the testing program applied them to the plays; the testing program checked that the list had not changed. The updated run follows this same separation of selection and scoring. It is nevertheless a rerun of an existing study, not a fresh blind experiment: the researcher already knew the earlier play results.3

Features are selected and weighted using correspondence, then applied to the plays. Updated analysis, 3 October 2026; see notes 2–3.
Features are selected and weighted using correspondence, then applied to the plays. Updated analysis, 3 October 2026; see notes 2–3.

For each play, a feature counted as present or absent. Repeating it fifty times brought no more credit than using it once. A score of 0.40 means that the play contains 40 per cent of the available weighted feature credit. It does not mean that 40 per cent of its language resembles Neville, or that Neville has a 40 per cent probability of being its author.

What the plays showed

The comparison contains the same 36 First Folio plays and 200 other plays as Chapter 12. Pericles, Edward III and The Two Noble Kinsmen are excluded from both groups. Their average scores were 0.394 and 0.330 respectively. Shakespeare’s advantage was a difference between groups, not a clean division between individual plays: Henry VIII had the highest complete-play score, followed closely by Chapman’s The Tragedy of Charles Duke of Byron.4

Each dot is one complete play; vertical strokes mark group means. A square identifies Henry VIII (full play); a diamond shows its Shakespeare part separately, excluded from group means. Scores measure weighted feature presence, not authorship probability. The distributions overlap substantially. Updated analysis; see note 4.
Each dot is one complete play; vertical strokes mark group means. A square identifies Henry VIII (full play); a diamond shows its Shakespeare part separately, excluded from group means. Scores measure weighted feature presence, not authorship probability. The distributions overlap substantially. Updated analysis; see note 4.

The twenty highest-scoring plays

The twenty highest-scoring complete plays are numbered 1–20 below, before adjustment for date and length. Henry VIII (Shakespeare part) appears unnumbered at its score position. Bold titles identify First Folio plays and this supplementary part; the part is excluded from group statistics and model fitting, so Henry VIII is counted only once.

Rank Play Score
1 Henry VIII (full play) 0.5935
2 The Tragedy of Charles Duke of Byron 0.5865
Henry VIII (Shakespeare part) 0.5630
3 Henry V 0.5532
4 All’s Well That Ends Well 0.5405
5 The Hector of Germany, or The Palsgrave, Prime Elector 0.5386
6 When You See Me You Know Me 0.5378
7 Henry VI, Part 2 0.5267
8 Richard III 0.5071
9 The Conspiracy of Charles Duke of Byron 0.5033
10 2 Edward the Fourth 0.5011
11 The Famous History of Sir Thomas Wyatt 0.4895
12 King Lear 0.4879
13 Hamlet 0.4862
14 The Dumb Knight 0.4824
15 1 Sir John Oldcastle 0.4806
16 Henry IV, Part 2 0.4749
17 A King and No King 0.4748
18 The Massacre at Paris 0.4747
19 The Revenge of Bussy D’Ambois 0.4726
20 A Knack to Know a Knave 0.4726

Original analysis, updated 3 October 2026; see note 4.

Eight of the top twenty plays belong to the Shakespeare group: 40 per cent of the leading twenty, against about 15 per cent of the full sample. Shakespeare is therefore substantially more prominent at the top of the ranking than his share of the sample would suggest. These are the unadjusted rankings; the model below accounts for date and length.

The main model allowed for a curved date trend and work length. The Shakespeare advantage persisted after those adjustments: 0.0327 score units, equivalent to 0.44 standard deviations. Its survival across the deletion checks shows that the pattern extends beyond any one unusually close play or feature.

The strongest standardized difference appeared in diplomatic and epistolary language: 0.45 standard deviations, compared with 0.30 for legal and institutional language and 0.29 for the remaining lexical patterns. Sequences allowing intervening words also produced larger standardized differences than exact strings of three to five words. The strongest observed affinity thus lay in the language of correspondence and in flexible word sequences. These subsidiary comparisons did not meet the .05 threshold after adjustment for multiple testing.5

Diplomatic and epistolary language has the largest standardized difference. The q values adjust for three category tests; none is below .05. Updated analysis; see note 5.
Diplomatic and epistolary language has the largest standardized difference. The q values adjust for three category tests; none is below .05. Updated analysis; see note 5.

The principal test compared the observed difference with 20,000 constrained reshufflings preserving date and length structure. It gave p = .1169, above the preset .05 cutoff. Under that procedure, about twelve per cent of the shuffled results matched or exceeded the observed difference. This is not the probability that the authorship hypothesis is true or false. A separate bootstrap interval was positive. It was calculated by repeatedly resampling plays within the two author groups, whereas the permutation test used reshufflings constrained by date and length. These different procedures need not give identical conclusions; the protocol assigned the primary decision to the permutation test.6

How much is diplomatic language?

An additional comparison tested whether the signature recognized the language of government outside drama. It paired 719 EarlyPrint texts identified through state, diplomatic, treaty, or legal metadata with texts of similar date and length. The government-related group scored higher, especially in legal and institutional language.7

That result matters directly to the Neville case. The language linking his correspondence with Shakespeare belongs in part to a documentary world he inhabited. It also leaves a live alternative explanation: dramatic treatment of diplomacy, government, and elite relationships may draw on that shared language without identifying its author. The calibration measures sensitivity to register; it does not determine how much of the Shakespeare-group difference register explains.

Two further checks preserved the direction of the result. Excluding Henry VIII, Timon of Athens and Titus Andronicus reduced the standardized difference to 0.40 and gave p = 0.1867. Pericles was already excluded from the main comparison. Restricting the correspondence control to 1580–1620 gave an adjusted difference of 0.0360 and p = 0.0766. These are sensitivity checks, not independent replications. The main model also lacked genre and author-group adjustments, which would help assess the influence of dramatic subject matter and multiple works by the same comparison writer.8

Dots are adjusted score differences; bars are 95% work-bootstrap intervals. The p values come from separate permutation tests. All exceed the .05 threshold. These procedures answer related questions differently; the intervals do not override the primary test. Updated analysis; see notes 6 and 8.
Dots are adjusted score differences; bars are 95% work-bootstrap intervals. The p values come from separate permutation tests. All exceed the .05 threshold. These procedures answer related questions differently; the intervals do not override the primary test. Updated analysis; see notes 6 and 8.

The earlier hand-picked parallels produced a larger effect, but Shakespeare had helped determine their selection. The value of the 180-feature study is that its program selected language from correspondence before scoring the plays. The earlier parallels remain examples to assess individually; their larger estimate cannot replace this test.

The study establishes an observed pattern across a broad comparison set: language selected for its prominence in Neville’s correspondence occurs more strongly in the Shakespeare group, and the adjusted advantage survives every reported single-play and single-feature deletion. That is a substantive result, not merely another list of attractive quotations. It supports taking the Neville–Shakespeare linguistic relationship seriously alongside the historical evidence. The primary test did not reach the preset significance threshold, and this design cannot decide whether authorship, professional language or dramatic subject matter explains the affinity.

Notes and sources


  1. Original analysis for this book, rerun 3 October 2026 using corpus version 26, the same deduplicated XML as Chapter 12, and the established selection and testing procedure. The adjusted coefficient is 0.0326594; divided by the residual standard deviation from the date-and-length model it is 0.443865. The unadjusted cross-group winning fraction is 0.690972. These measure different aspects of the result. This rerun changes the source corpus and restricts the play selection; it is not independent replication.↩︎

  2. Parsed Corpus of Early English Correspondence, second edition (2022), compiled by the CEEC Project Team; annotated by Ann Taylor, Arja Nurmi, Anthony Warner, Susan Pintzuk, and Terttu Nevalainen; revised, corrected, and lemmatized by Beatrice Santorini. See the public repository and documentation. The control comprises 1,412 lemmatized letter units dated 1560–1625 and 665,762 tokens. The Neville inventory contains 171 documentary records after removal of the duplicate 1606 letter; excluding three French records leaves 168 analytical records (165 letters and three longer documents) and 138,057 normalized lemma tokens. Nested quoted and enclosed material remains. The XML and record selection are the same as Chapter 12, but this study retains its established normalization, including numerical tokens; Chapter 12’s alphabetic-only count is therefore lower. Neither count is the separately defined descriptive census in Chapter 11.↩︎

  3. Original analysis: 180 features in fifteen cells, formed by three language categories and five feature classes, with twelve features per cell. The classes were single lemmas, consecutive two-lemma strings, consecutive three-to-five-lemma strings, and ordered three- and four-lemma sequences omitting between one and three words in total. Positive distinctiveness weights were limited and normalized within each cell; cells contributed equally. The protocol and checksum record constitute an internal freeze, not independently verified public preregistration. In the updated run the feature set was rebuilt using the established source-side rules before rescoring the plays; all 180 selected features were retained, with slightly revised weights after deduplication. The earlier target results were already known. The illustrative features in the text are actual selected patterns, not independent historical quotations. The numerical results are original research for this book; public deposit of the full study materials remains outstanding. Some scored features overlap, so removing one feature does not remove every feature containing the same expression. The reported deletion checks concern individual features, not whole families of overlapping expressions.↩︎

  4. Original work-score analysis, 236 plays dated 1590–1615: 36 First Folio plays and 200 others. Shakespeare-group mean 0.394180; comparison mean 0.329628. Henry VIII (full play) scored 0.593528; its Shakespeare part, 0.562960; The Tragedy of Charles Duke of Byron, 0.586460. The supplementary part uses the database’s Shakespeare division; no contiguous or skip-gram feature bridges an omitted passage. The table scores are not length-matched; the full play has more opportunities to contain a feature. Only complete plays enter the group models. The selection follows Pervez Rizvi’s play catalogue, part of his Collocations and N-grams collection of 527 plays, with Pericles, Edward III and The Two Noble Kinsmen excluded from both groups. The database’s dates and attribution divisions are used as supplied. Rizvi identifies EarlyPrint and Folger Digital Texts as his sources. This study applies its own normalization; full text-by-text reproducibility materials remain to be deposited.↩︎

  5. Rerun secondary analyses. Diplomatic/epistolary language: standardized effect 0.452, adjusted q = 0.1540; legal/institutional: 0.299, q = 0.2983; remaining lexical category: 0.290, q = 0.1633. Benjamini–Hochberg adjustment was applied within this three-test family. Ordered three- and four-word sequences had standardized effects 0.483 and 0.422, both with q = 0.1812 within the separate five-class family. The consecutive three-to-five-word class had effect 0.274, q = 0.3528. None of these adjusted secondary tests met .05; they do not replace the primary result.↩︎

  6. Rerun primary analysis: score regressed on First Folio-group membership, centered year, year squared, and log token count. Freedman–Lane residual permutations operated within twenty period-and-length blocks. Of 20,000 permutations, 2,338 matched or exceeded the observed coefficient; the add-one calculation is 2,339/20,001 = 0.1169442. The 95% interval from 10,000 work resamples within groups was 0.002146–0.063705. The interval and permutation test use different procedures; the protocol assigned the primary decision to the latter. Leave-one-work coefficients ranged from 0.025345 to 0.036999; leave-one-feature coefficients from 0.031141 to 0.034180.↩︎

  7. Rerun calibration using EarlyPrint: 719 metadata-selected proxy texts matched by date and length to 719 other texts. Mean paired score difference 0.059540; bootstrap 95% interval 0.051701–0.067477; two-sided sign-flip p = 0.000050. The legal/institutional component’s mean difference was 0.142353. Metadata selection supplies proxies for register, rather than individually verified genre classifications.↩︎

  8. Rerun sensitivity analyses. The three-work exclusion left 233 plays, including 33 First Folio plays: coefficient 0.028848, standardized effect 0.399361, p = 0.186741; work-bootstrap 95% interval -0.001361–0.057986. The narrower correspondence control produced coefficient 0.035981, standardized effect 0.494853, p = 0.076646; interval 0.006421–0.066860. Neither is an independent author-contrast replication. Supplementary Henry VIII parts do not enter these analyses.↩︎

Note

Search the book

Inspect the document