Blind Neville–Shakespeare Affinity Study: 180 Frozen Features Across 239 Plays
Topic: Blind Neville–Shakespeare Affinity Study
Governing result — 28 August 2026
This study is original quantitative research developed for the 2026 book. It is now a first-class AI Topic rather than an unrecorded book-only result.
The design compared Henry Neville's English correspondence with 1,412 historical control letters, selected source-distinctive language without using any play result, froze 180 weighted features in fifteen balanced cells, recorded the registry hash, and only then scored 239 plays dated 1590–1615. The target corpus contained 37 Shakespeare-group works and 202 comparison plays.
The Shakespeare group scored higher:
| Measure | Result |
|---|---|
| Shakespeare mean | 0.412633 |
| Comparison mean | 0.341913 |
| Adjusted group coefficient | 0.038781 |
| Standardized coefficient | 0.502424 SD |
| Work-bootstrap 95% interval | 0.007403–0.070829 |
| Probability that a random Shakespeare work exceeds a random comparison work | 69.8421% |
The primary confirmatory decision did not pass:
Of 20,000 blocked Freedman–Lane permutations, 1,095 produced a coefficient at least as large as the observed one. With the add-one correction, p = (1,095 + 1) / (20,000 + 1) = 0.054797. The frozen one-tailed gate was p < .05.
That number must never be rounded down or redescribed as statistically significant. The correct evidence state is: substantial, stable, positive exploratory corroboration that narrowly failed its own formal primary gate.
Freeze and source controls
- The frozen source run used Neville corpus v15: 168 documentary records in inventory, three French records excluded, 165 English analytical records, and 135,185 normalized lemma tokens. The later canonical v22 corpus must not be retroactively substituted into this historical freeze.
- The historical control was the lemmatized PCEEC2 layer, restricted to 1560–1625: 1,412 letters and 665,762 tokens. Public controls: PCEEC2 repository and Oxford Text Archive PCEEC record.
- The target population was frozen at 239 plays. Public corpus identities and limitations are documented by Folger's Early Modern English Drama project, data resources, and Shakespeare downloads.
- The frozen registry contains 180 rows. SHA-256:
abc35113bb160697ad001d6630ec1f9e8b3adfe5f07f156e585e344623bffd18.
Quoted source passages
This is a computational study. Historical quotations are not evidence units in the primary test; features were counted mechanically from declared corpora. Any illustrative Shakespeare quotation belongs in the relevant play topic and must use the Folger text.
What the protected secondary results show
| Protected stratum | Standardized effect | Bootstrap 95% interval | Permutation p | Corrected q |
|---|---|---|---|---|
| Diplomatic and epistolary | 0.588 | 0.0107–0.0689 | .00690 | .0207 |
| Legal and institutional | 0.298 | −0.0226–0.1293 | .2755 | .2755 |
| Lexical and relational | 0.414 | 0.0066–0.0373 | .0291 | .0436 |
The diplomatic/epistolary stratum is the strongest protected component. Ordered discontinuous patterns also carried more of the group difference than long exact strings:
| Feature class | Standardized effect | Permutation p | Corrected q |
|---|---|---|---|
| Ordered four-word skip-grams | 0.568 | .0137 | .0490 |
| Ordered three-word skip-grams | 0.563 | .0196 | .0490 |
| Exact two-word strings | 0.329 | — | .2442 |
| Single words | 0.355 | — | .2471 |
| Exact three-to-five-word strings | 0.233 | — | .4335 |
This favors a distributed pattern of order, syntax, and phraseological frames over a theory of long verbatim borrowing.
Robustness and limits
- Leave-one-play coefficients ranged from 0.0318 to 0.0430; leave-one-feature coefficients ranged from 0.0373 to 0.0403. No single work or feature created or reversed the effect.
- Excluding Henry VIII, Pericles, Timon, and Titus left a positive coefficient of 0.0338 (0.447 SD) but weakened the permutation result to
p = .1225. - Rebuilding the source signature with a narrower 1580–1620 PCEEC2 control produced coefficient 0.0354 and
p = .0722. This is a same-corpus-family sensitivity, not an independent replication. - George Chapman's The Tragedy of Charles Duke of Byron, not a Shakespeare-group play, had the highest raw score in the population. The method is therefore not an exclusive Shakespeare detector.
- EarlyPrint calibration showed that the signature recognizes diplomatic, state, and legal register outside Shakespeare. EarlyPrint is consequently both a calibration source and a warning against mistaking register for personal identity.
- The primary model did not include genre or named author as covariates. Comparison-author portfolios are not fully independent.
- The source and target corpora are edited, incomplete historical samples. The result cannot identify the unique cause of the affinity.
External register calibration
The frozen signature was applied to 719 EarlyPrint state, diplomatic, ambassadorial, treaty, or legal proxy texts, each matched to a generic text of similar date and length. Twenty thousand paired sign flips produced:
| Outcome | Mean proxy-minus-generic difference | 95% interval | Sign-flip p |
|---|---|---|---|
| Combined signature | 0.06453 | 0.05637–0.07282 | < .0001 |
| Diplomatic and epistolary | 0.00935 | — | .0028 |
| Legal and institutional | 0.14747 | — | < .0001 |
| Lexical and relational | 0.03676 | — | < .0001 |
This proves that register is a measurable alternative explanation. It simultaneously confirms that the signature belongs to Neville's documentary world and prevents an author-exclusive interpretation.
The earlier curated Top 30
Twenty-eight of the thirty previously curated Neville–Shakespeare parallels could be represented in the play database. They yield a standardized group effect of 1.31, blocked-permutation p < .0001, and a conditional matched-set tail of .0018. These are selection-conditioned exploratory results because the Shakespeare corpus participated in feature selection. The smaller blind estimate of 0.50 empirically shows how human curation can amplify a real underlying affinity.
Poetry boundary
The Sonnets, Venus and Adonis, Lucrece, and The Two Noble Kinsmen were scored descriptively under the frozen registry but were not included in the 239-play primary inference. Without an equivalently processed non-Shakespeare poetry comparison set, their scores cannot establish unusualness and support no inferential poetry claim.
Book-safe formulation
A source-side signature frozen before the target plays were opened finds a substantial and distributed Neville–Shakespeare group affinity. Its strongest protected component is diplomatic and epistolary language. The preregistered primary test narrowly missed the fixed threshold at p = .054797, and external calibration shows that elite administrative register is a measurable alternative explanation. The result is affirmative exploratory corroboration, not a unique authorship classification.
Citations
- PCEEC2 public repository.
- Oxford Text Archive PCEEC record.
- Folger Early Modern English Drama: About and Data.
- Folger Shakespeare downloads.
- EarlyPrint.
Notes on access
The source corpora are publicly identified above. The frozen registry, manifests, scripts, and full machine-readable output remain in the private project pending public deposit. Numerical results should therefore be cited as original analysis, with the public corpus identities supplied, rather than as claims already published by an external scholar.