Dataset honesty

Coverage & scope

Every number below is computed from committed data at build time — not hand-typed. Regenerate with npm run coverage and npm run variant-index after changing witness text.

Last computed: 14 Sept 2026, 11:52

Students and debaters: start with Use this site for pathways and claim discipline before quoting our counts in an argument. Researchers: see Methods for scope windows and reproducibility.

The Ehrman question — and what we answer

What Ehrman asked

Bart Ehrman's talking point is about volume: nobody has counted all variants across the entire Greek New Testament manuscript tradition (~5,700+ witnesses). The scale is enormous — and that claim is about the full tradition, not one papyrus fragment.

What Gurry added

Peter J. Gurry (NTS 2016) extrapolated roughly 500,000 distinct readings from ~3% of the text — excluding spelling and nomina-sacra abbreviations, including nonsense and singular readings. That is an estimate, not a census.

What this site answers

A defined census of extant-letter disagreements in Greek NT witnesses overlapping 1400 CE, compared word-by-word to open SR GNT via CNTR transcriptions:

108 witnesses · 37,592 variation units

Classified by kind (omission, addition, substitution, orthography, transposition), browsable in the variant explorer or witness compare.

What we do NOT claim

  • A full-tradition NT census (~5,700+ manuscripts)
  • Gurry's ~500,000 as our own number
  • ECM or NA apparatus judgments
  • Heuristic intentional tags as settled scholarship (see experimental section below)

Real disagreements — not abstract numbers

Each row below is one counted variation unit from our CNTR slice. Click through to browse all 37,528 in the variant explorer.

Download open census

Machine-readable export of our 1400 CE Greek NT variation units vs SR GNT — the same counted index that powers the variant explorer. This is not a full New Testament tradition census; it covers CNTR transcriptions in this repository only. Whole missing blocks (e.g. John 7:53–8:11) are taught on Famous passages, not as variant rows.

Columns: unit_id, witness, verse, kind, witness_reading, sr_reading. Licenses and provenance: DATA.md (CNTR CC BY-SA 4.0, SR GNT CC BY-SA 4.0).

How to use this site

For whole missing or inserted passages (not word-level rows), see Famous passages — block-level issues taught separately from the variant census.

  1. Pick a corpus on the timeline (Greek NT, Qurʾān, or Nag Hammadi).
  2. Browse counted disagreements in the variant explorer — filter by kind, witness, or book.
  3. Read block-level famous passages (Pericope Adulterae, Mark's ending, Comma Johanneum, etc.) on Famous passages.
  4. Open a witness card: photograph (when legal), diplomatic text, and variant strips under verses.
  5. For full critical work, follow links to CNTR, INTF / NTVMR, or IGNTP.

New to manuscripts?

A quick primer for students and curious readers. No prior coursework required.

Manuscript, papyrus, codex
A manuscript is any handwritten book or scrap before printing. Papyrus is plant-fiber sheets (common in Egypt). A codex is a bound book of pages — the ancestor of modern books. Most early NT scraps are papyrus fragments; later witnesses are often parchment codices.
Paleographic dates are ranges
Scholars rarely know the exact year a scribe wrote. Cards show a range (e.g. 125–175 CE) from handwriting style, not a single “birth year” for the text.
Witness vs modern critical text
A witness is one physical copy — P52, Codex Vaticanus, a Qurʾān leaf. A critical text (here, SR GNT from CNTR) is an editor’s reconstruction of the earliest recoverable Greek, built from many witnesses. Comparing one witness to SR GNT shows where that copy differs, not “the right reading” by itself.
Variation unit vs reading
A variation unit is one spot where witnesses disagree (e.g. a word in John 1:1). Each possible spelling there is a reading. Our site counts word-aligned variation units in surviving text only — classified as orthography, omission, addition, substitution, or transposition (mechanical v1 rules). Intentional-vs-error labels are not assigned yet.
Tiny real example from our data
P129, 1 Cor 7:35: witness letters μων vs SR GNT αυτων (substitution).
Block-level famous passages
Some well-known stories (e.g. the woman caught in adultery, John 7:53–8:11) are whole blocks missing from early witnesses — they do not show up as neat rows in our word-level census. See Famous passages for those issues, separate from the variant explorer.
Qurʾān regional rasm (not Greek variants)
Word-level Greek Variants count CNTR disagreements in the NT. Early Qurʾān witnesses share one Uthmanic Text Type; medieval sources record ~35–40 tiny regional rasm differences between Syria, Medina, Basra, and Kufa — not alternate Qurans. See Uthmanic regional rasm.
Why lacunae matter
Fragments have holes. If software treated every gap as a deliberate omission, you would get absurd “variant” counts. We only compare extant letters — lacunae ([...]) and editorially supplied reconstruction (~) are excluded. That is why our numbers are smaller and more honest than naive string diffs.

Greek New Testament (1400 CE)

108 witnesses · 37,592 variation units vs SR GNT · 66 with leaf image

108witnesses in dataset
89with CNTR transcription
66with leaf image
37,592variation units vs SR GNT

Extant text compared

388,200 Greek word tokens in extant (non-supplied) runs across all CNTR verses we store, including lazy-load overflow files. Lacunae and editorially supplied letters are excluded.

Taxonomy census (mechanical classifier)

Word-aligned variation units in extant runs vs SR GNT (CNTR), using scripts/lib/variant-classify.mjs. Lacunae and reconstructed supplied text are not counted. Intentional-vs-error labels are not assigned in this mechanical pass.

37,592 total variation units — browse in the explorer →

  • Median per witness: 6
  • Maximum per witness: 17562
  • Witnesses with ≥1: 86

By kind

KindCount% of total
Orthography

Spelling or pronunciation spelling only (itacism, movable nu, nomina sacra, diacritic/breathing-insensitive).

17,76847.3%
Omission

SR GNT has word(s) the witness lacks in the aligned extant span.

2,5506.8%
Addition

Witness has word(s) SR GNT lacks in the aligned extant span.

3,5309.4%
Substitution

Different lexical content at the same aligned position (not explainable as orthography alone).

13,38035.6%
Transposition

Same letters or word multiset in different order (cheap mechanical check only; may miss complex cases).

3641%

By book (Matthew → Revelation)

Per-book totals of word-aligned disagreement units in our CNTR slice — computed from the same census as the variant explorer, not hand-typed.

BookTotal% of sliceSubstitutionOrthographyOmissionAdditionTransposition
Matthew4,42911.8%1,3552,33228442335
Mark2,7647.4%9031,36521026323
Luke5,56514.8%1,9692,53642757162
John6,32216.8%2,5182,57550365472
Acts4,18811.1%1,3582,20023136039
Romans1,8584.9%6799021171564
1 Corinthians2,0885.6%7311,01815416223
2 Corinthians1,3363.6%4736796910510
Galatians7151.9%28633933507
Ephesians7051.9%30232030485
Philippians4961.3%17923926493
Colossians3771%12919020353
1 Thessalonians3330.9%10915728345
2 Thessalonians1730.5%688211111
1 Timothy1530.4%50819130
2 Timothy1300.3%47647120
Titus920.2%28484111
Philemon440.1%2511440
Hebrews1,3063.5%4416616712017
James4061.1%12422721313
1 Peter7952.1%23343446739
2 Peter4591.2%13225323483
1 John5011.3%12330532383
2 John600.2%20224140
3 John560.1%1734230
Jude2260.6%121747231
Revelation1,9515.2%91160718121735
All books37,592100%13,33117,7552,5503,528364

Experimental: intentional vs error tagging

Provisional only. Model-assisted error / intentional / uncertain labels on a subset of non-orthography variation units. Provisional hypotheses only — not ECM or NA judgments. These labels are model-assisted hypotheses — not ECM, NA28, or IGNTP judgments. Do not cite tag counts as scholarship.

5,463 units tagged of 19,773 taggable non-orthography units (27.6% of taggable subset)

  • Error: 2,099
  • Intentional?: 36
  • Uncertain: 3,328

INTF Liste papyri overlap

Of 106 Gregory-Aland papyri whose Liste range overlaps 1400 CE, we include 106 (100%).

CNTR gaps in our build

No CNTR transcription file for: P103, P16, P65, P78, P80, P12, P10, P50, P62, P7, P99, P105, P112, P127, P140, P54, P56, P93, P94.

Qurʾān (1–100 AH)

19witnesses
15/19with leaf image (79%)

No NA-style variant census for Qurʾān in v1. Cards show reference rasm and library links, not a full manuscript collation.

For evidence-backed regional rasm within the Uthmanic Text Type (~35–40 reports in Cook / Sidky), see Uthmanic regional rasm — distinct from our per-card excerpt display and from the Ṣanʿāʾ lower-text companion codex.

Hebrew Bible & Septuagint (250 BCE–400 CE)

24witnesses in seed
9 + 15Hebrew DSS + Greek LXX
16/24with leaf image (67%)

Hand-curated DSS + LXX papyri/uncials overlapping the Hebrew/LXX timeline. Diplomatic excerpts + WEB English on cards — not BHQ or Göttingen apparatus dumps.

Medieval Masoretic Text (MT) codices are centuries later than Qumran Hebrew; Dead Sea Scrolls show earlier Hebrew diversity. The Septuagint (LXX) is a separate early Greek translation stream — compare both to the site's Greek NT slice (1–400 CE), not as a single 'original Bible' line.

Switch to Hebrew / LXX on the timeline for cards with Leon Levy / library links (IAA scans not rehosted).

Nag Hammadi (~300–400 CE)

10tractate witnesses
10/10with leaf image (100%)

No Thomas↔NT collation in v1. These are Coptic Gnostic tractates, not Greek New Testament witnesses.