The Wider Index

The Epstein files are bigger than any one copy of them. This page catalogs every open corpus of OCR'd Epstein-files text we could find — community mirrors, institutional releases, and our own OCR — and rates each one on three axes. The chart grid and table below recompute from whatever you filter; click a row for detail.

The known universe: 1,394,531 DOJ documents (~2,770,154 pages, EFTA-numbered per page), the two House Oversight estate releases (31,621 pages), 11,830 court records, and 16 BOP videos. This catalog tracks 37 sources and records 9 dead or vaporware leads so the next person does not chase them again.

Scale — documents per source (log, top 12 of 37)
4.1MEpstein-Files (embedded-text): 4,110,1453.5MSOTA OCR extract (DS1-12): 3,500,0002.1MEpstein-Pipeline / epsteinexposed.com: 2,146,0001.4MEpstein-research-data / epstein-data.com: 1,425,5651.4Mepstein-data: 1,424,6731.4MJmail open data: 1,413,4171.4MDOJ Epstein Library: 1,394,5311.4Mtommycarstensen.com/epstein (JSONL text): 1,380,9871.4Mepstein-files-ocr-complete: 1,380,9321.3MPlainSite USDOJ-EFTA collection: 1,336,982550kepstein-doj-document-index: 550,308400kepstein.dugganusa.com: 400,000
Fidelity vs usability (bubble = coverage)
12345135Epstein-Files (embedded-text) — F2 C3 U3SOTA OCR extract (DS1-12) — F3 C4 U2Epstein-Pipeline / epsteinexposed.com — F4 C5 U3Epstein-research-data / epstein-data.com — F3 C5 U5epstein-data — F4 C5 U4Jmail open data — F5 C4 U5DOJ Epstein Library — F3 C5 U2tommycarstensen.com/epstein (JSONL text) — F3 C5 U4epstein-files-ocr-complete — F4 C5 U5PlainSite USDOJ-EFTA collection — F1 C4 U3epstein.dugganusa.com — F2 C2 U2CourtListener / RECAP — F4 C3 U3epstein-files.org — F3 C2 U1ep-nov-12.greg.technology — F3 C2 U2House Oversight official containers — F5 C3 U3jmail-* annotated family — F4 C1 U4DocumentCloud (MuckRock) — F3 C2 U4epstein-files-20k — F2 C2 U4epstein-ranker-dataset — F2 C1 U3Estate Release 7 OCR'd PDFs — F2 C1 U2Epstein_case_leaked_OCR_results — F3 C1 U2FULL_EPSTEIN_INDEX — F2 C2 U4epstein-fbi-files — F5 C1 U4Pinpoint EFTA collection — F4 C1 U2LLM-structured email sets — F3 C1 U3COURIER Pinpoint collection (NOV 12) — F3 C1 U2EpsteinFiles — F2 C1 U4epstein-docs.github.io — F3 C0 U3Senate committee text — F5 C1 U4NARA / Presidential libraries — F3 C1 U2epstein-emails (Qwen VL) — F3 C1 U3fidelity →
Engine families in view
Undisclosed10Open VLM5Mixed5No OCR5Commercial OCR4Native text3Classic OCR3Commercial VLM2
License posture in view
None stated17Public domain / CC07Permissive5Terms-gated4CC-BY3CC-BY-NC1
When sources appeared (first seen)
2025-10: 1 source25-1012025-11: 2 sources25-1122025-12: 1 source25-1212026-02: 1 source26-0212026-03: 1 source26-0312026-08: 1 source26-081
Document-level vs page-level sources
docs = pagesEpstein-research-data / epstein-data.com: 1,425,565 docs / 2,900,000 pagesDOJ Epstein Library: 1,394,531 docs / 2,770,154 pagesepstein-files-ocr-complete: 1,380,932 docs / 2,700,000 pagesThis site (marker OCR): 46,582 docs / 70,160 pagesep-nov-12.greg.technology: 31,621 docs / 31,621 pagesHouse Oversight official containers: 31,621 docs / 31,621 pagesEstate Release 7 OCR'd PDFs: 20,000 docs / 20,000 pagesEpstein_case_leaked_OCR_results: 11,924 docs / 34,662 pagesepstein-fbi-files: 8,150 docs / 33,295 pagesepstein-docs.github.io: 400 docs / 2,000 pagesdocuments (log) →↑ pages (log)
Cross-validation: pairwise character error rate
gregkabassmarkerteylerjmail567ishumilingregkabassmarkerteylerjmail567ishumilingreg vs kabass: CER 7.64%greg vs marker: CER 0.00%greg vs teyler: CER 9.17%greg vs jmail567: CER 12.27%kabass vs greg: CER 7.64%kabass vs marker: CER 38.75%kabass vs teyler: CER 5.91%kabass vs jmail567: CER 56.25%kabass vs ishumilin: CER 40.53%marker vs greg: CER 0.00%marker vs kabass: CER 38.75%marker vs teyler: CER 9.17%marker vs jmail567: CER 63.99%marker vs ishumilin: CER 75.41%teyler vs greg: CER 9.17%teyler vs kabass: CER 5.91%teyler vs marker: CER 9.17%jmail567 vs greg: CER 12.27%jmail567 vs kabass: CER 56.25%jmail567 vs marker: CER 63.99%jmail567 vs ishumilin: CER 75.86%ishumilin vs kabass: CER 40.53%ishumilin vs marker: CER 75.41%ishumilin vs jmail567: CER 75.86%lower (blue) = closer agreement; n=450
Profile (select a row)
Fidelityecosystem median: 3.0median 3.0Coverageecosystem median: 2.0median 2.0Usabilityecosystem median: 3.0median 3.0grey = filtered-set median · colored = selected source
37 of 37 sources
PreviewSourceCoversDocsEngineLicenseF·C·USeenStatus
no shot
Epstein-Files (embedded-text)
Nikityyy (HF)
DS1DS2DS3DS4DS5+10
4.1MNative text
pypdfium2 embedded-text extraction — not OCR; scans yield nothing
Permissive2 / 3 / 3live
no shot
SOTA OCR extract (DS1-12)
Daniel Voyce (certant.ai)
DS1DS2DS3DS4DS5+7
Undisclosed
undisclosed 'SOTA' system; '64 x 3090 GPUs' claimed
None stated3 / 4 / 2fragile
Preview of Epstein-Pipeline / epsteinexposed.comEpstein-Pipeline / epsteinexposed.com
stonesalltheway1
DS1DS2DS3DS4DS5+11
2.1MOpen VLM
Not named. OCR appears only in paid-tier marketing ("Enhanced exports with OCR & cross-references", "$99 Full Dataset (OCR snippets ...)").
Terms-gated4 / 5 / 3live
Preview of Epstein-research-data / epstein-data.comEpstein-research-data / epstein-data.com
rhowardstone (solo)
DS1DS2DS3DS4DS5+11
1.4MClassic OCR
Not named ('OCR-extracted text' offered without engine credit).
CC-BY-NC3 / 5 / 5live
Preview of epstein-dataepstein-data
kabasshouse (HF); frontend epstein.academy
DS1DS2DS3DS4DS5+9
1.4MMixed
Gemini 2.5 Flash Lite (856,028 files) + Tesseract (531,279, mostly DS9); per-file engine attribution
CC-BY4 / 5 / 4live
Preview of Jmail open dataJmail open data
Jmail team (lukeigel et al.)
DS8DS9DS10DS11Emails+1
1.4MCommercial VLM
Reducto (commercial document-parse AI); video by Kino AI
Public domain / CC05 / 4 / 52026-02-24live
no shot
DOJ Epstein Library
US Department of Justice
DS1DS2DS3DS4DS5+11
1.4MMixed
mixed: natives + DOJ-side OCR (unproven)
Public domain / CC03 / 5 / 22025-12-01live
Preview of tommycarstensen.com/epstein (JSONL text)tommycarstensen.com/epstein (JSONL text)
Tommy Carstensen
DS1DS2DS3DS4DS5+11
1.4MUndisclosed
undisclosed (pdf_text extraction)
None stated3 / 5 / 4live
no shot
epstein-files-ocr-complete
ishumilin (HF)
DS1DS2DS3DS4DS5+7
1.4MUndisclosed
Proprietary automated OCR pipeline by Wild Ma-Gässli (wildma.ch)
Public domain / CC04 / 5 / 52026-03-19live
no shot
PlainSite USDOJ-EFTA collection
Think Computer Corp
DS1DS2DS3DS4DS5+7
1.3MNo OCR
none (explicitly)
Terms-gated1 / 4 / 3live
no shot
epstein-doj-document-index
LayerDynamics (HF)
DS1DS2DS3DS4DS5+7
550kNo OCR
None stated· / 3 / 3live
Preview of epstein.dugganusa.comepstein.dugganusa.com
DugganUSA
DS1DS2DS3DS4DS5+7
400kUndisclosed
Not mentioned.
None stated2 / 2 / 2live
Preview of CourtListener / RECAPCourtListener / RECAP
Free Law Project (501c3)
Court
111kMixed
Tesseract (documented since 2012) + natives
Permissive4 / 3 / 3live
no shot
This site (marker OCR)
epstein-index (us)
DS1DS2DS3DS4DS6+12
46.6kMixed
marker-pdf 1.5.5 + surya-ocr 0.12.1 for the 14,961 DOJ/FOIA documents; the 31,621 estate-release pages serve the ep-nov-12 mirror's OCR text (engine undisclosed)
Public domain / CC0· / 2 / 4live
Preview of epstein-files.orgepstein-files.org
Andrew Walsh ('tsardoz', Sifter Labs)
DS1DS2DS3DS4DS5+7
33.9kUndisclosed
AI (unnamed)
None stated3 / 2 / 1vapor
no shot
ep-nov-12.greg.technology
Greg Sadetsky
House OctHouse Nov
31.6kUndisclosed
No OCR engine found named anywhere (Kagi forum/web lenses, HN). Closest HN neighbor project (epsteinsphone.org by HN user toon-noot, repo github.com/Toon-nooT/epsteins-phone-reconstructed) says only: "I used an OCR + vision-LLM pipeline to extract individual messages from the email screenshots" — a different project, often confused with this one.
None stated3 / 2 / 2degraded
no shot
House Oversight official containers
House Oversight Committee
House OctHouse Nov
31.6kNative text
estate natives + extracted text (not OCR)
Public domain / CC05 / 3 / 32025-10-17live
no shot
jmail-* annotated family
567 Labs (HF)
EmailsDS8DS9DS10DS11+1
28.0kCommercial VLM
Reducto (via jmail) + LLM annotations (severity, relevance, description)
None stated4 / 1 / 4live
no shot
DocumentCloud (MuckRock)
MuckRock /-TMU
Grand juryCourt
27.2kClassic OCR
DocumentCloud OCR (Tesseract family)
CC-BY3 / 2 / 4live
no shot
epstein-files-20k
teyler (HF); reuploads rbinrs, aurora2424, nglif
House Nov
25.0kClassic OCR
Tesseract
None stated2 / 2 / 4live
no shot
epstein-ranker-dataset
linogova (Kaggle)
House Nov
25.0kUndisclosed
locally-run AI (unspecified)
None stated2 / 1 / 3live
no shot
Estate Release 7 OCR'd PDFs
civicanger.com ('Default Account')
House Nov
20.0kUndisclosed
undisclosed ('OCR isn't 100% accurate')
None stated2 / 1 / 22025-11-15live
no shot
Epstein_case_leaked_OCR_results
LovenSar (HF)
DS1DS2DS3DS4DS5+7
11.9kOpen VLM
Local LLM: glm-ocr via Ollama + qwen3.5 router
None stated3 / 1 / 2removed
no shot
FULL_EPSTEIN_INDEX
theelderemo; mirrors genevera, B00MEMBERX
House NovDS1FBI VaultBOP video
10.0kUndisclosed
unnamed ('OCR made mistakes')
Permissive2 / 2 / 4live
no shot
epstein-images (FBI scans)
nynxz (HF)
FBI Vault
8.4kNo OCR
None stated· / 1 / 3live
no shot
epstein-fbi-files
svetfm (HF)
FBI Vault
8.2kCommercial OCR
AWS Textract with per-record ocr_confidence
CC-BY5 / 1 / 4live
no shot
Pinpoint EFTA collection
Brian Allen (private Google account)
DS1DS2DS3DS4DS5+7
7.5kCommercial OCR
Google OCR/NLP
Terms-gated4 / 1 / 2live
no shot
LLM-structured email sets
notesbymuneeb / Hannah2704 / pupepps; KillerShoaib; bru02 (HF)
Emails
5.1kOpen VLM
LLM structuring over released emails (incl. jmail re-scrape)
None stated3 / 1 / 3live
no shot
COURIER Pinpoint collection (NOV 12)
Camaron Stevenson, COURIER
House Nov
2.9kCommercial OCR
Google Pinpoint
Terms-gated3 / 1 / 2live
no shot
EpsteinFiles
Mark Ramm
House Oct
2.9kUndisclosed
OCR pipeline included (ocr_documents.py)
Permissive2 / 1 / 42025-11-14live
Preview of epstein-docs.github.ioepstein-docs.github.io
epstein-docs org
DS1FBI Vault
400Open VLM
OpenAI-compatible vision model returning text + entities
Permissive3 / 0 / 3degraded
no shot
Senate committee text
Senate Finance / Judiciary
Native text
native
Public domain / CC05 / 1 / 42026-08-04live
no shot
Epstein-Files (raw mirror index)
yung-megafone + community
DS1DS2DS3DS4DS5+7
No OCR
None stated· / · / 5live
no shot
NARA / Presidential libraries
National Archives
Mixed
varies
Public domain / CC03 / 1 / 2live
no shot
epstein (raw HF mirror)
AdcloseNN (HF)
DS1DS2DS3DS4DS5+7
No OCR
None stated· / 4 / 2live
no shot
epstein-emails (Qwen VL)
to-be (HF)
House NovEmails
Open VLM
Qwen 2.5 VL 72B
None stated3 / 1 / 3live
no shot
EpsteinFiles (Mistral OCR claim)
vikash06 (HF)
FBI Vault
Commercial OCR
Mistral OCR (claimed, unverifiable)
None stated· / 0 / 1card-only

How to read the ratings

Each source carries three 1–5 ratings from our census of 2026-08-28. They measure the source as a text provider, not the underlying documents.

Fidelity —
how trustworthy the OCR text is: engine quality, redaction handling, provenance. A source that deliberately rewrites text (for example, PII-redacted exports) scores below a verbatim engine.
Coverage —
how much of the known corpus the source covers. The denominator is the DOJ release plus the House estate pages.
Usability —
whether you can actually get and use the text: license, bulk access, structure, and whether rows map back to EFTA numbers or filenames.

Scale numbers are the maintainer's own claim unless we marked them verified against the files. Where the two censuses disagreed, we kept the figure we could check. "Undisclosed" engine means the maintainer never named it — not that there is no engine.

Cross-validation

Ratings describe sources; cross-validation tests them. We sample a few hundred pages that several sources have all OCR'd, align each pair's texts, and measure a normalized character error rate between them. We add two symbolic checks that need no judgment: whether digit runs (dates, amounts, EFTA stamps) survive, and whether the expected stamp regexes still match. Pairwise error is symmetric and reference-free: a high rate means the two engines disagree, and the heatmap shows who disagrees with whom.

We only believe a claimed problem when we can verify it independently — by checking a third source's take on the same page or by reading the page image itself. Findings that survive that check appear here; the rest stay out.

Current battery: 450 sampled pages in 4 strata, 24 pair-stratum rows scored, seed 42. Alignment on raw image basenames: House NOV/OCT filenames are shared exactly by marker md stems and greg (22,903 nov / 8,718 oct rows); teyler keys are 'IMAGES-<dir>-<STEM>.txt' (covers NOV 100%); jmail567 doc_id equals the OCT image stem (3,680 rows, all OCT); kabass HouseOversightEstate rows carry renumbered DOJ-OGR keys with no filename mapping, so they join via the HOUSE_OVERSIGHT_###### stamp found inside their own OCR text (6,794 unique stems; doc-level rows were cut to the '--- page N ---' segment containing the stem). DOJ joins on EFTA numbers: marker frontmatter eftaNumber vs ishumilin document_id vs kabass DataSet1 file_key vs jmail-doj doc_id prefix. Four batteries: nov (160 pages in 4 sources + 40 in 3), oct (80 in 4 + 20 in 3), doj (100 pages from the 3,145-doc 4-source intersection, photo-skewed) and doj-text (all 50 text-rich docs, marker ocrChars>=1000, in the same 4-source intersection). Normalization: lowercase, collapse whitespace, strip soft hyphens (U+00AD/U+2010); CER = rapidfuzz Levenshtein / max(len); digit Jaccard over multisets of digit runs >=3; stamps matched case-insensitively (EFTA\d{8} | HOUSE_OVERSIGHT_\d{6}). RNG seed 42; sampling code in /tmp/crossval.

Results of the 2026-08-28 battery — every finding below passed an independent check

  • The estate-release engines agree closely. On the 200-page Nov-12 stratum, pairwise median character error runs 2.6–5.0% and digit runs survive at 0.95–0.98 Jaccard across Tesseract (teyler), Gemini 2.5 Flash Lite (kabasshouse), and the text this site serves.
  • Provenance audit (verified by direct diff): the 31,621 estate-release pages this site serves carry the ep-nov-12.greg.technology mirror's OCR text. The marker|greg pair reads CER 0.000 because they are the same text, not because two engines agree. Our own marker-pdf OCR covers the other 14,961 ingested documents (DOJ, FOIA, and misc groups).
  • One image-verified reversal: on EFTA00003047, a rotated photograph of an emergency-contact form, our marker output shatters the rotated lines into per-letter fragments while Reducto (jmail) reads the page cleanly. Rotation is a real marker weakness on photo-style pages.
  • Photo pages measure differently: ishumilin stores image references instead of text on them, ishumilin and kabasshouse keep stamp-only text on 71–73% of the photo-skewed stratum, and jmail's DOJ drop carries VLM image descriptions rather than OCR there. A CER of 1.0 in this stratum means one side had no text — not that an engine failed.
  • Known limits: ishumilin is page-level while the other sets are document-level, so the text-rich DOJ stratum mixes levels and its high error rates (0.46–0.82 mean) must not be read as engine quality. The official NOV TEXT layer was not fetched within the battery's time-box and is next in line.

What we hold

Mirrored copies behind this site: the full Jmail open-data parquet set, the ishumilin and kabasshouse Hugging Face corpora, the small specialist sets (svetfm's Textract FBI files, the teyler 20k lineage, LovenSar's GLM output, the 567-labs annotated family), rhowardstone's structured-data mirror, and the only surviving copy of the ep-nov-12.greg.technology OCR database. We also hold both House Oversight official containers — the NOV container includes its TEXT and NATIVES; the OCT container does not. Ask before redistributing any of these: several carry no license or a non-commercial one, and this page records that rather than deciding it.

Method and provenance

The census ran on 2026-08-28 against live sites, APIs (Hugging Face, CourtListener, DocumentCloud, Wayback), and the maintainer's own documentation. Every claim in the underlying research carries its source URL, and screenshots show each site as we captured it. The catalog regenerates fromdata/catalogs/ecosystem/ via scripts/build_ecosystem_catalog.py; the ratings methodology follows the fleet's controlled-prose rules for evidence grading. The engine provenance of our own served text is documented indocs/ocr-engine-provenance.md. Corrections and new sources are welcome — the EFTA number or filename is the universal key, so adding a source is usually just metadata.