Training Sets

1 dataset

NameTypeSizeUploaded
Concordance-EN v10

Dense vocabulary-coverage English prose corpus. 225MB / 40.6M words / ~54M tokens / 497,737 paragraphs. Built from Simple English Wikipedia (ns=0) + 9,459 Project Gutenberg books; SCOWL-derived 593,704-word target vocabulary, 308,956 covered (52%). No LLM used - deterministic pipeline. English-only via function-word-ratio test; Wiktionary and 1,051 Gutenberg reference works (Webster's, Roget's, censuses) excluded. Verified: 0.00% HTML entities, 0.00% orphan markup, 0.54% non-sentences. Card + reproduction: https://huggingface.co/datasets/ajaxdavis/concordance-en SHA-256 8b54bb7fa519941f2d085352cff0c4c6655c2117eb27fadca44b521bc03e230b

txt225.0 MB2026-08-07