Foundation corpora
Petabyte-scale, deduplicated human-authored text and code for foundation models.
- Versioned, checksummed shards
- Lineage for every record
- Domain and era mixing controls
Human-authored, provenance-documented collections for pretraining, fine-tuning and evaluation — each with document-precise timestamps, sealed lineage and a data quality report.
Sample records show the metadata that ships with every record. Full slices, checksums and rights details are released under NDA with per-record audit rights.
Peer-reviewed science, technical monographs and standards — OCR-verified with formula-safe extraction and citation-preserving structure.
doc_id STM-1996-11872 source_title Physics journal backfile, Vol. 53 captured 1996-03-14 · pre-synthetic checksum ████████████████ gated tokens 3,844 lineage sealed · LS-77190 rights ████████████ gated
Also available on request: Multilingual · Speech & audio transcripts · Enterprise documents Ask about a collection
Every product shares the same sourcing, curation and validation standards — packaged for how you'll use it.
Petabyte-scale, deduplicated human-authored text and code for foundation models.
Written or verified by vetted experts and sized for small and mid-size models.
Never-published test sets, so benchmarks measure your model rather than its memory.
Data you can defend comes with its paperwork. Each release includes the documents your ML, legal and security teams will ask for.
Composition, intended uses, known limitations and collection methodology, in a standard format.
Purity score and the six measured dimensions for the exact release you receive. See a sample.
Signed, per-record provenance from source to shard — verifiable with your own tooling.
License class per slice, rights-clearance documentation and indemnification terms.
What changed between releases and why, with deprecation notices ahead of removals.
Shard inventory with SHA-256 checksums, schemas and loader examples for your format.
30 minutes on model goals, domains, scale and compliance constraints.
Full slice inventory, sample records, lineage ledgers and licensing terms.
A production-representative slice with its full validation report.
Versioned delivery to your infrastructure, with updates and support.
Run your existing corpus through the Pureset pipeline and get a quality report, cleaned shards and a remediation plan.
Complete slice inventory, sample records and lineage ledgers — shared under NDA after a 30-minute scoping call.