Dataset catalog

Domain collections, ready for diligence.

Human-authored, provenance-documented collections for pretraining, fine-tuning and evaluation — each with document-precise timestamps, sealed lineage and a data quality report.

Gated preview · full slice access under NDA
Flagship collections

Preview the catalog.

Sample records show the metadata that ships with every record. Full slices, checksums and rights details are released under NDA with per-record audit rights.

Peer-reviewed science, technical monographs and standards — OCR-verified with formula-safe extraction and citation-preserving structure.

1.34Ttokens
28.4Mdocuments
612archival sources
1974–2021coverage span
Temporal coverage · 1974–2021
188019502022 · synthetic era
Sample record · metadata view Gated
doc_id        STM-1996-11872
source_title  Physics journal backfile, Vol. 53
captured      1996-03-14 · pre-synthetic
checksum      ████████████████ gated
tokens        3,844
lineage       sealed · LS-77190
rights        ████████████ gated
Full slices include complete metadata, per-record lineage and audit tooling. Request access: STEM Corpus

Also available on request: Multilingual · Speech & audio transcripts · Enterprise documents Ask about a collection

Products by training stage

The right data for each phase of a model's life.

Every product shares the same sourcing, curation and validation standards — packaged for how you'll use it.

Pretraining

Foundation corpora

Petabyte-scale, deduplicated human-authored text and code for foundation models.

  • Versioned, checksummed shards
  • Lineage for every record
  • Domain and era mixing controls
Fine-tuning

Instruction & preference sets

Written or verified by vetted experts and sized for small and mid-size models.

  • Instruction, preference and domain sets
  • Human-authorship attestation
  • Deduplicated against public benchmarks
Evaluation

Held-out evaluation sets

Never-published test sets, so benchmarks measure your model rather than its memory.

  • Contamination-screened by construction
  • Domain-specific suites on request
  • Refreshed to stay unpublished
In every delivery

What ships with every dataset.

Data you can defend comes with its paperwork. Each release includes the documents your ML, legal and security teams will ask for.

Data card

Composition, intended uses, known limitations and collection methodology, in a standard format.

Data quality report

Purity score and the six measured dimensions for the exact release you receive. See a sample.

Lineage ledger

Signed, per-record provenance from source to shard — verifiable with your own tooling.

License & indemnity docs

License class per slice, rights-clearance documentation and indemnification terms.

Changelog & versioning

What changed between releases and why, with deprecation notices ahead of removals.

Delivery manifest

Shard inventory with SHA-256 checksums, schemas and loader examples for your format.

How access works

From first call to production license.

  1. 1
    Scoping call

    30 minutes on model goals, domains, scale and compliance constraints.

  2. 2
    Private catalog under NDA

    Full slice inventory, sample records, lineage ledgers and licensing terms.

  3. 3
    Pilot slice

    A production-representative slice with its full validation report.

  4. 4
    Production license

    Versioned delivery to your infrastructure, with updates and support.

Have your own data?

Run your existing corpus through the Pureset pipeline and get a quality report, cleaned shards and a remediation plan.

About data audits

Get the private catalog.

Complete slice inventory, sample records and lineage ledgers — shared under NDA after a 30-minute scoping call.