Platform

One accountable pipeline — from source to training shard.

Pureset combines licensed sourcing, deterministic curation and independent validation in a single pipeline. Every step is measured, documented and sealed, so the data you train on is data you can defend.

Step 1 · Source

Sources with a paper trail.

Before a record becomes a training token, it was a bound volume, a filed report, a signed letter — or work commissioned from a verified expert. Pureset licenses the sources where that trail still exists and captures it with immutable certificates.

Stacks of printed newspapers in a press archive CAPTURED ≤ 2021
Periodical & press archives
licensed mirror · 1998–2021
Editorial workspace with printed manuscripts and reference material OCR-VERIFIED
Academic monographs
formula-safe extraction · 1974–2021
Authored documents and handwritten notes on a desk RIGHTS-CLEARED
Authored correspondence
indemnified · 1881–2021

Licensed archives

Periodicals, monographs, filings and correspondence, captured from the physical and digital archives that hold them.

Pre-2022 web captures

Crawls with document-precise capture dates, filtered down to verifiably human-authored sources.

First-party partnerships

Publishers, institutions and communities licensing their collections to Pureset directly.

Commissioned expert data

Instruction, preference and evaluation data written by vetted specialists, with human-authorship attestation.

Step 2 · Curate & verify

A pipeline built for provenance, not volume.

Four deterministic stages carry every record from its source to a training-ready, cryptographically sealed shard. Select a stage to inspect its controls.

STAGE 01

Licensed archives, scans and first-party holdings are captured with immutable certificates at the point of origin.

→ capture manifests · SHA-256 at source
STAGE 02

MinHash LSH clustering, quality classifiers and exact-match sweeps remove noise without discarding rare human signal.

→ ≤ 1% residual duplicate rate
STAGE 03

Source, era and demographic distributions are measured against documented baselines and rebalanced before release.

→ distribution report per slice
STAGE 04

Every shard is hashed into a signed provenance ledger — verifiable end to end by your own audit tooling, not just ours.

→ ledger seal · customer-side verification
STAGE 01Throughput 1.2 TB/hrCapture certificates issued at sourceFormats: PDF · XML · TIFF · DjVu · EPUB · HTML
Step 3 · Validate

Six dimensions of clean.

“Clean” is a measurable property. These are the checks every release must pass — and the numbers that ship in its quality report.

Noise

Boilerplate, OCR errors, spam and machine-generated text removed with quality classifiers and perplexity filters tuned to keep rare human signal.

Reported: tokens filtered per slice

Duplication

Exact, near-duplicate and semantic duplicates removed within the corpus and against prior releases, so nothing is learned twice.

Target: ≤ 1% residual near-duplicates

Bias

Source, era, region and demographic distributions measured against documented baselines, rebalanced before release and reported per slice.

Reported: distribution vs. baseline

Contamination

N-gram and embedding screening against public benchmarks such as MMLU, GSM8K and HumanEval, so evaluations measure capability — not memorization.

Reported: overlap per benchmark suite

PII

Names, contact details and identifiers detected and scrubbed, with residual flags checked by human reviewers on audited samples.

Reported: flags per 10k-document audit

Provenance

Every record traced to its source, capture date and license, with each custody hand-off signed into a ledger.

Target: 100% lineage coverage
Data quality report

Measured, not asserted.

Every release ships with a quality report your team can check — and re-verify with its own tooling.

  • Purity score with a per-dimension breakdown
  • Residual duplication and benchmark contamination results
  • Bias distribution against documented baselines
  • PII audit results and license class per slice
  • Ledger seal and checksums for customer-side verification
data quality report · ps-stem-2021 · release 3.2 Passed
Purity score

Composite of six measured dimensions across 28.4M documents.

Noise6.1% of tokens filtered97
Duplication0.9% residual near-dups99
Bias balancewithin ±5% of baseline93
Contamination0.02% n-gram overlap99
PII0 flags · 10k-doc audit100
Provenance100% of records sealed100
Documents by capture decade
1970s1980s1990s2000s2010–21
License · enterprise-clear Formats · Parquet · JSONL Residency · US / EU
ledger seal LS-77190 · sha-256 7d44e0b1…52d8sample · values illustrative
Delivery & integration

Delivered the way your training stack expects.

Versioned, checksummed releases in open formats — pushed to your infrastructure, never locked into ours.

Open formats

Sharded, schema-documented files that load straight into standard data loaders.

JSONLParquetWebDatasetArrow

Your infrastructure

Delivered to private cloud buckets or a warehouse share, or over SFTP for air-gapped environments.

S3GCSAzure BlobSFTP

Versioned releases

Semantic versions, changelogs and deprecation notices. Pin a version for reproducible training runs.

Provenance API

Verify any shard's lineage programmatically — or with your own audit tooling. Documentation is shared with licensed teams during onboarding.

FAQ

Questions teams ask first.

Isn't pre-2022 data stale?

Not for what it's used for. Language, reasoning and domain knowledge written before the synthetic era remain the richest source of verifiably human signal.

For recent knowledge, Pureset pairs archival data with commissioned expert data and dated, attested collections — and every record carries its capture date, so you decide exactly what belongs in your mix.

What does “clean” actually mean?

Six measured dimensions — noise, duplication, bias, contamination, PII and provenance. Each ships as a number in the data quality report, not as a marketing adjective.

Can you clean data we already have?

Yes. A data audit runs your corpus through the same pipeline and returns a quality report, deduplicated shards and a remediation plan.

How do you handle licensing and rights?

Sources are licensed, public-domain or commissioned, and rights are cleared and documented before ingestion. Enterprise licenses include rights documentation and infringement indemnification. More on the trust page.

How long does a pilot take?

Most pilot slices ship within three weeks of the scoping call, together with their full validation report.

Is there public pricing?

No. Every engagement is scoped to your domains, scale and compliance needs. A scoping call takes 30 minutes.

See the pipeline on your data.

Scope a pilot slice with a data specialist — or send us a sample of your own corpus for a data audit.