Licensed archives
Periodicals, monographs, filings and correspondence, captured from the physical and digital archives that hold them.
Pureset combines licensed sourcing, deterministic curation and independent validation in a single pipeline. Every step is measured, documented and sealed, so the data you train on is data you can defend.
Before a record becomes a training token, it was a bound volume, a filed report, a signed letter — or work commissioned from a verified expert. Pureset licenses the sources where that trail still exists and captures it with immutable certificates.
Periodicals, monographs, filings and correspondence, captured from the physical and digital archives that hold them.
Crawls with document-precise capture dates, filtered down to verifiably human-authored sources.
Publishers, institutions and communities licensing their collections to Pureset directly.
Instruction, preference and evaluation data written by vetted specialists, with human-authorship attestation.
Four deterministic stages carry every record from its source to a training-ready, cryptographically sealed shard. Select a stage to inspect its controls.
Licensed archives, scans and first-party holdings are captured with immutable certificates at the point of origin.
→ capture manifests · SHA-256 at sourceMinHash LSH clustering, quality classifiers and exact-match sweeps remove noise without discarding rare human signal.
→ ≤ 1% residual duplicate rateSource, era and demographic distributions are measured against documented baselines and rebalanced before release.
→ distribution report per sliceEvery shard is hashed into a signed provenance ledger — verifiable end to end by your own audit tooling, not just ours.
→ ledger seal · customer-side verification“Clean” is a measurable property. These are the checks every release must pass — and the numbers that ship in its quality report.
Boilerplate, OCR errors, spam and machine-generated text removed with quality classifiers and perplexity filters tuned to keep rare human signal.
Exact, near-duplicate and semantic duplicates removed within the corpus and against prior releases, so nothing is learned twice.
Source, era, region and demographic distributions measured against documented baselines, rebalanced before release and reported per slice.
N-gram and embedding screening against public benchmarks such as MMLU, GSM8K and HumanEval, so evaluations measure capability — not memorization.
Names, contact details and identifiers detected and scrubbed, with residual flags checked by human reviewers on audited samples.
Every record traced to its source, capture date and license, with each custody hand-off signed into a ledger.
Every release ships with a quality report your team can check — and re-verify with its own tooling.
Composite of six measured dimensions across 28.4M documents.
Versioned, checksummed releases in open formats — pushed to your infrastructure, never locked into ours.
Sharded, schema-documented files that load straight into standard data loaders.
Delivered to private cloud buckets or a warehouse share, or over SFTP for air-gapped environments.
Semantic versions, changelogs and deprecation notices. Pin a version for reproducible training runs.
Verify any shard's lineage programmatically — or with your own audit tooling. Documentation is shared with licensed teams during onboarding.
Not for what it's used for. Language, reasoning and domain knowledge written before the synthetic era remain the richest source of verifiably human signal.
For recent knowledge, Pureset pairs archival data with commissioned expert data and dated, attested collections — and every record carries its capture date, so you decide exactly what belongs in your mix.
Six measured dimensions — noise, duplication, bias, contamination, PII and provenance. Each ships as a number in the data quality report, not as a marketing adjective.
Yes. A data audit runs your corpus through the same pipeline and returns a quality report, deduplicated shards and a remediation plan.
Sources are licensed, public-domain or commissioned, and rights are cleared and documented before ingestion. Enterprise licenses include rights documentation and infringement indemnification. More on the trust page.
Most pilot slices ship within three weeks of the scoping call, together with their full validation report.
No. Every engagement is scoped to your domains, scale and compliance needs. A scoping call takes 30 minutes.
Scope a pilot slice with a data specialist — or send us a sample of your own corpus for a data audit.