Skip to main content

Lab Notebook · Training Corpus

Dolphin — 50 billion tokens, sharded, manifested, and handed over honestly

Dolphin Data curated a 50-billion-plus-token training corpus and Dolphinprism formalised the handover — moving training from ad-hoc experimentation to a structured lab process, with an honest record of exactly what was and wasn't verified.

JR
Jon RossFounder, vocabotics — 15 years building safety-critical systemsLab report · dated 27 April 2026
Verified by a human. Drafted with AI, verified by a human. Jon Ross, 27 Apr 2026
Living document. Reviewed 27 Apr 2026
Entry date
27 April 2026
Category
Models
Lead over the world
a working baseline, not a lead
Access
🔓 Public
ResearchProven

A tokenisation run died at shard 368 of 1,500 — and the handover letter says so, right next to the parts that are 100% verified.

202,906
domain-specific-language training pairs, 100% verified
60–80%
academic-text verification coverage
3,261 files / 51GB
corpus scale (dolphin_data)

Honest evaluation

Proven

Proven as a working, honestly-scoped handover — the DSL channel is 100% verified and the corpus fed a real downstream training run; the dead tokenisation run and partial academic verification are disclosed, not hidden.

What would prove or disprove it further

What would prove or disprove it further: resume and complete the academic-text verification pass and the failed Cosmopedia tokenisation run, then check whether the now-fully-verified corpus changes downstream training outcomes measurably. A corpus that "was mostly fine anyway" would validate the honest-partial-handover approach; a meaningful downstream difference would show the disclosed gaps mattered in practice, not just on paper.

The evidence — full reasoning behind the verdict

Verdict: proven. The hypothesis being tested here was narrower than "build a perfect corpus" — it was that training data could move from ad-hoc experimentation to a structured, honestly-recorded lab process. That's proven: a 50-billion-plus-token corpus was sharded and manifested, a 202,906-pair domain-specific-language channel reached 100% verification, and the handover artefact itself became a real input to a later training run (the Dolphin trainer's Recipe-C), not a document that sat unused.

The mechanism was in what got written down, not just what got curated. A dead tokenisation run at shard 368 of 1,500, and academic-text verification stalled at 60-80%, are both logged directly in the handover record next to the fully-verified parts, so a downstream user knows exactly what they're getting rather than being told everything is done.

Runnable proof — see it work

A corpus is only as useful as the record of what's actually in it. This one shipped with that record attached.

What it was

Dolphin Data (/dolphin_data/) is a sharded .tokbin training corpus: academic text, FineWeb-Edu, ArXiv, Gutenberg, code, and a fully verified domain-specific-language training channel. Dolphinprism (/dolphinprism/) is the formal handover artefact — a letter, charter corrections, and council records — moving training from ad-hoc experimentation to a structured lab process with a named owner and a named state.

What we built

Research

The realised corpus: 50 billion-plus tokens, sharded and manifested. Within it, a 202,906-pair domain-specific-language channel reached 100% verification — every pair checked, not sampled. Academic text sat at 60–80% verified. Scale: roughly 3,261 files and 51GB for the corpus itself; the handover artefact was small by comparison, 20 files and 229KB. The handover completed over 26–27 April 2026.

What we learned — including the honest negative(s)

  • A curated corpus is not a complete one. The handover record states this plainly rather than rounding "curated" up to "done" — academic-text verification stopped at 60–80%, and the record says so.
  • A dead run is still useful information. A Cosmopedia-source tokenisation pass died partway through, at shard 368 of 1,500. Rather than quietly re-running it and pretending it always worked, the failure point was logged into the handover record — future auditors know exactly where to resume.
  • Sharded and manifested beats merely large. The real value of this corpus for anyone downstream wasn't its raw token count — it was that every shard is traceable back to a source and a verification state.
  • A handover is a discipline, not a formality. Moving training data from ad-hoc experimentation to a structured lab process meant writing down who now owns the corpus, what state it's actually in, and what still needs checking — the kind of unglamorous documentation work that determines whether a 50-billion-token corpus is reusable a year later or just a large pile of files nobody trusts.

Where it went / status

Handed over and consumed. This corpus became the real input to the Dolphin trainer's Recipe-C run inside the wider Unification project — a live, improving training pipeline, not a demo corpus sitting unused on disk.

What is still open — kept visible

The honest edges, next to the wins. This is what turns 🔬 into 🟢 — honestly.

  • A Cosmopedia-source tokenisation run died partway through, at shard 368 of 1,500 — flagged in the handover record, not hidden.
  • Academic-text verification sits at 60–80%, not complete, at the time of handover.
  • 'Curated' is not 'complete' — the handover record itself says so, plainly.

Where this connects

Sources

  1. vocabotics project audit — Dolphin Data / Dolphinprism handover (50B+ tokens, verified DSL channel), April 2026vocabotics internal project history · as of April 2026

    We use cookies.