- Entry date
- 27 April 2026
- Category
- Models
- Lead over the world
- a working baseline, not a lead
- Access
- 🔓 Public
A tokenisation run died at shard 368 of 1,500 — and the handover letter says so, right next to the parts that are 100% verified.
- 202,906
- domain-specific-language training pairs, 100% verified
- 60–80%
- academic-text verification coverage
- 3,261 files / 51GB
- corpus scale (dolphin_data)
Honest evaluation
Proven as a working, honestly-scoped handover — the DSL channel is 100% verified and the corpus fed a real downstream training run; the dead tokenisation run and partial academic verification are disclosed, not hidden.
What would prove or disprove it further
What would prove or disprove it further: resume and complete the academic-text verification pass and the failed Cosmopedia tokenisation run, then check whether the now-fully-verified corpus changes downstream training outcomes measurably. A corpus that "was mostly fine anyway" would validate the honest-partial-handover approach; a meaningful downstream difference would show the disclosed gaps mattered in practice, not just on paper.
The evidence — full reasoning behind the verdict
Verdict: proven. The hypothesis being tested here was narrower than "build a perfect corpus" — it was that training data could move from ad-hoc experimentation to a structured, honestly-recorded lab process. That's proven: a 50-billion-plus-token corpus was sharded and manifested, a 202,906-pair domain-specific-language channel reached 100% verification, and the handover artefact itself became a real input to a later training run (the Dolphin trainer's Recipe-C), not a document that sat unused.
The mechanism was in what got written down, not just what got curated. A dead tokenisation run at shard 368 of 1,500, and academic-text verification stalled at 60-80%, are both logged directly in the handover record next to the fully-verified parts, so a downstream user knows exactly what they're getting rather than being told everything is done.
Runnable proof — see it work
A corpus is only as useful as the record of what's actually in it. This one shipped with that record attached.
What it was
Dolphin Data (/dolphin_data/) is a sharded .tokbin training corpus:
academic text, FineWeb-Edu, ArXiv, Gutenberg, code, and a fully verified
domain-specific-language training channel. Dolphinprism (/dolphinprism/)
is the formal handover artefact — a letter, charter corrections, and council
records — moving training from ad-hoc experimentation to a structured lab
process with a named owner and a named state.
What we built
ResearchThe realised corpus: 50 billion-plus tokens, sharded and manifested. Within it, a 202,906-pair domain-specific-language channel reached 100% verification — every pair checked, not sampled. Academic text sat at 60–80% verified. Scale: roughly 3,261 files and 51GB for the corpus itself; the handover artefact was small by comparison, 20 files and 229KB. The handover completed over 26–27 April 2026.
What we learned — including the honest negative(s)
- A curated corpus is not a complete one. The handover record states this plainly rather than rounding "curated" up to "done" — academic-text verification stopped at 60–80%, and the record says so.
- A dead run is still useful information. A Cosmopedia-source tokenisation pass died partway through, at shard 368 of 1,500. Rather than quietly re-running it and pretending it always worked, the failure point was logged into the handover record — future auditors know exactly where to resume.
- Sharded and manifested beats merely large. The real value of this corpus for anyone downstream wasn't its raw token count — it was that every shard is traceable back to a source and a verification state.
- A handover is a discipline, not a formality. Moving training data from ad-hoc experimentation to a structured lab process meant writing down who now owns the corpus, what state it's actually in, and what still needs checking — the kind of unglamorous documentation work that determines whether a 50-billion-token corpus is reusable a year later or just a large pile of files nobody trusts.
Where it went / status
Handed over and consumed. This corpus became the real input to the Dolphin trainer's Recipe-C run inside the wider Unification project — a live, improving training pipeline, not a demo corpus sitting unused on disk.
What is still open — kept visible
The honest edges, next to the wins. This is what turns 🔬 into 🟢 — honestly.
- A Cosmopedia-source tokenisation run died partway through, at shard 368 of 1,500 — flagged in the handover record, not hidden.
- Academic-text verification sits at 60–80%, not complete, at the time of handover.
- 'Curated' is not 'complete' — the handover record itself says so, plainly.
Where this connects
Sources
- vocabotics project audit — Dolphin Data / Dolphinprism handover (50B+ tokens, verified DSL channel), April 2026vocabotics internal project history · as of April 2026