Skip to main content

A vocabotics Guide

AI You Can Prove — the vocabotics capability brochure

The company capability brochure, cut from the ratified dashboard: four capabilities each carrying its own dated, measured proof number, trust certified at three layers, safety by construction, fifteen years of rail heritage — and exactly how to verify every claim yourself.

JR
Jon RossFounder, vocabotics — 15 years building safety-critical systemsDeep dive · 8 min read · reviewed 11 July 2026
Verified by a human. Drafted with AI, verified by a human. Jon Ross, 11 Jul 2026
Living document. Reviewed 11 Jul 2026
Shipped & measuredDownload the guidePDF · 2.5 MB

Most AI companies ask you to believe them. This brochure is built so you don't have to: every capability carries the actual measured number behind it, the date it was measured, the battery it was measured on, and an honesty tier — 🟢 shipped, 🔬 research, 🔭 vision — that says plainly how finished it is.

An earlier edition of this brochure led with "confabulation driven from 87% to 0.0%". We could not find a benchmark, corpus, sample size or baseline model behind either number, so we withdrew both rather than restate them, and replaced them with the counts we can name. What that correction cost us in headline is the point of the document.

The four capabilities

① It can't bluff — a property under test, and it has already failed once. Ask our knowledge system something it can source and it answers with the citation; ask it something it can't and it refuses rather than improvising. On the clean battery and two honesty batteries we wrote, it emitted zero confident-wrong answers, and the no-bluff floor held under a 16-turn adversarial interrogation. On a 44-turn battery of colloquial phrasing, the default path emitted nine. And the first instrument we built to question it from outside our own batteries found two confident lies on its first run — we retracted "it never lies" in our own design-of-record on 2026-08-01 and published the episode at /lab/never-lies-retracted. The demo: watch it refuse, then cite Canberra — with provenance — from a local knowledge store. (🔬 2026-07-02, retracted in part 2026-08-01)

② It stays fast when the context is huge. Our wave-based model architecture decodes at 2× llama.cpp's speed at 128,000 tokens of context — flat where a transformer climbs. Long-context regime specifically; not a claim of being fastest everywhere. (🟢 2026-07-02)

③ Proven-correct code, any chip, AI-written. Everything is written in a small language built for machines to write and machines to check. Every compile runs the program twice — native machine code and an independent typed interpreter — and the results must agree, value for value. The gate has never been disabled. The corpus holds 1,053 certified-equivalent functions and not one more. Yield depends entirely on the corpus, so we never publish it without one. (🟢 captured 2026-07-10)

CorpusCertified yieldBasis
Numeric C, pre-filtered to the long/double subset14 of 15 · 93%🟢 Measured, one run — the corpus the converter is built for
Pure algorithm, maths and codec cores (sort, hash, bignum, DCT, CRC, matrix)60–80%🔭 Projected by domain — not yet run end to end
Data-structure libraries (btree, hashmap, arena)35–55%🔭 Projected by domain — pointer glue drops out
Network, storage and OS glue (TCP state machines, VFS, blit)15–35%🔭 Projected by domain — mostly syscalls and state
Idiomatic concurrency, actors, GPU kernels~5%🔭 Projected by domain — effectively out of reach today
Wild, unfiltered C — 5 permissive repos, 1,423 repo-authored functions6.7% gate-green · 1.9% certified🟢 Measured, harvest run 3, 2026-07-10

The bottom rung is low because the converter refuses a plain int clamp(int, int, int) outright — C's 32-bit int wraps where our uniform 64-bit integer does not, so the two are not equivalent and it will not say they are. Most C ever written is typed int. The refusal is the product: a converter that certified 90% of wild C would be telling you something it could not check.

④ It runs on hardware you already own. The cognitive stack trains at a median 21 ms step in a 377 MiB peak on one consumer RTX 3090, over 250 steps — a training figure, not a serving one, and we publish no serving latency for this mind. A real open-weights model runs full interactive generation on our own stack at 151.7 tok/s, identical to the reference token-for-token across the 20 generated tokens of one prompt. Caveats, printed at the same size: that model is Qwen2.5-0.5B in f32, and 20 tokens on one prompt is an agreement rather than a decode-parity result. The head-to-head we ran separately, on an int4 short-context bench, puts us at 305 tok/s against llama.cpp's 313 — level, not ahead. 7B decode runs at parity, prefill is 5–30× behind, and llama parity holds on 2 of 3 prompts. (🟢 2026-07-02, caveats 🔬)

Buying on the four capability numbers above

AI reliability: Go

The vision organs (a world you own; new senses)

AI reliability: Amber

Trust, certified at three layers

Not "fastest" — a rival can copy a kernel. The claim is narrower and harder:

  1. The gate — code proven correct: every compile self-verified, native ≡ interpreter, never disabled. Mechanical trust.
  2. The calculated mind — cognition that reasons or refuses: honesty by architecture, not a bolt-on filter. Architectural trust.
  3. The cull — the founder archived 57 projects in one day, with a manifest and a one-command recovery left in the repo, and the archive keeps the failures visible. You cannot fake having subtracted. Cultural trust.

The honesty ledger

Every published claim is tiered: 🟢 shipped (measured, gate-clean, dated), 🔬 research (measured, partial, gaps named), 🔭 vision (designed, unbuilt). One law governs movement: the bigger claim goes UP a tier — never down.

The receipts: disproven projects stay published in full (including a retracted ternary-7B result — the retraction is part of the record); the chapter census is published as an audit, not a boast — 570 chapters audited, ~59% confirmed genuinely real, 60 stubs quarantined; and the corpus figure above is one function lower than the previous public number, because a false equivalence was found and quarantined. We printed the smaller number.

Safety by construction — and where we learned it

Run our compiler with a safety profile and it statically rejects code that allocates after initialisation, loops without a provable bound, recurses, exceeds length limits, skimps on assertions, or ignores a checked return — six checkers modelled on NASA's Power-of-Ten rules — and issues a certificate on code that passes. A memory leak is a compile error, not a review finding. (🟢 2026-07-02; WCET certificates are the designed next step, 🔭.)

The discipline comes from fifteen years delivering safety-critical systems for global rail and metro — including a $4m metro control system won and delivered — the world where a wrong announcement in an emergency is a safety incident, not a bad user experience.

What we sell today

We apply the same honesty tiers to the price list as to the research, which makes for uncomfortable reading. The company was incorporated on 20 August 2026 and has not yet delivered a paid engagement, so only the free scorecard can carry 🟢.

  • AI Readiness Scorecard — free, five minutes, three concrete next steps. Live and self-serve. 🟢
  • Strategy Audit — from £2,500: where AI genuinely helps your business, and where it doesn't yet. A plain-English plan, not a sales pitch. Offered and scoped; no client engagement delivered yet. 🔭
  • Compliance-grade AI Audit — for regulated firms: documentation an assessor expects, in the format your quality system already uses. We hold no certification ourselves and have not yet delivered this to a client. 🔭
  • Training & Workshops — calm, practical sessions for teams and communities. None delivered to a paying client yet. 🔭
  • No-Bluff Build — from £25k: cite-or-refuse knowledge tools, agentic workflows with a STOP button. Labelled honestly: the underlying systems are prototypes. 🔬

How to verify us

The close of this brochure is a checklist, not a slogan — about fifteen minutes:

  1. /proof — recorded gate transcripts from the live repo, including the one where a real 32-bit wraparound made the gate say DIFFER instead of papering over it.
  2. /proof/ledger — the census: 570 chapters audited, repaired, genuine and quarantined, counted in public.
  3. /dashboard — the flagship figures with tier and caveat attached, recomputed from the archive on every build.
  4. /lab — the dated R&D notebook. Read a disproven report in full, then decide what a 🟢 from us is worth.

If any of them disappoints you, you'll have learned something true about us — which is the point.

Sources

  1. The Demo — AI you can prove (recorded gate transcripts)vocabotics · as of 2026-07-10
  2. The Ledger — 570 chapters auditedvocabotics · as of 2026-07-10
  3. The Dashboard — the R&D archive, by the numbersvocabotics · as of 2026-07-10
  4. The Lab Notebook — the dated R&D archivevocabotics · as of 2026-07-10

    We use cookies.