Skip to main content

Lab Notebook · Newtraining — Non-Transformer Line

Newtraining — an honest plateau, then a real break, hunting for something not a transformer

An explicit bet that transformers are pattern matchers, not minds: byte-level field dynamics, dense-associative memory, no backprop, no tokenisation, no attention. A hard-kill gate was set before the run started. The result was a flat plateau at 0.155 held-out byte-accuracy — broken, via bigram conditioning, to 0.248 (peak 0.375). It did not clear the hard-kill bar. We say so, and explain what the plateau turned out to be.

JR
Jon RossFounder, vocabotics — 15 years building safety-critical systemsLab report · dated 24 April 2026
Verified by a human. Drafted with AI, verified by a human. Jon Ross, 24 Apr 2026
Living document. Reviewed 24 Apr 2026
Entry date
24 April 2026
Category
Models
Access
🔓 Public
ResearchDisproven

A hard-kill bar was named before the run started. The result didn't clear it — and that's exactly the kind of result this lab is built to publish.

0.155
flat-matcher plateau, held-out byte-accuracy
0.248 (peak 0.375)
after bigram conditioning broke the plateau
6.56 → 5.37
bits-per-byte compression
D6 >= 0.35
the hard-kill gate — not cleared

Honest evaluation

Disproven

The pre-registered hard-kill gate (D6 >= 0.35) was not cleared; the readout-capacity-ceiling diagnosis is a real, useful finding, not a substitute pass.

What would prove or disprove it further

What would prove or disprove the diagnosis further: build a stronger readout mechanism on top of the same substrate (the field-dynamics/dense-associative-memory core, unchanged) and re-run against the same D6 >= 0.35 gate. If a better readout clears the gate, the readout-capacity diagnosis is confirmed and the substrate itself is vindicated; if it still misses, the ceiling was likely in the substrate after all, and the diagnosis here would need to be revised.

The evidence — full reasoning behind the verdict

Verdict: disproven. The hard-kill gate — D6 must reach at least 0.35 — was set before the run started specifically so there would be no room to move the bar after seeing the result. It didn't reach 0.35. By its own pre-registered bar, this experiment did not solve anything, and that's the plain reading, not a softened one.

The run measured a flat plateau at 0.155 held-out byte-accuracy, broken via bigram conditioning to 0.248 (peaking at 0.375), with bits-per-byte compression improving from 6.56 to 5.37. That the plateau moved at all, and moved specifically when bigram conditioning was added, points to a credible diagnosis: the ceiling sat in how the representation was read back out, not in whether the representation itself could hold structure. A real, useful finding — but a diagnosis of why it fell short, not a pass on the bar it was measured against.

Set the goalposts before the run, not after. That's the whole discipline behind this experiment — and the honest headline is that it didn't clear them.

What it was

Newtraining is an explicit research bet that "transformers are pattern matchers, not minds" — a line pursuing byte-level field dynamics and morphogenetic, dense-associative memory, deliberately avoiding backpropagation, tokenisation, and attention altogether. The mandate, stated up front: categorically different, not ten percent higher on a benchmark. A hard-kill gate — D6 must reach at least 0.35 — was named before the experiment began, so there would be no room to move the bar after seeing the result.

What we built

Research

Three days of focused experimentation, 23-25 April 2026, across 437 files and 137MB. The first measured result was a flat plateau: 0.155 held-out byte-accuracy, stuck. Rather than abandon the run, the team tried bigram conditioning — and the plateau broke, moving accuracy to 0.248 (peaking at 0.375) and compressing bits-per-byte from 6.56 down to 5.37.

What we learned — including the honest negative

The hard-kill gate was not cleared. D6 needed to reach 0.35; it didn't. Read that plainly: this experiment did not solve anything, and no result here should be read as a working non-transformer model.

What it did do is locate something real: the plateau turned out to be a readout-capacity ceiling, not a ceiling in the underlying substrate. That's a meaningfully different diagnosis than "the approach doesn't work" — it says the representation itself was learning something, and the bottleneck was in how that representation got read back out, not in whether it could hold structure at all. Bigram conditioning breaking the plateau supports that reading directly.

This is exactly the kind of result this lab exists to publish honestly: a negative against the stated bar, with a specific, credible account of why, rather than either burying the miss or dressing it up as a win.

Where it went / status

Closed as a time-boxed experiment against its own hard-kill gate, and archived as a negative-but-informative result. The readout-capacity diagnosis is the useful thing to carry forward into the wider non-transformer architecture line — it says where to look next, even though this particular run didn't clear its own bar.

What is still open — kept visible

The honest edges, next to the wins. This is what turns 🔬 into 🟢 — honestly.

  • The hard-kill gate (D6 >= 0.35) was not cleared — this did not solve anything, and we say so plainly rather than reframing a miss as a win.
  • The plateau turned out to be a readout-capacity ceiling, not a substrate ceiling — a real finding, but a diagnosis, not a fix.
  • 437 files, 137MB, active over three days (23-25 April 2026) — a short, sharply time-boxed experiment, not a mature research programme yet.

Where this connects

Sources

  1. vocabotics project audit — Newtraining / non-transformer, non-backprop experiment (hard-kill gate D6, 23-25 Apr 2026)vocabotics internal project history · as of April 2026

    We use cookies.