Sixteen hours into a 33-hour experiment, my dog Sylvie pulled the power cable out of my PC. My first thought was that the whole thing was ruined. It wasn’t, and the experiment’s answer turned out more interesting than the one I’d bet on.

Sylvie, a grey long-haired dog, asleep and sprawled half off her dog bed

Sylvie, unrepentant.

The question

Most language models read text through a tokenizer: a fixed dictionary of word pieces chosen before training starts. A newer family of models skips the tokenizer. They read raw bytes and learn where to group them into chunks. The one I used is based on H-Net (Hwang, Wang and Gu, 2025; code), which does this in stages: bytes into small chunks, small chunks into bigger ones. I didn’t use their code. I wrote a small version from scratch in PyTorch, with ordinary Transformer layers instead of their Mamba-2 ones and about 34 million parameters instead of their 1.3 billion, so treat my results as being about my model, not theirs.

Code seemed like the perfect place to test that. Python has a real grammar with tokens, expressions, statements and blocks, and a parser can tell you exactly where each one starts. So:

  1. Does a byte model trained on Python rediscover the grammar? Do its learned chunk boundaries line up with statements, beyond what a trivial rule like “cut at every new line” already gets?
  2. If you nudge it toward the grammar, does it train more efficiently? Same compute budget, better predictions?

Writing the answer key first

Before training anything, I wrote down the hypotheses, the exact pass/fail thresholds, and my own guesses. Then I committed that file to git. It’s called pre-registration, and it’s standard in medicine and psychology. It’s rare in hobby ML, where it’s very easy to run an experiment, look at the numbers, and then decide what counts as success.

The bar for the first question was concrete: “cut at every new line” matches Python statements with a score of 0.80. To count as rediscovering grammar, the model had to get close to that. Under 0.20 would be a clear no.

My guesses, written down on September 22:

  • Question 1: 40% no. I was skeptical the model would find statements on its own.
  • Question 2: 50% yes. Code is far more structured than prose, so handing the model the grammar felt like it should help.

The setup

All of this ran on one RTX 3080 Ti in my basement:

  • Model: about 34 million parameters, trained from scratch on raw bytes of Python.
  • Data: 2.65 GB of Python from open-source projects, split by project so forks of the same repo can’t land in both training and test.
  • Conditions: five ways of choosing chunk boundaries. The model learns them (one stage or two); the model learns them with a gentle nudge toward Python’s real token boundaries; token boundaries from a parser are imposed; or boundaries from a dumb regex that cuts wherever the character type changes are imposed.
  • Fairness: every run got exactly the same compute budget, about 3 hours each. The learned conditions ran with three different random seeds, so one lucky run can’t carry a conclusion.
  • Total: 11 runs, about 33 GPU-hours.

Sylvie

The queue ran overnight. By the next morning, five runs had finished: one seed of every condition. Sixteen hours in, 38 minutes into run six, the power cable came out.

Here’s what I’d built in beforehand without really thinking about Sylvie:

  • Every finished run saved a checkpoint and its results before the next one started. I had Claude Code check the damage: all five checkpoints loaded cleanly. Nothing that had finished was lost.
  • The queue was resumable. It skips any run with a saved checkpoint, so restarting it picks up at the first unfinished run.
  • What I lost: the 38 minutes of run six. It only saves at the end, so that run had to start over. The partial run got moved aside so it couldn’t mix into the results.

Restart, 20 more hours, and 11 out of 11 runs finished.

The lesson isn’t clever: save your work in pieces, and make restarting boring. Most of the time that’s for crashes. This time it was for Sylvie.

What the model learned

Before the verdict, here’s what the trained model actually does. I fed it a small Python function and recorded where each stage cut. | is a first-stage chunk, ‖ a second-stage chunk:

‖def |area(|r):
|    |if |r |< |0:
|        |raise ‖ValueError(|"neg"|)
|    |return |3.14 |* ‖r |* |r

Nobody told it what a Python token is. It saw bytes, nothing else. Its first stage still found def , area(, ValueError(, the whole string "neg", and 3.14 as one number. That’s roughly word-level, which is about what I expected.

The second stage is the interesting one. It cuts rarely, and not at statements.

The verdicts

Question 1, does it rediscover grammar? No. The second stage matched statements with a score of 0.089, against 0.80 for the new-line rule, and all three seeds landed clearly in the “no” band. It isn’t a budget problem: the second stage cut at nearly the same rate as statements occur, so it could have lined up with them. It put its boundaries somewhere else, and that didn’t change at any point during training. Whatever it learned to group, it isn’t Python’s grammar.

Question 2, does nudging it toward grammar help? No, and if anything it hurt. The difference fell inside the “no effect” range I’d set in advance, so that’s the official answer. But the nudged model was a little worse in all three seeds, by about five times the seed-to-seed noise. It’s small, and consistent.

The surprise. The best result of all five conditions came from the dumb regex: cutting wherever the character type changes, imposed as a fixed rule. It beat the learned model and the parser-based boundaries. It’s only one seed, so it’s a lead rather than a finding. My best guess at why: in actual code, the regex and the parser agree 99% of the time. They differ inside strings and comments. The parser treats a whole docstring as one token, a single enormous chunk, while the regex cuts English prose into words. On code with lots of docstrings, that probably matters.

How my guesses did

  my guess result
Doesn’t rediscover grammar 40% ✔
Nudging toward grammar helps 50% ✘

One right, one wrong. The one I got wrong is the one I’d have argued for hardest. This is why the guesses go in writing before the data exists: otherwise I’d remember being “pretty sure it might not help.”

What it means (and doesn’t)

This is a 34-million-parameter model, a tiny fraction of the size of the models these ideas are aimed at. A “no” at this scale says nothing about a model a thousand times larger. What it does say: at small scale, on code, learned chunking didn’t find grammar, gentle guidance didn’t help, and a simple fixed rule did best. That fits a pattern that keeps showing up in ML: earlier work found much the same on English text, where given boundaries did about as well as learned ones.

Next I’ll run the regex condition with more seeds to see whether its lead is real, and dig into what the second stage is grouping, since it clearly learned something. Sylvie will be supervised.

Questions, or a guess about what that second stage is doing? I’m @roho.foo on Bluesky, or email nick_notes@icloud.com.

There’s also a short anime version, with a closing song.