On June 29 my research project passed six milestones in one day. Every metric was green. The README said “thesis validated.”

Two days later I wrote a retraction into that same README, and ran the test that should have come first. It killed the hypothesis.

This post is about both halves: the idea, how I fooled myself, and what I found once I stopped.

The idea

Big language models keep what they know in their weights: billions of numbers tuned during training. To know more, you train a bigger model. That’s expensive, and you can’t open the model up to fix one wrong fact.

Brains seem to split things differently. Your genome is far too small to spell out every connection in your brain. It carries priors and learning rules, and the rest gets built from experience. (That argument is from a 2019 paper by Anthony Zador.)

So the bet was to split the jobs:

  • Facts go in a knowledge graph, a database of simple statements like “KeyError is a kind of LookupError.” You can read it, edit it and check it.
  • Reasoning goes to a logic solver, a program that chains facts together exactly. If A is a kind of B and B is a kind of C, then A is a kind of C, every time, for free.
  • The language model stays small and only works at the edges: turning text into facts, and breaking ties the solver can’t settle.

If that worked, a small model plus structure should match a much bigger model on its own. That was the claim.

The one before

This wasn’t my first try. The project before it was called Tilda, and it was a mess. I was inventing my own words for what I was doing, so I couldn’t tell when something already had a name, and I kept reinventing the wheel. I also kept adding variables, so many that there was never going to be a clean result at the end. So the new project got rules:

  1. Standard terms only. If the field already has a word for something, use that word. Then you can look it up and find out that someone solved it decades ago.
  2. Write the pass/fail bar before you run anything. Every milestone gets a metric and a “kill” threshold, committed to git before the data exists. That forces each experiment to ask one question, not twenty. It’s called pre-registration, and I wrote about it with my dog’s help earlier this week.

Six milestones in a day

I built it with Claude Code, one milestone at a time: algorithms and data structures first, then a codebase, then my homelab’s live config, then computer science as a field, then questions where sources actually disagree. Each domain was harder to check than the last. Each one passed.

Except one.

The failure I’m still proud of

Milestone five asked a question I care about a lot: can the model tell when it shouldn’t answer? Some questions are settled. Some are genuinely contested among experts. A trustworthy system should answer the first kind and say “this is disputed” on the second.

The first version abstained on 62% of the settled questions, which was over my kill bar. I found a real bug (the models were bad at double negatives in my probe question), fixed it, and re-ran on a fresh set of questions, so I couldn’t just tune until it passed. Over-abstaining dropped to 38%. But now the model only caught 17% of the contested questions.

The finding: small models aren’t uncertain about contested questions. They’re confidently opinionated. You can’t get “this is disputed” by asking the model how sure it is. You need outside sources that disagree with each other.

That one went into the record as a failure, and it’s the most useful thing from the whole first round.

Where I fooled myself

The other milestones were green because I’d written the tests, the right answers and the planted errors myself. The solver was checking whether it agreed with facts I’d typed in. Of course it did.

Two things were missing:

  • Outside ground truth. The grader and the author were the same person.
  • A baseline. I never asked whether a plain large model, with no solver at all, would do just as well.

Around then, Anthropic’s Fable model came back after a few weeks offline. I’d barely gotten to use it the first time, because I was nearly out of usage, so I rushed to throw the project at it before it could disappear again. I asked for a blunt review. It found both problems, and I couldn’t argue with either one. For about an hour I seriously considered a career in retail.

So the README got a banner saying these milestones tested the plumbing, not the idea. Then I designed the real test.

Doing it properly

I needed questions where the right answer is computed, not typed by me. Python happens to have one built in: issubclass. Is KeyError a kind of Exception? Python will tell you, and it’s never wrong.

I took 127 classes from Python’s standard library (exceptions, number types, collections, file types) and built every “is A a kind of B” question with its true answer. Some are one hop (KeyError → LookupError). Some take six.

The real hypothesis, written down before anything ran: a plain model gets worse as the chain gets longer, while the solver doesn’t. If the gap grows with depth, the solver earns its place.

The models were Qwen 2.5 at five sizes, from 0.5 billion to 14 billion parameters, all running locally.

What the small models did

The first pilot turned up something funny. Each question was multiple choice, and I shuffled which letter meant “true” on every question. The smallest model answered “A” 96 times out of 96. Without the shuffle, that strategy would have looked like a decent score.

The 1.5B and 3B models were worse in a more interesting way: they denied almost everything. The 3B model said a tuple is not an object. In Python, everything is an object. It wasn’t guessing; it had a bias toward “no.” That’s a scary property in any system that uses small models to double-check things.

At 7B the models suddenly became competent. It was a cliff, not a slope.

Can the model write down the facts?

The plan needed the small model to build the knowledge graph: “list the direct parent classes of X.” Checking a fact is one thing. Writing it down exactly is another.

It couldn’t. The best model, 14B, got 60% of the parents. Half of what it wrote was true but wrong: asked for IndexError’s parent, it said Exception, skipping LookupError in between. It knows roughly where things go in the tree, but it doesn’t keep the exact steps. It even flipped a few upside down.

That meant the small model can’t supply the solver’s facts. Half of the plan was already dead.

The main test

400 questions, 100 per chain length, three ways of answering:

  • Model alone: just the question.
  • Model with facts: the question, plus the true facts it needs, padded out to 30 facts with true ones it doesn’t need.
  • Solver: the same facts, chained by the logic solver. No model at all.
  1 hop 2 hops 3 hops 4+ hops
Solver 1.00 1.00 1.00 1.00
14B alone 0.85 0.85 0.87 0.89
7B with facts 0.97 0.94 0.91 0.96

The 14B model didn’t get worse with longer chains. If anything it got slightly better. My pilot, with 24 questions per cell, had shown it declining. At 100 per cell, that was just noise.

And the model with facts nearly matched the solver. Once you hand a model the right facts, it chains them fine.

Hypothesis falsified. The kill bar I’d written down fired, and it stayed fired.

How my guesses did

The guesses went into the pre-registration files before each run:

  my guess result
The model can write down exact facts 95% ✘
The 14B model gets worse with longer chains 85% ✘
The hypothesis holds 60% ✘
The model with facts keeps up with the solver 40% ✔

Only one right, and it was the one I thought was less likely. I was sure about the other three.

What survived

The failure split the idea into pieces, and the pieces came out differently:

  • Keeping facts outside the model: supported. The 7B model with facts beat the 14B model alone at every chain length, in less than half the time.
  • The solver as the thinker: not needed at this scale. Given the facts, the model reasons nearly as well.
  • The model as the fact-writer: unreliable. It knows the shape of things, not the exact steps.

The solver isn’t smarter than the model. It’s free and exact: zero tokens, the same answer every time, and it can do some things in-context reasoning can’t, like spotting three claims from three sources that are each fine alone but contradict each other together.

None of the “facts outside the model” half is new. Retrieval-augmented models have been showing it for years. What I got was a clean, controlled look at which half of my idea was doing the work. That’s less exciting than “thesis validated,” but it’s true.

What I’d tell myself on June 29

If everything passes on the first day, be suspicious. Ask who wrote the answer key. Ask what the boring alternative would score. And write your guesses down, because you’ll be wrong about the thing you’re surest of.

The next project, a router that picks the right tool for each question, came straight out of this one. It has outside ground truth from day one.

The code is public: neurosymbolic-kb on GitHub. The full write-up has all the numbers, and every model answer is committed, so the tables rebuild without running a single model. Questions, or a domain where you think the hypothesis would survive? I’m @roho.foo on Bluesky, or email nick_notes@icloud.com.