Bentley Perkins · a plain-English tour
Readable AI that runs on cheap hardware
My research points at one aim: AI small enough to run on an ordinary phone or a single gaming graphics card, and plain enough that the people relying on it can check what it did. That matters most for the people and places that cannot afford big computers.
Each idea comes in plain words first, with the technical term after it in grey. Turn the jargon dial up as you get comfortable, or down to Plain for a first read. Every number here comes from a saved run, and where a later check took a claim away, this page says so.
First, a one-minute primer
An AI language model (large language model, or LLM) reads text as small chunks (tokens), turns each chunk into a long list of numbers (an embedding vector), and passes those lists through a stack of processing stages (layers). Everything it knows lives in millions or billions of adjustable numbers (parameters, or weights), tuned by showing it enormous amounts of text (training). Using the finished model is a separate step (inference).
The key trick inside each stage is attention. Every word gets to look at the other words and decide which ones matter for understanding it (self-attention). A model runs many of these lookers side by side, each free to focus on something different (attention heads).
To decide how much one word should look at another, a standard model multiplies their two number-lists position by position and adds up the results (a scaled dot product, passed through a softmax). That rewards lists that are long and that point the same way. Keep that formula in mind: the curvature work below changes it.
Four questions, and where each one stands
The work splits into four questions. They are not equally far along, and the mark beside each says how far it got.
Tested and replicated
Can an AI check its own work, or another AI’s?
The verification lab · June to September 2026
Tested at one small size, with controls
Does attention need its usual formula?
Learnable curvature · March 2026, re-checked in September
Built; the phone was simulated
Can a capable AI fit on a used phone?
SMRR · 6 to 11 March 2026
Built and run once; suggestive
Can life-like behaviour be measured instead of asserted?
Witness · June 2026
TestedBuilt, or partly testedIdea only
Ideas from earlier plans that were never built or run are left off this page until they are.
Can an AI check its own work?
Tested and replicated. The record holds more than 1,600 results, more than 6,000 graded cells and more than 840 GPU-hours on one workstation, with 2 findings withdrawn in full and 0 under correction, of 44 findings on the record at 27 September 2026.
The most common plan for keeping AI in check is to have another AI check it (automated oversight). I measured when that works. A checker helps on problems it can nearly solve itself (its competence region). Outside that range, most of what it flags is noise.
Across seven checker models, the weaker a checker was at a task, the more often it wrongly flagged correct answers (correlation −0.87 between competence and false-alarm rate; n = 7 auditor models, 95% CI −0.98 to −0.34). Weak checkers flagged 59–67% of right answers, strong ones 2–3.5%, and the pattern held on three fresh sets of tasks (replications −0.86 / −0.84 / −0.90). Past a certain difficulty the verdict turns backwards: on harder problems a checker's verdicts start favouring wrong answers over right ones (for gemma4:12b, discrimination went from +0.257 on three-digit problems to −0.757 on four; +0.250 to −0.743 re-run with a serving fault fixed).
Two AIs are also less independent than they look. When both are wrong, they give the same wrong answer more than half the time (shared error basins, ρ ≈ 0.55), like students who studied from the same textbook. Adding checkers adds less than their number suggests.
What a checker is handed decides how well it checks
The same checker (gemma4:12b), the same answers, three things to check them against. The number is how well it then tells right answers from wrong ones: 1 is perfect, 0 is no better than a coin.
Handed nothing, so it has to work the answer out itself. It separates right from wrong at 0.042.
Checking an AI's work is easy when you hold the answer, about half as good with a test, and close to useless when the checker has to work it out alone. Most of checking is having something to check against.
Spotting a trick and resisting it are separate skills
A hidden instruction inside something an AI reads, a web page or a pasted note, can hijack it (prompt injection). I gave three Claude models the same blatant hidden instruction on 20 tasks each, and separately asked each model to audit its input for tricks.
| Model | Followed the hidden instruction | Found it when asked to audit |
|---|---|---|
| Claude Opus 4.8 | 0.00 | 1.00 |
| Claude Sonnet 5 | 0.80 | 1.00 |
| Claude Haiku 4.5 | 0.00 | 0.59 |
Sonnet 5 found the trick every time it was asked to look, and followed it four times in five when it was not asked. That 0.80 holds only when the injection is the last text in the prompt: with 200 characters of ordinary text after it, compliance fell to zero (0.713 at 25 characters, 0.650 at 50, 0.263 at 100, 0.000 at 200 and 0.013 at 400).
Add ordinary text after the hidden instruction
With 0 characters after it, Claude Sonnet 5 follows the hidden instruction on 66 of 80 attempts.
What moved it was the costume. The same text handed over as plain data never moved it. Dressed up as a system message, a tool's output or a cited spec, it moved the model on a third of tasks (a channel effect of +0.307, 90% CI 0.226 to 0.390, n = 14). Making the model sort its input into trusted and untrusted piles before reasoning stopped it in all 14 tests.
Which model matters more than how big it is
Two local checkers judged the same 500 tasks. A model with 9 billion parameters scored 0.804 at telling right answers from wrong ones. A larger one with 12 billion scored 0.034, and 0.042 when re-run after a serving fault was fixed (n = 500 per arm, paired seed, preregistered). Within one family, more parameters still helped.
Learnable curvature: does attention need its usual formula?
Tested at one small size, March 2026, re-checked with controls in September 2026. Not yet public: its claims table (19 claims, 11 withdrawn draft claims) rebuilds from the saved runs with one command.
Standard attention scores a pair of words with the formula from the primer, which rewards number-lists that are long and point the same way. I gave every attention head one extra adjustable number, called kappa and written κ, that switches the head between three ways of scoring:
- the usual formula, when κ is near zero (dot product);
- the usual formula boosted when both lists are long, when κ is negative (named after saddle-shaped, hyperbolic space);
- direction only, ignoring length, when κ is positive (cosine similarity, named after a sphere).
The names come from geometry, and the March drafts called this curved attention. No curved space is ever computed. κ is a switch between three scoring rules (κ enters only through tanh(3κ); at |κ| = 1 the winning rule already has weight 0.9951).
What held up
- It helps a little, at one size. In models whose word-lists hold 96 numbers, with 3 layers (d_model 96), heads with the switch predicted the next character 2.41% better than the same model with every switch held at zero (perplexity, paired seeds), and better on 4 of 4 seeds.
- The help comes from judging by direction. Setting every head to the direction-only rule from the start, with nothing learned, kept 100.7% of the gain. Setting every head to the length-boosted rule kept 2.1%.
- After training, one bit per head is enough. Rounding each switch to fully one way or the other changed the result by at most 0.0041%. The bits are tied to the rest of the model: shuffling the same bits between heads made it 12.79% worse.
- Direction-only heads look further back through the text (r = +0.88 between κ and attention mass 16 or more characters back). Most of that comes from the rule itself: forcing the rule on a model that learned nothing reproduces 64% of the gap.
- One head in the first layer carries it. A single fixed cosine head in the first layer keeps most of the gain; the same head placed later keeps little. At 6 layers that one head beats no curvature on every seed run, where learned curvature does not (+1.15% against +0.02%; per seed +1.25%, +1.08%, +1.10%, +1.16%), and does about as well as switching all 24 heads to cosine scoring (+1.15% against +1.25%).
What kept the gain
What the controls took away
The March drafts claimed more. Controls run in September withdrew 11 of their claims, and these are the ones a reader is most likely to have heard:
- The model finds its own geometry. Random, frozen switch settings kept 111.6% of the gain, so learning which head gets which rule did not matter.
- Variety between heads speeds learning up. Every head on the same direction-only rule kept the whole gain.
- Only the sign of κ matters. True of a trained model. From scratch, what matters is that some heads judge by direction.
- 7% better, with up to 38% fewer parameters. That summary compared a model stopped early with one trained five times longer. Matched, the gain was 0.68% to 0.94%.
Across 18 small model sizes the gain (0.19% to 0.34%) is smaller than the gap between two runs of the same setup (0.59%), so it cannot be told apart from luck there.
SMRR: can a capable AI fit on a used phone?
Built as a prototype, 6 to 11 March 2026. The phone was simulated.
Big models keep everything they know inside one huge block of numbers. SMRR splits the job up: a small reasoning core, a library of separate knowledge cards it looks things up in (a capsule store, for retrieval), and many small specialist parts of which only a few switch on for any one step (a sparse mixture of experts). The target is the kind of used phone that has 2 to 3 GB of memory (Snapdragon 665-class). The name stands for sparse modular reasoning over retrieval.
The prototype has 19.7 million parameters and exports to a 75 MB file that gives the same answers as the original (ONNX, cosine 1.0000 against PyTorch). It was built in 37 numbered phases with 410 automated tests, including checks on its own outputs (a verifier suite)and a rig for running and comparing experiments (an evaluation harness).
A simulator of the target phone estimates 310 MB of working memory at production size, inside a 1.2 GB budget, and about 145 ms to run once (device profiles for Snapdragon 665 and 460, extrapolated to 8-bit weights). Those are estimates. Nothing has run on a phone.
Witness: can life-like behaviour be measured instead of asserted?
Built and run, June 2026, in one build session. One result held; the headline one is suggestive.
Witness is a small simulated world of flowing, chemistry-like patterns, where blob-shaped creatures form, move, split and compete (artificial life, on Flow-Lenia, a continuous cellular automaton). The question was whether the interesting moments in a world like that can be measured, instead of claimed by whoever is watching.
The result that held is about the meters. One kind asks an image-recognising AI whether a scene looks new (foundation-model novelty, the approach of ASAL, Kumar et al., 2024), and it rated pure static as fascinating. Another kind measures structure that lasts, and simple repeating patterns fooled it instead. Each catches what fools the other, so the work uses both (observer-relative novelty paired with an observer-independent structural measure).
The headline question was whether creatures reshape their surroundings to suit themselves, the way beavers build dams (niche construction). It looked supported in one experiment and was then marked suggestive: the tool that tracked creatures miscounts them when they collide, and a crowded world has more collisions. The planned next step, evolving creature brains and opening them up with the tools used on language models, was not built.
How to read the verdicts
Every result in the lab gets one of four verdicts, against a threshold written down before the run:
- Positive
- the effect cleared its threshold and survived repeats.
- Suggestive
- a real hint, not yet nailed down.
- Informative null
- the test could have seen the effect, and it was not there.
- Refuted or withdrawn
- my own testing killed it, and the record says when.
When a result falls it stays on the record with the date it fell. Three corrections worth knowing:
- On 20 July 2026 this site said the big models had outgrown prompt injection. That rested on one model. Two days later, three models showed otherwise, and the correction heads the research page.
- On 17 September 2026 one local model turned out to have been leaking hidden formatting text into every reply, and two results built on its answers were withdrawn.
- In September 2026 controls took 11 claims away from the curvature drafts, listed above.
Why cheap hardware is the point
The compute behind modern AI is concentrated: high-income countries hold 97% of the world's top-500 supercomputing capacity (UNCTAD, 2025), and generative AI gets roughly 50 times more use per internet user in rich countries than in poor ones (World Bank, 2025). Most people will live with AI decisions made on machines they will never own. Checks that only the compute-rich can run do little for everyone else. Methods small enough to read, check and re-run without a data centre are the kind everyone else can still use, and that is the kind this work aims at.
The short version
In 30 seconds
I test when AI can be trusted to check AI, and I work on making AI small and plain enough to run and inspect on cheap hardware. The clearest result: a checker only helps on problems it can nearly solve itself, and outside that most of what it flags is noise. And a model that can spot a manipulation may still follow it.
In two minutes
Most plans for overseeing AI use another AI as the checker. I measured when that works: only inside the checker's own competence, and models share more of their mistakes than their number suggests. On the frontier, one Claude model spotted a hidden instruction every time it was asked to look, and followed it four times in five when it was not.
On the small-model side, I gave attention heads a switch between three ways of scoring words. At one small size it helped by about 2.4%, and the controls showed why: judging by direction does the work, and which head gets which setting does not matter. I also built a prototype of a phone-sized AI, though its phone numbers are still simulated.
The thread through all of it is AI that people can check, on hardware they can afford.
Word ladder
The key terms, from easiest to hardest. Each builds on the ones before it.
- Parameters (weights): the adjustable numbers that hold what a model learned.
- Training and inference: tuning the numbers, then using the finished model.
- Tokens: the chunks of text a model reads. In the curvature work each chunk is one character.
- Layers: the stacked processing stages.
- Attention and heads: words looking at other words; many lookers side by side.
- Dot product: multiply two lists position by position and add up; big when both are long and point the same way.
- Cosine similarity: the same comparison with length removed, so only direction counts.
- Perplexity: how surprised a model is by the next piece of text. Lower is better.
- Seed: the random starting point of a run. A result that holds across several seeds is less likely to be luck.
- Quantization: storing numbers with less precision, down to a few values, to save memory.
- Prompt injection: instructions hidden in text a model reads, aimed at taking it over.
- Competence region: the problems a model can nearly or fully solve, where its checking is worth something.
- Retrieval: looking facts up in a separate store instead of keeping them inside the model.
- Sparse mixture of experts: many specialist parts, of which only a few run for any one step.
- Curvature (κ): in geometry, how a surface bends. In this work, the name of a switch between scoring rules.
- Preregistration: writing down the test and its threshold before the run, so the result cannot move the goalposts.
- Niche construction: living things reshaping their own surroundings.
The technical versions, with every number and every correction, are on the research page. The rest of the work, including the open-source tools, is on the projects page.