Bentley Perkins · collaboration
Where this work needs other people
The mission: AI that people can check, on hardware they can afford, and the civic tools and communities that make that matter.
I do most of this alone, so most of its blind spots are mine, and I cannot see them from here. Below is the whole range of the work, then four places I know it is stuck. I suspect the more useful help is somewhere I have not thought to look, and I would like to hear about that too.
Message Bentley
I read and answer every note. It opens in your own email app, and this page keeps nothing.
The range of the work
Six areas, 26 projects, one line from inside an AI model out to a piece of land. The research asks when AI can be trusted and how to make it small; the civic work is what that is for.
Research when AI can be trusted, and how to make it small
Checking AI
I measure, on one workstation, when one AI can check another. Each test is written down before it runs, and every correction stays on the record.
- The competence gate
- frontier-repro
- Boundary Intelligence
- Local verification engine
- Governed scale-up program
- Shoofly
Needs someone: Check my numbers, The experiment one workstation cannot run
Small, affordable AI
Models and tools small enough to run on a phone or a single graphics card, so the people relying on them can run them and check them.
Mathematics and science
Tools that say how far a result can be trusted, tried on a hard open problem in mathematics and on real systems: energy, and artificial life.
Civic work the records, groups and places the research is for
Shared records
Facts that carry their sources, on an open standard anyone can run, with no ads, accounts or tracking.
Needs someone: Run the civic tools where you live
Groups
Plain files a group keeps for itself: what it agreed, who is doing it, and what would change its mind.
Needs someone: Run the civic tools where you live
Places, land and film
Communities that already live well together, mapped from their own sources and filmed for a documentary about our shared future.
Needs someone: Land, community and film
The civic projects meet at the Futurism Institute. The research is argued for a general reader at Legible AI and kept in full, with every correction, on the research page.
Four openings
Check my numbers
The raw completions behind the frontier results are public in frontier-repro. Its scripts need Python and nothing else: no API key and no network. If a number does not reproduce, an issue on that repository is the fastest way to say so. For the local results, the result cards, preregistration templates and the full writeup go to anyone reviewing the work.
I have been wrong in public before: the research page lists nine corrections, each with what caught it. If you find the tenth, that is the most useful thing you could send.
Run the experiment one workstation cannot
The open question at the centre of the lab: do AI checkers from different companies share blind spots? Oversight schemes that stack several checkers assume they do not, and nobody here has measured it. It needs access to frontier models from more than one company and a second person on the grading. If you have either, write and say which.
And if you think it is the wrong question, I would rather hear that before the experiment than after it.
Run the civic tools where you live
Values Commons ranks everyday things by your values, on facts that carry their sources, with no ads and no tracking. Underneath it is an open standard, so a community can run its own version without asking anyone. The Futurism Institute holds the rest of that work.
I do not know your place. What broke when you tried it, and what does your community already use to decide things together?
Land, community and film
I want to work on sustainable land projects and film established community models for a documentary about our shared future. If your community has been living one of those models for years and would want it recorded, or you are starting a land project, I would like to hear what you are building.

Living Energy Farm, Virginia, from the draft film. You would know your place better than any film could, and I would want you shaping how it is told.
Four shapes I will not build at any price
With anyone, for any reason.
No amount of review bolted on afterwards makes these safe, because the review is the part they remove.
- Systems that modify or extend themselves
- No self-editing prompts, self-rewriting tools, or a system that changes its own instructions between runs. Every version a system runs is one a person approved.
- AI that builds or configures other AI unattended
- A model may draft a config or a script. A person reads it and installs it. Nothing generates a running system and puts it into service without that step.
- Improvement loops with no human in the cycle
- Evaluate, revise, redeploy, repeat is the shape I refuse most firmly. It removes the only reviewer at the exact point the system starts changing fastest.
- Agents that spawn or direct other agents
- Orchestration that a person cannot read as a single flow is orchestration nobody can audit. One flow, one owner, one place it stops.
This follows from what the lab measured: a checker cannot verify past its own competence, so a system cannot safely improve past it either. Anyone selling you a self-improving loop is selling you the part where nobody is watching.
What you might know that I don’t
Some of what this work needs, I cannot ask for by name, because I do not know it exists. These are questions I cannot answer alone.
If you have tried something like this beforeWhat happened? The history of attempts like this that did not work is what I most need and can least find.
- If you evaluate AI systems for a livingWhich of my methods would you not trust, and what would you do instead?
- If you organise where you liveWhat would a tool have to respect to be of any use there?
- If you work on law or policyWhich of these findings would change anything you do, and which would not?
- If you are young and curious about any of itWhat did not make sense? A page only experts can read has failed at half its job.
And if none of these is you: what can you see from where you stand that I cannot see from here?
Write to me

Tell me what you are working on, what you know that this work is missing, or where you think I am wrong. Say which opening you mean, or none of them.
Bentley PerkinsSeattle
futurisminstitute@gmail.comRead the research first
Not one of the four, and just curious? Send a note instead.