Skip to content
Bentley PerkinsAn independent commonsFuturism Institute
Bentley Perkins: shoulder-length fair hair with sunglasses pushed up into it, a grey knitted scarf, and a red and blue patterned shirt, standing in a room with graffiti-covered walls and ceiling.

Bentley Perkins · independent researcher and builder

Welcome!

I’m developing a set of protocols to unify resilient infrastructure standards online and beyond. I’m looking to work on the development of sustainable land projects and film the journey of established community models into a documentary about our shared future.

Beneath our polycrisis, we have a communication and consensus challenge, and therefore a platform design challenge. The present focus on developing powerful AI minds leads us to a power and citizenship enigma that requires our voices to be adequately formed and heard.

If we assume catastrophic scenarios of AI development, then our authorship of our reality and the AI that influences it could disappear. Now is the opportunity to leverage global networking and translating technology to experiment and refine robust dialogue, democracy and cooperation. Within a rigorous system, global priorities and values can emerge. Defining global citizenship and what civic participation looks like in that paradigm is the first question it can start to answer.

I’m approaching this challenge by developing a digital ecosystem of literacy and connectedness, decentralization where possible and connecting the unified approach through living spaces of culture that I then record into the documentary.

Seattle and remote · two founding-client engagements open, as of August 2026

Build a workflow with meRead the research

Research

Legible-AI lab

The finding

Self-correction is competence-gated. Verification fails where it's needed most.

Most plans for overseeing AI lean on one move: use another AI to check the first one. I measured when that move works. A checker helps inside its own competence region, and outside it, its objections are mostly noise. The models also turn out to be less independent than they look. When two of them are wrong, they hand you the same wrong answer more than half the time, like students who studied from the same textbook. I derived this from a minimal model first, then confirmed it on seven real ones, then replicated it on fresh tasks.

Read the full result →New here? The plain-English tour →
Instrument · the competence gatemodel
0501000%50%100%objections mostly noisemostly realthe gatechecker competence on this task →share of its objections that are real

At 38 competence, 4 of 20 objections are real. most of what it flags here is noise.

An illustrative model of the measured pattern, not the data: the measured curves live on the research page.

−0.87correlation between an auditor's competence and its false-alarm rate on correct work, low-competence auditors flag 59–67% of right answers vs 2–3.5% for competent ones (63 cells, seven models; replicated −0.86 / −0.84 / −0.90)
ρ ≈ 0.55how often "independent" models give the same wrong answer when both are wrong, a verifier's catch rate tracks the derived law 1 − (1 − q)·ρ at corr 0.83
0.0 → 0.69held-out accuracy on a task best-of-N sampling never solved, self-repair at a matched call budget, replicated; it works only where the model has a foothold, and can regress where it doesn’t

What this covers: the core auditor and repair results are local to open-weight models and programmatically graded tasks. The frontier runs cover three Anthropic models on one injection probe, so they are single-vendor and small (20 tasks each, 14 for the authority test). I published a wrong conclusion from those runs on July 20 and corrected it on July 22; the correction is at the top of the research page. Nothing here tested cross-vendor auditing, shared basins, or real software-engineering work. I claim the measurements, not more than that. The rigor standard behind them has refuted seven of my own headline hypotheses in writing, so the surviving results have earned some trust. Methods, result cards, and the full writeup are on the research page, and I'm glad to share the underlying evidence directly.

Selected work

Research and the tools it produced

The tool repositories, governor and lattice-commit, are public. Research code, data, and reproduction bundles are available to reviewers and collaborators on request.

Now Model identity overwhelms parameter count in the paired verifier test.

Work with me

Open to research collaboration, funding conversations, and reviewer access to the full evidence.

Reviewers looking at the civic half: the grant brief for Values Commons, with its scope, ask, and six-month plan, is at valuescommons.org/funders.

futurisminstitute@gmail.com