Boundary Intelligence
A multi-project program (986+ pre-registered trials) finding that simple mechanisms at system boundaries can match or beat far larger internal-model approaches. The report includes five negative results.
Synthesis · 986+ trials
Bentley Perkins · independent researcher and builder
I’m developing a set of protocols to unify resilient infrastructure standards online and beyond. I’m looking to work on the development of sustainable land projects and film the journey of established community models into a documentary about our shared future.
Beneath our polycrisis, we have a communication and consensus challenge, and therefore a platform design challenge. The present focus on developing powerful AI minds leads us to a power and citizenship enigma that requires our voices to be adequately formed and heard.
If we assume catastrophic scenarios of AI development, then our authorship of our reality and the AI that influences it could disappear. Now is the opportunity to leverage global networking and translating technology to experiment and refine robust dialogue, democracy and cooperation. Within a rigorous system, global priorities and values can emerge. Defining global citizenship and what civic participation looks like in that paradigm is the first question it can start to answer.
I’m approaching this challenge by developing a digital ecosystem of literacy and connectedness, decentralization where possible and connecting the unified approach through living spaces of culture that I then record into the documentary.
Seattle and remote · two founding-client engagements open, as of August 2026
The finding
Most plans for overseeing AI lean on one move: use another AI to check the first one. I measured when that move works. A checker helps inside its own competence region, and outside it, its objections are mostly noise. The models also turn out to be less independent than they look. When two of them are wrong, they hand you the same wrong answer more than half the time, like students who studied from the same textbook. I derived this from a minimal model first, then confirmed it on seven real ones, then replicated it on fresh tasks.
At 38 competence, 4 of 20 objections are real. most of what it flags here is noise.
An illustrative model of the measured pattern, not the data: the measured curves live on the research page.
What this covers: the core auditor and repair results are local to open-weight models and programmatically graded tasks. The frontier runs cover three Anthropic models on one injection probe, so they are single-vendor and small (20 tasks each, 14 for the authority test). I published a wrong conclusion from those runs on July 20 and corrected it on July 22; the correction is at the top of the research page. Nothing here tested cross-vendor auditing, shared basins, or real software-engineering work. I claim the measurements, not more than that. The rigor standard behind them has refuted seven of my own headline hypotheses in writing, so the surviving results have earned some trust. Methods, result cards, and the full writeup are on the research page, and I'm glad to share the underlying evidence directly.
Selected work
A multi-project program (986+ pre-registered trials) finding that simple mechanisms at system boundaries can match or beat far larger internal-model approaches. The report includes five negative results.
Synthesis · 986+ trialsA completely local confidence and verification layer: behavioral agreement, risk-aware escalation, independent checks, and an honest refusal when the available stack cannot establish an answer.
Private alpha · external receipt openAn open programme building failure-aware tools for AI-assisted mathematics, tested against Robin’s inequality. Every result is labelled with the evidence ceiling it was actually established at, and the claims this project has withdrawn stay on the page. RH remains open.
RH open · results labelled by ceilingA local, OpenAI-compatible execution gateway (~13,300 LOC, 133 tests): tiered model routing, async task management, and a containment validator that caught 100% of adversarial proposals in testing.
Working softwareA git-native, test-gated wrapper for LLM code-repair loops. Restores reliable multi-file repair where a bare loop silently reverts good fixes. Packaged and benchmarked on 22 workspaces.
Alpha · PyPIA keyword classifier at the request boundary that routes each turn to the cheapest adequate model tier, 23–65% cost reduction at 95–96% quality in measured sessions.
Open source · MITThe tool repositories, governor and lattice-commit, are public. Research code, data, and reproduction bundles are available to reviewers and collaborators on request.
Now Model identity overwhelms parameter count in the paired verifier test.
Open to research collaboration, funding conversations, and reviewer access to the full evidence.
Reviewers looking at the civic half: the grant brief for Values Commons, with its scope, ask, and six-month plan, is at valuescommons.org/funders.
futurisminstitute@gmail.com