I help small teams turn repetitive research, proposal, intake and reporting work into controlled systems, with evidence, permissions and human approval built in. Then I write down what the system cannot do, and hand you the instructions for turning it off.
Seattle and remote · two founding-client engagements open, as of August 2026
One repetitive process becomes a controlled system. The main offer, and the one to read first.
A map of the workflow as it runs today, before anything is automated
One bounded AI-assisted system for that single process
Up to two integrations with tools you already pay for
Human approval required before any consequential action
Ten to twenty evaluation cases built from your real examples
A record of every input, action, output and failure
A staff walkthrough and a written handoff
A known-limitations sheet, and shutdown instructions
50% to begin, 50% at handoff. Third-party API and software costs are billed to you directly, not marked up.
Agent Reliability Audit
$4,500 founding rate
For a team already running an AI system that nobody has tried hard to break.
Where the system states things its evidence does not support
Prompt-injection paths through any untrusted input it reads
Tool permissions wider than the task requires
What happens on failure, and whether anyone is told
Whether your checkers are as independent as their number implies
A prioritised fix list, and one safeguard implemented
Fixed price. Findings are yours; nothing is published without your written agreement.
AI Workbench Clinic
$300 per session
One person, one working session, one system they use every week made materially better.
A working session on your actual files and your actual process
A workflow map you keep
Configured templates and instructions
One follow-up adjustment within two weeks
Paid in advance. The most common route in, and often how an organisation finds the workflow worth a sprint.
Managed Stewardship
$500 to $1,000 per month
After a system is installed: watching it, correcting it, and telling you when it drifts.
Monitoring of the evaluation set against live behaviour
A monthly report with what changed and what degraded
A fixed allowance of improvement work
Model and cost routing kept current as prices move
Only offered after a sprint or audit. There is nothing to steward before that.
How a sprint actually runs
Map the workWhat happens now, who touches it, where it breaks, and what a good result looks like.
Establish the baselineHow long it takes and how often it goes wrong today. Without this, no improvement can be claimed.
Build the smallest useful systemOne process. Not a platform. Small enough to finish and to understand.
Test against real casesYour examples, not invented ones. Including the awkward ones you would rather not send.
Train, document, measureYour team runs it. The limitations are written down. The baseline is measured again.
Stage two is the one most engagements skip, and skipping it is why so many AI projects cannot say whether they worked. If we do not measure the before, there is no after to compare it to.
What I will not build
Named up front, because a boundary that only appears once you have paid is not a boundary. I do not install systems that take these actions without a person in the loop:
Moving money, or any financial transfer
Final legal, medical, or clinical determinations
Hiring, firing, or performance decisions about a person
Destructive operations against a production database
Sending mass communication that no person reviewed
Anything where I cannot see how it would be checked
And four shapes I will not build at any price
The list above is about actions a system must not take on its own. These are different: they are system shapes, and no amount of review bolted on afterwards makes them safe, because the review is the part they remove.
Systems that modify or extend themselves
No self-editing prompts, self-rewriting tools, or a system that changes its own instructions between runs. Every version a system runs is one a person approved.
AI that builds or configures other AI unattended
A model may draft a config or a script. A person reads it and installs it. Nothing generates a running system and puts it into service without that step.
Improvement loops with no human in the cycle
Evaluate, revise, redeploy, repeat is the shape I refuse most firmly. It removes the only reviewer at the exact point the system starts changing fastest.
Agents that spawn or direct other agents
Orchestration that a person cannot read as a single flow is orchestration nobody can audit. One flow, one owner, one place it stops.
This is not caution borrowed from a policy document. It follows from the finding this practice is built on: a checker cannot verify past its own competence, so a system cannot safely improve past it either. Anyone selling you a self-improving loop is selling you the part where nobody is watching. If that is the engagement you want, I am the wrong person, and I would rather say so on the price page than in week two.
There is also a limit on me. I have not run a client engagement of this kind before: the research programme behind it is five months old and has published seven of its own refuted hypotheses, but the consulting practice starts with you. That is what the founding rate is for, and it is why the first two engagements are priced below what the work is worth.
Bring one workflow.
Not a strategy conversation. One process your team repeats, where it currently breaks, and what a good result would change. That is enough for me to tell you whether a sprint fits, and to say so plainly if it does not.
That address is a plain mailbox rather than a form, and it is the same one the rest of the work uses. A dedicated one at this domain is on the list; publishing it before it can receive mail would have been the first broken promise of the engagement.
I reply to every message that describes a real process, including to say no. If a sprint is the wrong shape for what you need, I would rather tell you in week zero than in week two.