Hey Kaleb & Co. · Implementation Pack
The review that finds what you can't see from inside your own work — then hands you the findings and stops, so you stay the one making the call. This is the method behind this week's self-review loop: a prompt you can run today, and the engine underneath it.
The pack for the essay Claude F*cked With My Head.
Start today · The self-review loop
The portable version — tool-agnostic, so it works in any capable model, not just Claude Code. Drop it at the end of an AI-assisted task and it reviews its own work adversarially before it reaches you. Count how many problems it catches on its own. That number is the work you were doing by hand.
Paste it at the end of your prompt, or after the model hands you a draft.
Before you hand this back, review your own work. First, in one line, name what a great version must achieve — that's the bar. Then critique it as a skeptical expert trying to break it. Flag each issue as CRITICAL (wrong, unsafe, or misses the ask), MAJOR (materially weakens it), or MINOR (polish) — quote the exact text and say why. Fix the Critical and Major issues, changing the minimum: preserve my intent, scope, and voice, and flag anything that's a judgment call instead of rewriting it. Then re-read only the revised version as if you'd never seen it, and catch what the first pass rationalized away. Stop when only Minor issues remain. Hand back the revised work, the issues you left, and the judgment calls you flagged.
One honest limit — and one job that stays yours. This is the model checking its own work, so it shares its own blind spots. For anything high-stakes, run the review in a fresh chat or a second model — a clean read catches what a self-review talks itself out of. And it doesn't replace you: the loop clears the banal checks so your time and talent aren't spent there, but you still read the output and confirm it's right. The tool refines; you decide.
Go deeper · The /council skill
Here's the rule the whole skill is built on: findings and fixes are separate steps, with your decision in between. A reviewer that diagnoses and repairs in the same breath quietly trains you to skip the diagnosis. That's the atrophy machine — the exact thing this edition is about.
So the split is deliberate. Finding the flaw is the hard part you can't do from inside your own frame — worth handing off. Fixing it is where your judgment actually lives — hand that over every time and it goes away. The council delivers the findings, then asks how you want to work them:
The one exception: consent, privacy, security, and anything that could harm someone not in the room are never coached — they're stated outright, with the fix written in, because the cost of you not getting there lands on a participant, not on you.
The artifact
Two more things it does before it trusts itself. It audits its own findings against your actual work — models import facts that were never there and get numbers wrong, so every load-bearing finding gets checked, and the ones that don't survive are shown, not quietly dropped. And it runs in three modes: council to test whether it's right, panel to check it's ready for a team, gate to hold it against a standard. It never signs off — every read is advisory, and you own the call.
The single-pass version — a decorrelated read, a self-audit, findings only, then it stops and asks. Swap in your work and one line of stakes. For the most independent read, run it in a fresh chat or a second model.
You are running a review council on the WORK below. Find what I can't see
from inside my own frame — then hand the findings back so I keep the
judgment. Follow this exactly.
STAKES: [one line — what this decides, what being wrong costs]
WORK:
[paste the plan, guide, draft, analysis, or decision here]
1. READ IT FROM FOUR ANGLES, committing fully to each before the next:
- Correctness / risk: where does it break, what fragile assumption is
load-bearing?
- Framing: is this even the right problem, or a good answer to the
wrong one?
- Coverage: what's missing, unclaimed, or assumed?
- The people affected: consent, privacy, who's absent from the sample,
who bears the cost of being wrong?
Ground every point in the actual work. Generic advice that would fit
anything is the failure mode — cut it.
2. AUDIT YOURSELF before showing me anything. For each load-bearing point
(one whose removal would change your verdict), go back to the WORK and
check it: did you import a fact that isn't there? Is every number right?
Mark each CONFIRMED, CORRECTED (give the correction), or OVERTURNED (cut
it — but list what you cut and why). Don't quietly drop your own misses.
3. FINDINGS ONLY — do NOT rewrite the work. Return:
- READ: Ship / Revise / Rework / Stop, one line with the reason. Any
consent, privacy, legal, security, or harm blocker means it can't be
"Ship."
- MUST FIX FIRST: consent / privacy / legal / security / harm findings
ONLY, each with the concrete fix written out (the actual clause, the
field to cut). Skip this section if there are none.
- WHAT BROKE: each finding in priority order — the problem, why it
matters here, its audit mark, and ANCHORED (point at the line) or
NORMATIVE (an opinion about good practice). No fixes yet.
- WHAT HOLDS UP: briefly, what survived.
4. THEN STOP AND ASK how I want to work the findings:
- Give me the fixes.
- Coach me through them — one at a time, ask don't tell.
- Coach the top one (the finding with the most judgment in it), fix the
rest.
Wait for my answer. Don't fix and ask in the same message.
Stay adversarial, but don't invent problems — if it's genuinely sound, say
so briefly and stop. Everything you return is advisory; I own the call.
The blind-review lineage traces to Andrej Karpathy's LLM Council and multi-agent-debate research; the findings/fixes fork and the source audit are my own. Full sourcing travels with the download.
Run it on the thing you're most sure about. That's where it earns its keep.
Why it's built this way
The council isn't a separate philosophy bolted onto the work — it's my HEARTS framework running as software. Every move maps to a lens:
Human-Led, Amplification over Automation. The fork keeps the judgment yours — it surfaces what you can't see, and never takes over the fix.
Rigorous, Transparent. It audits its own findings against the source, and shows you the ones it overturned instead of hiding them.
Safe, Experience-Focused. It scrubs participant data before any model sees it — and weighs the experience of all three people HEARTS centers: the participant, the practitioner, and the partner.
The findings are the gift. The fix is yours to make.
Work with me
I help UX research teams put AI into their workflows with intention and integrity — without becoming the only human holding the whole thing up. If that's the problem you're sitting with, let's talk.

More from the Library
A six-lens self-check for AI-assisted research — plus an adversarial prompt that turns any model into an outside HEARTS reviewer.
Open the kit → Aug 26, 2026A readiness quiz and path checklists for deciding whether to build or buy your research repository.
Open the kit → Aug 20, 2026Seven checks to run and ten questions to ask your vendor before you connect an MCP server to your research.
Open the kit →