Design and review
How much design a change deserves, and the tools for each rung: architect, arena, swarm, interrogate, and a second model on your work.
One attempt at a hard design locks in the first shape the model thought of. These tools spend more effort where a wrong shape costs more. Most changes need none of them.
How much design a change deserves
| The change | Reach for |
|---|---|
| Small and finished, but you're unsure | interrogate, or a blast-radius check |
| Crosses function boundaries or moves ownership | architect, which runs an arena of sketches |
| A standalone decision where independent attempts help (a name, a format, an algorithm) | an arena directly |
| A coverage matrix, parallel checks, or a race | a swarm |
| A contested design that's expensive to reverse | architect, then interrogate or a second model before shipping |
Architect
you > architect the import pipeline before writing code. i care most about how callers use it.
you > architect with checkpoint. show me before implementing.The caller's usage comes first: the quickstart a consumer reads and two or three real call sites, with the types derived from them. At least two structurally different shapes get sketched, each screened for four red flags: a shallow module, information leakage, temporal decomposition, and pass-through methods. The deepest interface wins. A one-page rationale ships with the sketch, covering the usage, the shape, the tradeoffs accepted, and the alternative that lost. When implementation keeps fighting the sketch, the sketch is scrapped and redesigned smaller rather than patched.
Arena
N attempts at the same brief run in parallel, each in its own worktree or folder. A read-only judge scores them against a rubric the candidates never saw. Kevin reads every candidate, picks the one a maintainer can extend most easily as the base, grafts the best idea or two from the others, and verifies the result. When every candidate converges, that's the answer. When they wildly diverge, the brief was underspecified and gets reframed.
Swarm
N workers cover separate slices ("one worker per package") or race the same brief under a rule declared up front. Each reports PASS, ISSUES, or BLOCKED with evidence, and you get one table back instead of N transcripts. A missing result is a gap, never a pass.
Interrogate
you > interrogate the whole branch, skeptically. no nitpicks unless it's a real bug.Several reviewers attack the same diff and intent from independent angles: correctness, root cause versus symptom, structure, verification, complexity, and security, plus a strict code-quality lens that hunts for the reframing that deletes whole branches. Kevin then judges as the lead, not an aggregator. It sorts every finding into Act on, Consider, Noted, or Dismissed, with a reason for each dismissal so you can overrule it, and maps where the reviewers agreed. Nothing is applied automatically.
A second model on your work
Reviewers inside one session share a model's blind spots. For model diversity, adversarial-review briefs a different model or host on your branch, working tree, or plan, then verifies each of its findings against the code and fixes the real ones, all on one dossier. See Second-model review.
The comment pass
Before any diff reaches you, a reviewer that didn't write the code strips comments that narrate what the next line does. It keeps license headers, public API contracts, and constraints forced by something outside your control. A comment explaining a surprise in your own code comes back as a flag to fix the code instead, and a lint suppression that protects correctness gets flagged for removal.
Principles
Twenty-three engineering principles, each a short rule with a test the agent runs on its own work. Name one to steer.
Pull requests
The engineer skill's three pull-request playbooks. Review a teammate's PR, answer the review on yours, and walk the room through yours. Nothing posts; you paste.