Verification
What counts as proof. The proof ladder, a generated verification skill per repo, forensics, blast radius, and hillclimbing a metric.
A green build proves the code compiles. Kevin checks the real artifact: it runs the feature path, reads the stored value, and drives the page, the simulator, or the terminal you actually use. A wrong-surface or inconclusive result is reported as exactly that, never as a pass.
The proof ladder
Every load-bearing claim gets taken as far down this ladder as is cheap, and Kevin says where it stopped:
- You said so. Worth nothing on its own.
- You pointed at the line: a real
file:line, or the library's own source. - You walked the failure step by step and showed it can't be reached.
- You ran it: a script or test that calls the real code and fails loudly if you're wrong.
- You reproduced it in the running app.
A verification skill for your repo
you > give this repo a verification skillThe engineer skill interviews the repo, not you, to work out what a user touches, how the app starts, what can drive it, and what evidence proves behavior. It then writes a repo-local verify-<app> skill with Launch, Doctor, Drive, Evidence, and Cleanup sections, plus a feature map with one file per user-facing feature. It runs the new skill end to end once before handing it over, because a verification skill that never ran is a draft. A maintenance pass later rechecks every mapped feature from source and live, and fixes the map or reports a real regression.
Blast radius
you > blast radius of this diffListing callers is what grep is for. This finds the breakage grep won't show: a wire format, a database column, timing, a flag, code three hops away. Most risky-looking changes are safe because of one fact. Kevin names that fact and proves it by running code, then lists the confirmed risks, the cleared ones, and the cheapest check to run before merging.
Forensics
- Runtime forensics captures a live signal (a CPU profile, a heap snapshot, a trace), reduces it to the smoking gun, and proves the mechanism with a probe before mapping it to source.
- Trace forensics takes a profile someone already captured, makes it queryable, narrows to the hot frame or the retainer chain, and attributes it to source.
Both hand back a diagnosis, not a fix.
Hillclimb
For sustained improvement of one number:
- Build a measurement harness, prove it separates the slow case from the easy ones, then freeze it.
- Set a stop rule that pairs a target with a minimum number of attempts, so a lucky early win can't end the run.
- Loop one hypothesis at a time, keeping a change only when it moves the number past the noise with tests green.
Every attempt is logged, and the target never relaxes to declare victory.