Skill
One skill routes every kind of code task to a playbook, pull requests included, applies the engineering principles, and runs a comment pass before you see the diff.
engineer fires on its own when you ask for code work in a repo: "fix the duplicate write", "why is startup slow", "refactor the parser", "how does billing retry work". It reads the repo's AGENTS.md first, matches the task to a playbook, and copies that playbook's steps into its todo list, so a skipped step stays visible with its reason.
/engineer help prints the whole menu in your terminal: every playbook grouped by what you're doing, an example ask and what you get for each, the principles you can steer by name, and the checks that run every time.
Playbooks
| You ask for | Playbook | What it holds you to |
|---|---|---|
| "How does X work?", "where should this live?" | how | Entry point to exit, the key types, where things live, the gotchas |
| "Why is X like this?" | why | Git history, PRs, and your own session logs as evidence, every claim labeled direct, supported, inferred, or unknown |
| A read-only question or a choice | investigation | A cited answer, verdict first |
| A bug | bug fix | Reproduce it on the real surface, binary-search the cause, failing test before the fix when it's cheap |
| New or changed behavior | feature | Name the data shape first; a throughput checkpoint before fan-out |
Simplify, clean up, /engineer simplify <target> | simplify | An audit against eight criteria and the common smells, findings ranked before anything changes, then the host's built-in simplify pass over the applied diff |
| Refactor, reshape | refactoring | Pin the behavior first, subtract before adding, keep it only if the code got easier to read |
| "Why is this slow?" | performance | A baseline trace, eight strategy families, one change at a time |
| Push a metric to a target | hillclimb | A frozen harness, a stop rule, one commit per accepted win |
| A leak, spin, or glitch | runtime forensics | Capture the live signal and prove the mechanism before believing it |
| A profile or trace someone captured | trace forensics | Make it queryable, narrow to the hot frame, attribute it to source |
| Match two UIs pixel for pixel | visual parity | Baseline first, image diff to zero |
| "What could this break?" | blast radius | The one fact the change is safe because of, proven by running code |
| Settle a design by looking | prototype | A throwaway behind one switcher, a recommendation |
| Large or cross-cutting work | figure it out | A designed workflow, a falsifiable definition of done, a decision log |
| "Keep going until done" | autonomous run | A checkable finish line; a duration doesn't count |
| "Pause here" / "pick up where it left off" | pause safely, session pickup | A resume note; the prior trail treated as authoritative |
| Write a skill; test a skill change | authoring a skill, eval | Validated frontmatter and links; blinded candidates and judge |
| Give a repo a way to prove the app works | verification skill | A generated verify-<app> skill with a feature map, proven once before handover |
| "Review PR 142" (a teammate's) | pr review | Seven lanes held to the principles, every finding verified, paste-ready comments. See Pull requests |
| "Reply to the comments on my PR" | pr replies | Every thread judged against the code, accurate ones fixed uncommitted, a reply for each |
| "Prep my PR for standup" | pr walkthrough | A standup script, the diff tour, the questions with receipts, a recording runbook |
Every time
- The data shape is named before logic is written.
- A question you could answer by running something gets run, not asked.
- Every claim carries its evidence or a label: measured, inferred, or guess.
- "Done" means proven on the real artifact, with how far it was proven stated.
- Before any diff reaches you, a separate reviewer strips narrating comments and flags workarounds for a proper fix.
Simplify
/engineer simplify <target> (or "simplify this", "is this over-engineered?") audits the target end to end against eight criteria: elegance, simplicity, conciseness, accuracy, robustness, reliability, best practices, and not over-engineered. It also checks a list of common smells and hunts dead code, confirming each removal with a search first. Findings come back ranked (critical, worth simplifying, nits) with file:line and the concrete change, and nothing changes until you pick. What you pick is applied with the refactoring playbook's guardrails: behavior pinned first, removals before reshapes, equivalence proven. Where the host ships a built-in simplify pass (Claude Code's /simplify), it then runs over the applied diff, and any edit it makes that doesn't trace to one of the criteria is reverted.
How to ask
you > the export writes two rows after a retry. repro first, then fix and verify.
you > architect the import pipeline before writing code. i care most about how callers use it.
you > blast radius of this diff
you > /engineer help
you > /engineer simplify the webhook handler
you > review PR 142
you > /engineer walkthrough 142 --rehearse
you > keep going until the checker reports zero old callers. log your decisions.Say the goal and what done means. Naming a sequence of skills usually reorders steps the playbook already sequences. "Don't change any code yet" keeps a request read-only.