AgentLayer▸docs
Engineering

Pull requests

The engineer skill's three pull-request playbooks. Review a teammate's PR, answer the review on yours, and walk the room through yours. Nothing posts; you paste.

Code review is where a human gates what agents write, and most pull requests now arrive faster than anyone can read them. The PR playbooks exist so you walk into that conversation understanding the change better than anyone else in it, holding verified defects rather than impressions, with the words already written. They are part of the engineer skill, so a review is held to the same principles and playbooks Kevin uses on its own code: the blast-radius check, the bug-fix discipline, the comment pass. Whose PR it is picks the playbook, and you can name one directly.

PlaybookWhose PRAskWhat you get
pr reviewa teammate's"review PR 142", engineer review 142The change explained, verified findings anchored to file and line, and paste-ready comments
pr repliesyours"reply to the comments on my PR", engineer replies 142Every review thread judged against the code, the accurate ones fixed uncommitted, and a reply drafted for each
pr walkthroughyours"help me present #142", engineer walkthrough 142A two-minute standup script, the diff tour in scroll order, the questions reviewers will ask with receipts, and a recording runbook for the test video

All three run on the GitHub pack's read-only tools and share one foundation: resolve the PR, read the prior context (earlier reports, the task it serves, the concept articles for the domains it touches, the repo's conventions), understand the change before judging it, and build and test it locally in a worktree. Every report lands in <HOME>/reports/reviews/ and shows up in the next session's start context. Nothing is ever posted, commented, approved, or pushed by the agent: the GitHub pack cannot write, by design, so nothing reaches a teammate without you deciding it should.

Reviewing a teammate's PR

engineer review 142 starts from the stance that the PR is wrong until the code proves otherwise. The title, body, and commit messages are claims to test, not context to trust.

  1. Prior context first. Earlier reports on the same PR, the task it serves, the concept articles for the domains it touches, and the repo's own conventions file.
  2. Understand before hunting. The diff is read the way the how playbook reads code, for what changed grouped by concern, why, how the data flows, and the blast radius across every caller. The report's "How it works" section is written before any finding, because a reviewer who cannot explain the change cannot judge it.
  3. Local verification. The head is checked out on a throwaway worktree, never the author's branch, and the repo's own build, lint, format, and test scripts run there. A green check on GitHub is a claim; the local run is the verdict.
  4. Seven adversarial lanes in parallel. Correctness, domain invariants and authorization, security, regressions and blast radius, conventions, tests, and PR hygiene, each covering every file on the checklist. Each lane reads the engineering principles it enforces before it judges, and every finding names the principle or repo rule it applies, so the author can read the rule instead of arguing with the reviewer. The regressions lane runs the blast-radius playbook: it names the one fact the change is safe because of and says how far it was proven. The conventions lane holds the change to the engineering principles as well as the repo's written rules, so a missed simplification, a conditional tangled into an unrelated flow, a leaked wire type, or an old path kept alive beside its replacement is a finding. The tests lane flags hollow tests, the ones that would pass if every import returned nothing, and asks for a missing test only where the repo's test policy calls for one.
  5. Every finding verified by a fresh agent that owes the lane nothing: re-anchored by snippet, blamed to this PR, traced from an entry point, scored. High scores become findings, middling ones become questions for the author, low ones are listed as dropped so you can rescue one. Authorization, data loss, and concurrency claims never drop silently, and a preference without a concrete cost never scores as a finding.
  6. The report. A verdict, the walkthrough, ranked findings with a paste block each, questions, what was verified clean, the checks that ran, and a summary comment for the conversation tab.
you  > review PR 142
you  > /agent-kevin:engineer review 142 --quick    (diff only, no worktree, for triage)
you  > /agent-kevin:engineer review 142            (on your own PR: the full lanes as a self-review)

The throwaway review worktree is removed once the report is saved. One holding uncommitted work or unpushed commits stays, and a worktree on your own branch is never touched.

Answering the review on your own PR

you  > what did reviewers say on my PR
you  > address the review on my branch
you  > /agent-kevin:engineer replies 142

The replies playbook pulls every thread, judges each comment against the current head rather than the snapshot it was written on, fixes what is accurate in the working tree (uncommitted, so git status is the reviewable set), pushes back on what is wrong with receipts, and writes a reply for every thread in PR scroll order. A bot's confidence is not evidence; its comments get the same treatment, except that a bot's security, data, migration, or concurrency claim is never dismissed on a reading alone: it is disproved with a run or put to you. Each fix follows the bug-fix playbook at the size of the defect (reproduced, fixed at the cause, a failing test first when the repo wants one), gets the comment pass, and comes with a suggested commit message. A bounded self-pass over your own diff runs the correctness, invariants, and regression lanes, so you find your own defect before the reviewer does.

Presenting your own PR

engineer walkthrough 142 is for the PR you are about to defend at standup or on a recorded test video. It rebuilds the change hunk by hunk from the diff, the task, and the sessions that wrote it, classifies each hunk by how hard it is to defend, and verifies the build locally. The doc it writes is meant for a hidden second screen: a standup script, the diff tour in the order GitHub shows it, the questions a reviewer will ask with the receipts to answer them, and a scene-by-scene recording runbook that covers every path the change adds rather than the happy one. Every receipt is labeled by how well it is known, the way the why playbook labels evidence, and a hunk you cannot defend becomes a gap with an action (simplify, revert, ask, or fix the PR description), never a bluff.

you  > help me present #142
you  > /agent-kevin:engineer walkthrough 142 --rehearse   (the room's questions, one at a time, misses corrected against the code)
you  > /agent-kevin:engineer walkthrough 142 --check ~/Desktop/pr-142.mp4   (did the recording cover every scene)

A review of the same PR later reads the walkthrough as your stated intent and test plan and checks the diff against it.

A second model on your work

A teammate's PR gets the review above. Your own work, whether a branch, a working tree, or a plan with no code at all, gets a different model on it with adversarial-review, which verifies each finding with the same rubric the review lanes use. See Second-model review.

What the playbooks never do

  • Never post to GitHub or any chat surface. You paste.
  • Never commit on a teammate's branch or check it out; a review runs on a throwaway worktree at the PR head.
  • Never commit or push. Fixes on your own branch stay in the working tree for you to commit.
  • Never put private material in a paste block: nothing from the agent home's memory or notes, no secrets, no customer identifiers.

On this page