CLAUDE CODE GUIDE · CODE REVIEW & TESTING

Code Review Skills for Claude Code: A Practical Shortlist

Claude Code has an official pull-request review plugin; other Agent Skills cover narrower jobs such as diff analysis, static checks, or browser evidence. This guide compares what each actually does, which tools it needs, and where human review still matters.

Published Updated ORIGINAL · SOURCE-AWARE
01

Start with the review job—not the tool name

People searching for a Claude Code code-review Skill may mean a GitHub pull-request workflow, a focused explanation of a diff, a static-analysis guardrail, or proof that a changed user flow still works. These jobs need different tools and should not be collapsed into one score or package.

SkillSignal compares the documented job, access requirements, compatibility evidence, and output of each option. It is a source-based shortlist—not a runtime benchmark. Keep the final merge decision with a reviewer who understands the repository and release context.

02

Claude Code's official /code-review plugin: PR workflow, not portable Skill

Anthropic's official Claude Code plugin repository lists a /code-review command for automated pull-request review using specialized agents and confidence-scored findings. The upstream README says the plugin is included with Claude Code; check that the command exists in your installed version instead of relying on old marketplace instructions.

This is a Claude Code plugin command, not a portable SKILL.md folder. Its README lists GitHub CLI and GitHub integration as requirements, and describes posting review comments. Treat it as write-capable: inspect the target PR and available permissions first. Confidence filtering can reduce noise, but it cannot certify that a finding is correct or that a PR is safe to merge.

03

Add a Skill only for evidence the review still lacks

differential-review focuses on the changed lines and likely blast radius. It can help structure a diff review, but it cannot prove runtime behavior or replace repository knowledge. Confirm that its source supports the Agent you plan to use.

Semgrep provides rule-based static checks; CodeQL can analyze supported languages and configured databases. Their results depend on rules, language coverage, configuration, and triage. Neither is a universal security certificate, and they are not substitutes for a pull-request workflow.

webapp-testing is a separate role for browser-level evidence on local web applications. It can verify a specified flow, but a passing browser run does not cover every state, browser, accessibility path, or production dependency.

04

A four-step review stack that stays auditable

First, define the change boundary: target branch, expected behavior, sensitive files, and the decision the review must support. Second, if using Claude Code's PR plugin, verify the intended repository and GitHub write permissions; inspect the diff before adding broader automation. Third, run the smallest applicable static or behavioral check and preserve its command, scope, and output. Fourth, compare all evidence with the acceptance criteria and record unresolved risks instead of converting a tool result into automatic approval.

  • The core review explains changed behavior and likely blast radius.
  • Every automated check states its files, rules, language support, and exclusions.
  • A GitHub-connected review has explicit permission to post or withhold comments.
  • Behavioral evidence covers the changed flow, not an unrelated smoke test.
  • A named reviewer owns the final decision and unresolved exceptions.
05

Native fit matters less than a reproducible boundary

A Skill may provide a native path for one Agent and a portable SKILL.md path for another. Portability can be valuable, but it is not evidence of equal testing, tool access, or permission behavior. Review the selected Agent's install path and required tools before treating two environments as equivalent.

For review work, reproducibility is the stronger selection signal: can another reviewer see the same source revision, run the same bounded check, and understand why a finding matters? If not, adding more Skills usually increases ambiguity rather than confidence.

06

Selection checklist

Choose the smallest combination that answers the current review question. Add another Skill only when it owns a distinct role and produces evidence the reviewer will actually use.

  • What exact decision must the pull-request review support?
  • Which files, commands, network calls, and external services can each Skill access?
  • Is the upstream source pinned, and has the Skill been independently runtime-tested?
  • What result would make the reviewer stop, escalate, or reject the change?

NEXT STEP

Build a smaller, more auditable review plan

Start with the Claude Code workflow or the narrowest compatible Skill, then add only evidence your reviewer will use.