The Code-Review Behavior Read: Mentor, Gatekeeper, or Rubber-Stamp
You can score one engineer's public review activity across four named dimensions and reach a Mentor, Gatekeeper, or Rubber-Stamp verdict after discounting AI-written comments.
You are about to interview an engineer and you want to know, before the call, what their public code-review history actually proves about seniority and judgment. This guide is for engineering managers, technical founders, developer-relations leads, and technical sourcers who read GitHub profiles for a living. It gives you a four-dimension scoring model, a procedure to run it, and an explicit discount for the AI-written review comments that now contaminate the signal, so you land on one of three verdicts: Mentor, Gatekeeper, or Rubber-Stamp.
Most profile reads grade the code an engineer authored. The reviewer side of a profile is where mentorship and architectural judgment show up most clearly, and almost nobody grades it. That is the gap this document fills.
Why the reviewer side of a profile proves more than the author side
The code someone writes shows what they can build; the reviews they leave show what they can see in someone else's work. Reviewing forces judgment under constraint: a reviewer must decide what matters, say why, and decide whether to block a merge. That is the same judgment you are hiring for in a senior engineer, and it is visible on public GitHub without a single interview.
Google's own engineering practice is explicit that review comments should be clear and useful and should mostly explain why instead of what, and that reviewers should offer encouragement for good practices, because telling a developer what they did right can be more valuable for mentoring than telling them what they did wrong. That is a behavior, and behaviors leave a public trail: review state, review bodies, and inline comments pinned to specific diff lines.
The trail is pullable. Public review data comes from two REST endpoints. The reviews endpoint lists all reviews for a pull request in chronological order, and each review object carries a state, a body, the user, and an author_association. The review-comments endpoint lists inline comments, each attached to the diff_hunk it commented on. Both work unauthenticated for public resources. Review state is one of APPROVED, CHANGES_REQUESTED, COMMENTED, or PENDING, and submitted_at is how you derive time-to-first-review - but pending reviews are never submitted and carry no submitted_at, so they do not count.
The scarcity in that number is the point. In Refolk's index of professional profiles, only 208 US Engineering Managers name Code Review as an explicit skill, which means you almost never find the signal pre-labeled. You have to read it off review history yourself.
The four dimensions that separate a Mentor from a Rubber-Stamp
Score four dimensions, each from 0 to 3, each backed by a linked example comment. The dimensions are specificity, resolution-orientation, mentorship, and judgment-under-disagreement. A Mentor scores high across all four. A Rubber-Stamp scores near zero on specificity and mentorship and leaves bare approvals. A Gatekeeper sits between: real blocking judgment, thin on teaching.
| Dimension | What a 3 looks like | What it looks like when it lies |
|---|---|---|
| Specificity | Comment names the exact line and explains why it matters | Generic "looks good" or "fix this" with no diff anchor |
| Resolution-orientation | CHANGES_REQUESTED the author then resolved | Blocking comments that were ignored or dropped |
| Mentorship | Explains the why, prefixes polish with "Nit:", encourages | Only flaws, never context, never what was done right |
| Judgment-under-disagreement | Holds a position, concedes when shown new facts | Approves to avoid conflict, or digs in past evidence |
Each dimension proves something specific and fails in a specific way. Specificity proves the reviewer read the diff rather than the title; it lies when a reviewer pastes the same encouraging line across unrelated PRs. Resolution-orientation proves the review changed the code; it lies when a reviewer files CHANGES_REQUESTED as a reflex and the author merges around it. Mentorship proves the reviewer invests in the author, and Google's "Nit:" convention - prefixing low-importance polish so the author knows it is optional - is a cheap, reliable tell. Judgment-under-disagreement is the hardest to fake and the most senior: it is visible in a thread where the reviewer held a line, the author pushed back, and the exchange resolved on facts rather than rank.
Placing the verdict on two axes
How the AI-comment surge contaminates the signal, and how to discount it
Before you score anything, strip the AI. AI-written review comments now look exactly like a diligent human reviewer, and crediting a human for a bot's feedback is the single fastest way to overrate a profile. The discount is not optional - it is step zero of the read.
The scale of the contamination is real but not precisely published. No source gives the share of all GitHub review comments that are AI-authored. Strong proxies exist: CodeRabbit alone reports over 13 million pull requests reviewed across more than 2 million repositories, and is the most-installed AI app on GitHub and GitLab. For scale, Octoverse reports 518.7 million merged pull requests in public repositories, up 29% year over year, across more than 180 million developers. The direction is steep: cross-product AI-to-AI review grew more than two orders of magnitude from one quarter to the next in recent data.
| Metric | Value | Source type |
|---|---|---|
| CodeRabbit PRs reviewed | 13M+ | Vendor-published |
| CodeRabbit repositories | 2M+ | Vendor-published |
| AI-to-AI cross-product review share of agent PRs | ~1.6% | Independent study |
| CodeRabbit comment precision (acted on) | 49.2% | Independent benchmark |
Read that table with the source column in mind. The two vendor-published counts tell you reach, not quality. The independent Martian benchmark precision of 49.2% - roughly one comment in two leading to a code change - is your external anchor, because vendor-origin quality stats like a 1.7x-defect finding should be flagged as vendor-origin and not treated as ground truth.
Identify AI reviewers with two tiers of signal. Login alone is unreliable, so combine them.
Tier 1 - vendor account logins (exact match): coderabbitai[bot] devin-ai-integration[bot] gemini-code-assist[bot] github-copilot[bot] Tier 2 - body signatures (substring / regex): Co-Authored-By: Claude cursor.com/agents chatgpt.com/codex "summarize by coderabbit.ai" (HTML marker) CodeRabbit Walkthrough / Summary block
Apply both tiers; a hit on either flags the comment as AI-assisted and removes it from the human score.
What does the noise look like once you strip it? An independent one-month log of 28 pull requests gives a usable baseline for AI comment categories, and it is sobering.
| Category | Share |
|---|---|
| Quality improvements | 35% |
| Nitpicking | 21% |
| Useless/noise | 15% |
| Wrong assumptions | 13% |
| Security/critical | 3% |
Roughly a third of AI comments were genuine quality improvements and well over a third were nitpicking, noise, or wrong assumptions. If a profile's "prolific reviewer" reputation rests on comments that match this shape, you are reading a tool, not a person.
Running the AI filter and the evaluative split by hand across every repo a candidate touched is the slow part of this read. Refolk lets you ask for the people whose public review behavior already matches the pattern - change-requesters and mentors in a named ecosystem - so you start from a shortlist rather than a raw endpoint dump.
Counting comments without being fooled by volume
A high comment count is not a high judgment count. The most common false read is treating raw volume as evidence, and the data says why: non-evaluative comments - CI chatter, chatops commands, bare "lgtm" notices - account for 65.53% of human review comments on agent-authored PRs, against 93.56% on human-authored PRs. Either way, most comments steer rather than evaluate.
So separate evaluative comments from steering before you score anything. An evaluative comment judges the code: it names a problem, proposes an approach, or weighs a tradeoff. A steering comment moves the process: it triggers CI, pings a reviewer, or acknowledges a merge. Only the evaluative set feeds the specificity and mentorship scores.
Two API mechanics matter here. First, a review object can return state COMMENTED with an empty body even though inline comments exist, so counting the review body alone will misread a careful reviewer as silent. Pull the comments endpoint, not the review body. Second, pending reviews carry no submitted_at, so exclude them from any timing analysis.
Run the read: seven steps from endpoint to verdict
This is the procedure. Budget roughly 90 minutes per candidate the first few times; it compresses with practice. The one judgment call to settle before you start: whether an author's own first pass counts as a review. Some teams treat self-review as a legitimate step; Google-style gating counts only an independent reviewer. Decide, then apply it consistently, because counting self-reviews inflates the review count with the author's own passes.
Scoring one engineer's review behavior
- Pull the review eventsQuery the reviews and comments endpoints across repos the person reviewed on. Capture state, submitted_at, body, and inline comment counts per PR.
- Strip bot and AI-assisted commentsFilter [bot] and vendor logins plus body signatures (Co-Authored-By trailers, coderabbit markers). Produce a human-only comment set with an AI-assisted flag per thread.
- Classify evaluative vs non-evaluativeSeparate real feedback from CI chatter, chatops, and bare lgtm. Produce an evaluative-comment count not inflated by steering.
- Score the four dimensionsScore specificity, resolution-orientation, mentorship, and judgment-under-disagreement 0-3 each, with a linked example comment per score.
- Validate approval signalsSeparate self-merges and admin-bypass merges from true peer approvals and resolved change-requests. Produce an approval-quality flag.
- Check for suppressorsLook for private employer repos, internal tooling, and non-English review culture. Note a confidence discount on the verdict.
- Land the verdictCombine scores and the approval flag into Mentor, Gatekeeper, or Rubber-Stamp, plus the two strongest supporting threads.
From raw review events to scorable comments
- 100%All review comments pulled
raw endpoint output
- human-onlyAfter stripping bot and AI-assisted
removes vendor-authored feedback
- ~34% on agent PRsAfter removing non-evaluative
steering and CI chatter gone
- the signalScored evaluative comments
what the four dimensions read
The funnel is why volume misleads. By the time you remove bot comments and the 65-plus percent that is non-evaluative on agent PRs, the scorable set is a fraction of what a profile page advertises. Score the fraction, not the headline.
Validating approvals: why an APPROVED state is weak on its own
An approval proves peer-accepted judgment only when the reviewer is not the author and branch protection was actually enforced. GitHub does not let you approve your own PR, but you do not need to: with the right role you can check "Merge without waiting for requirements to be met" and merge unreviewed. Required-reviewer settings are also bypassable, and any reviewer can submit code on a PR during review and merge to main. So three common patterns are weak signals - self-merges, admin-bypass merges, and a maintainer waving through an outside contributor.
The strong signal is the opposite: a CHANGES_REQUESTED that the author resolved. That proves a peer found a problem and the code changed because of it. When you build the approval-quality flag, sort every approval into one of these buckets:
- Peer approval with branch protection active: strong, counts fully.
- Resolved change-request: strongest single signal, counts double in your read.
- Self-merge or admin-bypass merge: discard, proves nothing about judgment.
- Maintainer approving an outside contributor with no enforcement: weak, note it but do not lean on it.
An approval proves nobody objected; a resolved change-request proves a peer found the problem and the code moved.
How this read goes wrong: the seven failure modes
This is the section to read twice, because a scoring model that overclaims is worse than none. Each failure mode below produces a confident but false verdict. The check column is how you catch it before it reaches the hiring conversation.
| Failure mode | False verdict it produces | Check |
|---|---|---|
| Counting comment volume as judgment | "Prolific, therefore senior" | Reclassify evaluative vs steering; up to 65.53% is non-evaluative |
| Crediting a human for a bot | "Diligent reviewer" | Filter [bot] logins and body signatures first |
| Trusting an APPROVED state | "Peer-validated" | Confirm reviewer is not author and protection was active |
| Reading empty-body reviews as rubber-stamps | "Thin reviewer" | Pull the comments endpoint, not the review body |
| Mislabeling terse or non-English culture | "Rubber-Stamp" | Look for private or internal tooling before downgrading |
| Taking vendor benchmarks as ground truth | Overrated AI-assisted profile | Flag vendor-origin stats; anchor on the independent Martian figure |
| Counting self-review as review | Inflated review count | Drop author-authored reviews before scoring |
Two of these deserve extra weight. The suppressor problem is structural: private-repo work, internal review tools like Gerrit and Phabricator, and non-English "lgtm" review cultures all leave thin public trails, and the magnitude of each is not publicly established. A thin trail is low confidence, not a Rubber-Stamp verdict. When you hit a suppressor, note it and move the question to the interview rather than inferring an answer from silence.
The vendor-benchmark problem is about where numbers come from. CodeRabbit's own defect and precision figures are published by the vendor whose product they describe. The independent Martian benchmark - near 49.2% precision - is the external anchor. When a profile's review quality rests on AI-assisted comments, grade the comments against the independent baseline, not the vendor's.
Verify before you call it: the pre-verdict checklist
Run this checklist before you write down Mentor, Gatekeeper, or Rubber-Stamp. If any item fails, the verdict is not ready.
Before you record a verdict
- Bot logins and AI body signatures are filtered out of the scored comment set
- Comments are split into evaluative and non-evaluative, and only evaluative ones were scored
- Comment counts came from the comments endpoint, not the review body field
- Each of the four dimensions has a 0-3 score with a linked example comment
- Approvals are sorted into peer, resolved-change-request, self-merge, and admin-bypass buckets
- Self-reviews by the author were excluded per the rule you set up front
- Suppressors (private repos, internal tooling, non-English culture) are noted as a confidence discount
- The verdict carries its two strongest supporting threads
Keeping the read current as AI review tooling changes
The one part of this model with a short shelf life is the AI-reviewer filter, because vendor accounts and body signatures change as tools ship and rebrand. The dimension scoring does not expire - specificity, resolution-orientation, mentorship, and judgment-under-disagreement are durable - but the login and signature list in the template will drift. Treat it as a maintained list, not a fixed one.
Re-check the filter by sampling. Pull a batch of recent review comments from active public repos, group by author login, and look for any high-volume reviewer whose comments carry a Walkthrough or Summary block, a Co-Authored-By trailer, or a vendor URL you do not yet filter. Add new signatures as you find them. The two-tier framework itself is stable: a vendor-controlled account login plus a high-confidence body signature, combined because login alone is unreliable. One research corpus reached 94% attribution precision using commit signatures, bot markers, and PR metadata together, which is the standard to hold your own filter against.
Finally, keep the supply side in view when you decide how hard to read. Staff-level reviewer talent is not evenly distributed, and that shapes how much you can afford to discount.
| Market | Staff SWE profiles | Ratio vs US |
|---|---|---|
| United States | 32,796 | 1.00x |
| Germany | 1,151 | 0.035x |
In Refolk's index, the United States returns 32,796 Staff Software Engineer profiles against 1,151 in Germany - the US pool is roughly 28.5 times larger. In a deep pool you can afford a strict read and reject on a weak trail. In a thin pool, a suppressor-heavy profile deserves the benefit of the doubt and an interview, not a Rubber-Stamp verdict on absence of evidence. The model is the same everywhere; how aggressively you apply the discount depends on how many comparable reviewers you have to choose from.
Questions practitioners ask
How do I tell an AI-written review comment from a human one on GitHub?
Use two signals. First, the account login: known AI reviewers include coderabbitai[bot], devin-ai-integration[bot], gemini-code-assist[bot], and github-copilot[bot]. Second, body signatures: a Co-Authored-By: Claude trailer, a cursor.com/agents or chatgpt.com/codex URL, or CodeRabbit's auto-generated HTML marker and Walkthrough block. Login alone is unreliable, so researchers combine both; one corpus reached 94% attribution precision using commit signatures, bot markers, and PR metadata.
Does an APPROVED review on GitHub prove a peer validated the code?
Not on its own. You cannot approve your own pull request, but with the right role you can merge without waiting for requirements, and required-reviewer settings are bypassable by any reviewer who submits code during review. So self-merges, admin-bypass merges, and a maintainer waving through an outside contributor are weak signals. A CHANGES_REQUESTED that the author actually resolved is far stronger evidence of accepted judgment.
Why does the GitHub API show an empty review body when I can see comments in the UI?
A review object can return state COMMENTED with an empty body even though inline comments are visible. The review body and the inline comments are different fields. If you count only the review body you will misread a careful reviewer as a rubber-stamp, so always pull comment volume from the review-comments endpoint rather than the review object itself.
Can I score review behavior for an engineer who works mostly in private repos?
Partly. Private-repo work, employer-internal review tooling like Gerrit or Phabricator, and non-English review culture all suppress the public signal, and the size of each effect is not publicly established. Treat a thin public trail as low confidence, not as a Rubber-Stamp verdict. Note the suppressor explicitly and plan to probe review behavior in the interview instead of inferring it from silence.
How much of the AI comment signal is actually noise worth discounting?
Vendor and independent figures diverge, so treat them as a range. One independent benchmark put CodeRabbit comment precision near 49.2%, roughly one comment in two leading to a code change. An independent one-month log of 28 pull requests found 72% of findings relevant, with 21% nitpicking and 15% useless noise. The point is not the exact number but that a large share of AI comments is low-value, so crediting a human for them inflates the read.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.