Refolk
FrameworkEngineering and open source

The Code-Review Behavior Read: Mentor, Gatekeeper, or Rubber-Stamp

You can score one engineer's public review activity across four named dimensions and reach a Mentor, Gatekeeper, or Rubber-Stamp verdict after discounting AI-written comments.

15 min readLast reviewed October 7, 2026Read as Markdown

You are about to interview an engineer and you want to know, before the call, what their public code-review history actually proves about seniority and judgment. This guide is for engineering managers, technical founders, developer-relations leads, and technical sourcers who read GitHub profiles for a living. It gives you a four-dimension scoring model, a procedure to run it, and an explicit discount for the AI-written review comments that now contaminate the signal, so you land on one of three verdicts: Mentor, Gatekeeper, or Rubber-Stamp.

Most profile reads grade the code an engineer authored. The reviewer side of a profile is where mentorship and architectural judgment show up most clearly, and almost nobody grades it. That is the gap this document fills.

Why the reviewer side of a profile proves more than the author side

The code someone writes shows what they can build; the reviews they leave show what they can see in someone else's work. Reviewing forces judgment under constraint: a reviewer must decide what matters, say why, and decide whether to block a merge. That is the same judgment you are hiring for in a senior engineer, and it is visible on public GitHub without a single interview.

Google's own engineering practice is explicit that review comments should be clear and useful and should mostly explain why instead of what, and that reviewers should offer encouragement for good practices, because telling a developer what they did right can be more valuable for mentoring than telling them what they did wrong. That is a behavior, and behaviors leave a public trail: review state, review bodies, and inline comments pinned to specific diff lines.

The trail is pullable. Public review data comes from two REST endpoints. The reviews endpoint lists all reviews for a pull request in chronological order, and each review object carries a state, a body, the user, and an author_association. The review-comments endpoint lists inline comments, each attached to the diff_hunk it commented on. Both work unauthenticated for public resources. Review state is one of APPROVED, CHANGES_REQUESTED, COMMENTED, or PENDING, and submitted_at is how you derive time-to-first-review - but pending reviews are never submitted and carry no submitted_at, so they do not count.

208
US Engineering Managers in Refolk's index who list the explicit skill "Code Review"
Top employers are Google (4), Meta (3), and Cloudflare (2), which tells you how rarely this is named as a skill at all.

The scarcity in that number is the point. In Refolk's index of professional profiles, only 208 US Engineering Managers name Code Review as an explicit skill, which means you almost never find the signal pre-labeled. You have to read it off review history yourself.

The four dimensions that separate a Mentor from a Rubber-Stamp

Score four dimensions, each from 0 to 3, each backed by a linked example comment. The dimensions are specificity, resolution-orientation, mentorship, and judgment-under-disagreement. A Mentor scores high across all four. A Rubber-Stamp scores near zero on specificity and mentorship and leaves bare approvals. A Gatekeeper sits between: real blocking judgment, thin on teaching.

DimensionWhat a 3 looks likeWhat it looks like when it lies
SpecificityComment names the exact line and explains why it mattersGeneric "looks good" or "fix this" with no diff anchor
Resolution-orientationCHANGES_REQUESTED the author then resolvedBlocking comments that were ignored or dropped
MentorshipExplains the why, prefixes polish with "Nit:", encouragesOnly flaws, never context, never what was done right
Judgment-under-disagreementHolds a position, concedes when shown new factsApproves to avoid conflict, or digs in past evidence

Each dimension proves something specific and fails in a specific way. Specificity proves the reviewer read the diff rather than the title; it lies when a reviewer pastes the same encouraging line across unrelated PRs. Resolution-orientation proves the review changed the code; it lies when a reviewer files CHANGES_REQUESTED as a reflex and the author merges around it. Mentorship proves the reviewer invests in the author, and Google's "Nit:" convention - prefixing low-importance polish so the author knows it is optional - is a cheap, reliable tell. Judgment-under-disagreement is the hardest to fake and the most senior: it is visible in a thread where the reviewer held a line, the author pushed back, and the exchange resolved on facts rather than rank.

Placing the verdict on two axes

High mentorshipLow mentorship
Rubber-Stamp
Bare approvals, no diff-anchored feedback; do not credit as senior review
Gatekeeper-lite
Teaches but vaguely; probe whether comments actually land on the diff
Gatekeeper
Sharp, blocking, impersonal; strong judgment, weak on growing people
Mentor
Specific and encouraging; the signal you are hiring for
Low specificityHigh specificity
Specificity and mentorship together separate the three verdicts from each other.

How the AI-comment surge contaminates the signal, and how to discount it

Before you score anything, strip the AI. AI-written review comments now look exactly like a diligent human reviewer, and crediting a human for a bot's feedback is the single fastest way to overrate a profile. The discount is not optional - it is step zero of the read.

The scale of the contamination is real but not precisely published. No source gives the share of all GitHub review comments that are AI-authored. Strong proxies exist: CodeRabbit alone reports over 13 million pull requests reviewed across more than 2 million repositories, and is the most-installed AI app on GitHub and GitLab. For scale, Octoverse reports 518.7 million merged pull requests in public repositories, up 29% year over year, across more than 180 million developers. The direction is steep: cross-product AI-to-AI review grew more than two orders of magnitude from one quarter to the next in recent data.

MetricValueSource type
CodeRabbit PRs reviewed13M+Vendor-published
CodeRabbit repositories2M+Vendor-published
AI-to-AI cross-product review share of agent PRs~1.6%Independent study
CodeRabbit comment precision (acted on)49.2%Independent benchmark

Read that table with the source column in mind. The two vendor-published counts tell you reach, not quality. The independent Martian benchmark precision of 49.2% - roughly one comment in two leading to a code change - is your external anchor, because vendor-origin quality stats like a 1.7x-defect finding should be flagged as vendor-origin and not treated as ground truth.

Identify AI reviewers with two tiers of signal. Login alone is unreliable, so combine them.

AI-reviewer filter, run before scoring
Tier 1 - vendor account logins (exact match):
  coderabbitai[bot]
  devin-ai-integration[bot]
  gemini-code-assist[bot]
  github-copilot[bot]

Tier 2 - body signatures (substring / regex):
  Co-Authored-By: Claude
  cursor.com/agents
  chatgpt.com/codex
  "summarize by coderabbit.ai" (HTML marker)
  CodeRabbit Walkthrough / Summary block

Apply both tiers; a hit on either flags the comment as AI-assisted and removes it from the human score.

What does the noise look like once you strip it? An independent one-month log of 28 pull requests gives a usable baseline for AI comment categories, and it is sobering.

CategoryShare
Quality improvements35%
Nitpicking21%
Useless/noise15%
Wrong assumptions13%
Security/critical3%

Roughly a third of AI comments were genuine quality improvements and well over a third were nitpicking, noise, or wrong assumptions. If a profile's "prolific reviewer" reputation rests on comments that match this shape, you are reading a tool, not a person.

Running the AI filter and the evaluative split by hand across every repo a candidate touched is the slow part of this read. Refolk lets you ask for the people whose public review behavior already matches the pattern - change-requesters and mentors in a named ecosystem - so you start from a shortlist rather than a raw endpoint dump.

Counting comments without being fooled by volume

A high comment count is not a high judgment count. The most common false read is treating raw volume as evidence, and the data says why: non-evaluative comments - CI chatter, chatops commands, bare "lgtm" notices - account for 65.53% of human review comments on agent-authored PRs, against 93.56% on human-authored PRs. Either way, most comments steer rather than evaluate.

65.53%
Share of human review comments that are non-evaluative on agent-authored PRs
On human-authored PRs it climbs to 93.56%, so a raw comment count mostly measures chatter, not judgment.

So separate evaluative comments from steering before you score anything. An evaluative comment judges the code: it names a problem, proposes an approach, or weighs a tradeoff. A steering comment moves the process: it triggers CI, pings a reviewer, or acknowledges a merge. Only the evaluative set feeds the specificity and mentorship scores.

Two API mechanics matter here. First, a review object can return state COMMENTED with an empty body even though inline comments exist, so counting the review body alone will misread a careful reviewer as silent. Pull the comments endpoint, not the review body. Second, pending reviews carry no submitted_at, so exclude them from any timing analysis.

Run the read: seven steps from endpoint to verdict

This is the procedure. Budget roughly 90 minutes per candidate the first few times; it compresses with practice. The one judgment call to settle before you start: whether an author's own first pass counts as a review. Some teams treat self-review as a legitimate step; Google-style gating counts only an independent reviewer. Decide, then apply it consistently, because counting self-reviews inflates the review count with the author's own passes.

Scoring one engineer's review behavior

  1. Pull the review events
    Query the reviews and comments endpoints across repos the person reviewed on. Capture state, submitted_at, body, and inline comment counts per PR.
  2. Strip bot and AI-assisted comments
    Filter [bot] and vendor logins plus body signatures (Co-Authored-By trailers, coderabbit markers). Produce a human-only comment set with an AI-assisted flag per thread.
  3. Classify evaluative vs non-evaluative
    Separate real feedback from CI chatter, chatops, and bare lgtm. Produce an evaluative-comment count not inflated by steering.
  4. Score the four dimensions
    Score specificity, resolution-orientation, mentorship, and judgment-under-disagreement 0-3 each, with a linked example comment per score.
  5. Validate approval signals
    Separate self-merges and admin-bypass merges from true peer approvals and resolved change-requests. Produce an approval-quality flag.
  6. Check for suppressors
    Look for private employer repos, internal tooling, and non-English review culture. Note a confidence discount on the verdict.
  7. Land the verdict
    Combine scores and the approval flag into Mentor, Gatekeeper, or Rubber-Stamp, plus the two strongest supporting threads.

From raw review events to scorable comments

  1. All review comments pulled
    100%

    raw endpoint output

  2. After stripping bot and AI-assisted
    human-only

    removes vendor-authored feedback

  3. After removing non-evaluative
    ~34% on agent PRs

    steering and CI chatter gone

  4. Scored evaluative comments
    the signal

    what the four dimensions read

Every stage removes comments that would otherwise inflate a human reviewer's score.

The funnel is why volume misleads. By the time you remove bot comments and the 65-plus percent that is non-evaluative on agent PRs, the scorable set is a fraction of what a profile page advertises. Score the fraction, not the headline.

Validating approvals: why an APPROVED state is weak on its own

An approval proves peer-accepted judgment only when the reviewer is not the author and branch protection was actually enforced. GitHub does not let you approve your own PR, but you do not need to: with the right role you can check "Merge without waiting for requirements to be met" and merge unreviewed. Required-reviewer settings are also bypassable, and any reviewer can submit code on a PR during review and merge to main. So three common patterns are weak signals - self-merges, admin-bypass merges, and a maintainer waving through an outside contributor.

The strong signal is the opposite: a CHANGES_REQUESTED that the author resolved. That proves a peer found a problem and the code changed because of it. When you build the approval-quality flag, sort every approval into one of these buckets:

  • Peer approval with branch protection active: strong, counts fully.
  • Resolved change-request: strongest single signal, counts double in your read.
  • Self-merge or admin-bypass merge: discard, proves nothing about judgment.
  • Maintainer approving an outside contributor with no enforcement: weak, note it but do not lean on it.
An approval proves nobody objected; a resolved change-request proves a peer found the problem and the code moved.

How this read goes wrong: the seven failure modes

This is the section to read twice, because a scoring model that overclaims is worse than none. Each failure mode below produces a confident but false verdict. The check column is how you catch it before it reaches the hiring conversation.

Failure modeFalse verdict it producesCheck
Counting comment volume as judgment"Prolific, therefore senior"Reclassify evaluative vs steering; up to 65.53% is non-evaluative
Crediting a human for a bot"Diligent reviewer"Filter [bot] logins and body signatures first
Trusting an APPROVED state"Peer-validated"Confirm reviewer is not author and protection was active
Reading empty-body reviews as rubber-stamps"Thin reviewer"Pull the comments endpoint, not the review body
Mislabeling terse or non-English culture"Rubber-Stamp"Look for private or internal tooling before downgrading
Taking vendor benchmarks as ground truthOverrated AI-assisted profileFlag vendor-origin stats; anchor on the independent Martian figure
Counting self-review as reviewInflated review countDrop author-authored reviews before scoring

Two of these deserve extra weight. The suppressor problem is structural: private-repo work, internal review tools like Gerrit and Phabricator, and non-English "lgtm" review cultures all leave thin public trails, and the magnitude of each is not publicly established. A thin trail is low confidence, not a Rubber-Stamp verdict. When you hit a suppressor, note it and move the question to the interview rather than inferring an answer from silence.

The vendor-benchmark problem is about where numbers come from. CodeRabbit's own defect and precision figures are published by the vendor whose product they describe. The independent Martian benchmark - near 49.2% precision - is the external anchor. When a profile's review quality rests on AI-assisted comments, grade the comments against the independent baseline, not the vendor's.

Verify before you call it: the pre-verdict checklist

Run this checklist before you write down Mentor, Gatekeeper, or Rubber-Stamp. If any item fails, the verdict is not ready.

Before you record a verdict

  • Bot logins and AI body signatures are filtered out of the scored comment set
  • Comments are split into evaluative and non-evaluative, and only evaluative ones were scored
  • Comment counts came from the comments endpoint, not the review body field
  • Each of the four dimensions has a 0-3 score with a linked example comment
  • Approvals are sorted into peer, resolved-change-request, self-merge, and admin-bypass buckets
  • Self-reviews by the author were excluded per the rule you set up front
  • Suppressors (private repos, internal tooling, non-English culture) are noted as a confidence discount
  • The verdict carries its two strongest supporting threads

Keeping the read current as AI review tooling changes

The one part of this model with a short shelf life is the AI-reviewer filter, because vendor accounts and body signatures change as tools ship and rebrand. The dimension scoring does not expire - specificity, resolution-orientation, mentorship, and judgment-under-disagreement are durable - but the login and signature list in the template will drift. Treat it as a maintained list, not a fixed one.

Re-check the filter by sampling. Pull a batch of recent review comments from active public repos, group by author login, and look for any high-volume reviewer whose comments carry a Walkthrough or Summary block, a Co-Authored-By trailer, or a vendor URL you do not yet filter. Add new signatures as you find them. The two-tier framework itself is stable: a vendor-controlled account login plus a high-confidence body signature, combined because login alone is unreliable. One research corpus reached 94% attribution precision using commit signatures, bot markers, and PR metadata together, which is the standard to hold your own filter against.

Finally, keep the supply side in view when you decide how hard to read. Staff-level reviewer talent is not evenly distributed, and that shapes how much you can afford to discount.

MarketStaff SWE profilesRatio vs US
United States32,7961.00x
Germany1,1510.035x

In Refolk's index, the United States returns 32,796 Staff Software Engineer profiles against 1,151 in Germany - the US pool is roughly 28.5 times larger. In a deep pool you can afford a strict read and reject on a weak trail. In a thin pool, a suppressor-heavy profile deserves the benefit of the doubt and an interview, not a Rubber-Stamp verdict on absence of evidence. The model is the same everywhere; how aggressively you apply the discount depends on how many comparable reviewers you have to choose from.

Questions practitioners ask

How do I tell an AI-written review comment from a human one on GitHub?

Use two signals. First, the account login: known AI reviewers include coderabbitai[bot], devin-ai-integration[bot], gemini-code-assist[bot], and github-copilot[bot]. Second, body signatures: a Co-Authored-By: Claude trailer, a cursor.com/agents or chatgpt.com/codex URL, or CodeRabbit's auto-generated HTML marker and Walkthrough block. Login alone is unreliable, so researchers combine both; one corpus reached 94% attribution precision using commit signatures, bot markers, and PR metadata.

Does an APPROVED review on GitHub prove a peer validated the code?

Not on its own. You cannot approve your own pull request, but with the right role you can merge without waiting for requirements, and required-reviewer settings are bypassable by any reviewer who submits code during review. So self-merges, admin-bypass merges, and a maintainer waving through an outside contributor are weak signals. A CHANGES_REQUESTED that the author actually resolved is far stronger evidence of accepted judgment.

Why does the GitHub API show an empty review body when I can see comments in the UI?

A review object can return state COMMENTED with an empty body even though inline comments are visible. The review body and the inline comments are different fields. If you count only the review body you will misread a careful reviewer as a rubber-stamp, so always pull comment volume from the review-comments endpoint rather than the review object itself.

Can I score review behavior for an engineer who works mostly in private repos?

Partly. Private-repo work, employer-internal review tooling like Gerrit or Phabricator, and non-English review culture all suppress the public signal, and the size of each effect is not publicly established. Treat a thin public trail as low confidence, not as a Rubber-Stamp verdict. Note the suppressor explicitly and plan to probe review behavior in the interview instead of inferring it from silence.

How much of the AI comment signal is actually noise worth discounting?

Vendor and independent figures diverge, so treat them as a range. One independent benchmark put CodeRabbit comment precision near 49.2%, roughly one comment in two leading to a code change. An independent one-month log of 28 pull requests found 72% of findings relevant, with 21% nitpicking and 15% useless noise. The point is not the exact number but that a large share of AI comments is low-value, so crediting a human for them inflates the read.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next