Refolk
ReferenceEngineering and open source

The Pull-Request Signal Reference: What Code Really Proves

You will be able to point at any signal in a candidate's public pull request and state what it proves about engineering level and how it can mislead you.

15 min readLast reviewed October 1, 2026Read as Markdown

Key takeaways

  • In an eye-tracking study of programmers reviewing pull requests, attention split 57% to code, 32% to technical signals, and 10% to social signals, and participants fixated on social signals more than they recalled.
  • A pre-registered experiment with 1,026 engineers found identical code rated 9% less competent when the author was labeled as using AI, with women penalized 13% versus 6% for men.
  • A squash merge can record the merger rather than the original author as the commit author, so a clean single merge commit can erase the granularity and authorship you came to read.
  • Review difficulty rises with file count, not line count: a pull request touching 20 or more files reads as weak scoping even when the line count is small.
  • A separate study of 447 engineers in an AI-normalized org found a seniority label biased competence judgments while an AI-disclosure badge did not, so label bias is context-dependent, not universal.
  • Score the diff before you look at the profile; the measured bias lives in competence ratings, not quality ratings, and hiding the author label collapses the gap.

You are about to interview an engineer and you have their public pull requests open in another tab. This reference tells you what each thing in the diff and the review thread actually proves about engineering judgment, and the documented way each signal lies to you. It is for engineering managers, technical founders, developer-relations leads, and technical sourcers who read code to decide, not sourcers scanning stars and contribution graphs. Jump to the row you need and leave.

Most hiring content about GitHub stops at the profile surface: stars, the green contribution grid, pinned repositories. That is a sourcer's view. A technical reviewer reads inside the diff and inside the conversation, where the real signal lives and where the real traps are. The sections below give one row per in-PR signal, a procedure you can run with your hands, and a failure-mode section with real weight, because the ways this goes wrong are the most valuable part of any standard.

Where your attention actually goes during review

Your eyes spend less time on the code than you think. In an eye-tracking study of programmers reviewing pull requests, attention split roughly 57% to the code snippet, 32% to supplemental technical signals such as previous contributions and popular repositories, and 10% to social signals like the author's avatar and follower count. Critically, programmers fixated on social signals more than they recalled doing so.

43%
Share of review time spent off the code itself
In the eye-tracking study, 32% went to technical signals and 10% to social signals, and participants under-reported the social share.

The practical consequence is that "I only judge the code" is factually false for most reviewers. The mechanism is habitual signal-scanning that people do not notice, which is why hiding the author profile changes outcomes. Here is the measured allocation.

Signal classMeanMedian
Code57.15%64.23%
Technical32.42%28.45%
Social10.43%7.38%

Column source: eye-tracking study means and medians as published by denaeford.me and ResearchGate 346358689.

Read this table as a budget you are already spending whether you plan to or not. The goal of a disciplined review is to move the technical and social time onto things you chose to look at, and to keep the social fixations from leaking into your competence judgment. The rest of this reference is built to make that budget deliberate.

The in-PR signal rows: what each proves and how it lies

Each signal below is a thing you can point at inside a diff or a review thread. For each, the first clause is what it proves, and the second is what it looks like when it misleads you.

Commit granularity and messages

An atomic commit is the smallest meaningful change: it does exactly one thing and nothing more. A pull request built from atomic commits with intent-bearing messages proves the author thinks in discrete, reviewable units and can explain why, not just what. Good messages tell you why something changed; the what is already visible in the diff. The dominant named standard is Conventional Commits, formatted as <type>[optional scope]: <description>, a blank line, an optional body, and optional footers, with subjects in the imperative present tense ("add" not "added"), lowercase, no trailing period, under 72 and ideally under 50 characters.

How it lies: a clean single merge commit can be a squash that erased messy work-in-progress commits. One focused message beats many "fix typo" commits, but one dump commit can also be laziness dressed as tidiness. Read granularity on the branch, not the merge.

Test changes

A diff that adds or modifies tests alongside the logic proves the author treats behavior as something to pin down, not just produce. The strong version changes tests because the behavior changed and leaves a trail of what the new behavior is.

How it lies: tests that assert the implementation rather than the behavior, or tests added only to clear a coverage gate, read as rigor while proving little. Check whether a test would fail if the logic were subtly wrong.

Core-file handling: state, auth, money, concurrency

Core files are new logic, changed behavior, and anything touching state, money, authentication, or concurrency. How the candidate handles these proves engineering level more than any other signal, because this is where mistakes are expensive and irreversible. GitLab encodes this organizationally: database migrations or changes to expensive queries must be approved by a database maintainer, and authentication changes by the authentication team.

How it lies: core judgment can hide in a file you triaged as mechanical. A rename-only file with a logic edit buried inside is a classic.

File count and scoping

A tightly scoped pull request proves the author can carve a change into a reviewable unit. Review difficulty rises with the number of changed files, not changed lines; a "big" pull request in this sense means 20 or more changed files. Strong engineers submit database migrations as separate, dedicated pull requests before the application code that uses them, the expand-and-contract pattern.

How it lies: some cross-cutting changes legitimately span many files, and a low line count can still be a 25-file sprawl. Check whether migrations and config were split out before you read the file count as weak design.

Error handling and types

Explicit handling at boundaries proves the author anticipates failure. In TypeScript, the any type should be used sparingly because it removes static typing; unknown is the safer pattern because it forces a type check before use, and strict: true flags every any. The practitioner split is explicit types at function and method signatures, inference inside bodies.

How it lies: dense type annotations on inferable local variables read as thoroughness but are considered poor style. Over-annotation is fake rigor. Check whether annotations sit at boundaries or clutter bodies.

Review-thread responses

How the candidate responds to challenge proves judgment under pushback better than any clean diff can. The strong signal is revising when the reviewer is right, defending with a reason when the reviewer is wrong, and naming the tradeoff either way.

How it lies: agreeableness is not the same as judgment. An author who accepts every suggestion may lack conviction; one who argues every point may lack it in a different direction. Cite one specific exchange, not a vibe.

The author-label traps: seniority, AI, and the social halo

Identical code is rated differently depending on what you believe about who wrote it. This is the most dangerous category because it feels like judgment and is actually bias. The effect sizes below come from two named studies on identical code.

ConditionCompetence penalty
AI-use disclosed, overall9% lower
AI-use disclosed, men6% lower
AI-use disclosed, women13% lower
Seniority label (AI-normalized org)significant, magnitude not stated
AI disclosure (AI-normalized org)no detected bias

Column source: Gai/Hou/Tu (SSRN 5255039) and Microsoft Research VLHCC 2026.

Two things matter here. First, in the pre-registered experiment of 1,026 engineers, the bias moved competence ratings while quality ratings for the same code stayed flat. People judged the person, not the product, when author information was visible. Second, the two studies disagree on AI: in an AI-normalized org of 447 engineers, the AI-disclosure badge produced no detected penalty while a seniority label still biased both perceived competence and effectiveness. So AI-label bias is context-dependent, but seniority bias showed up in both settings.

The bias is not in the code. It is in what you believe about the hands that wrote it.

There is a subtler version. One eye-tracking study found that reviewers looked differently at novice versus senior authors yet accepted their pull requests at the same rate. The danger there is not always a wrong decision; it is miscalibrated confidence. You may reach the right verdict while trusting it for the wrong reason, which fails you on the next, closer call.

The defense is procedural, not willpower. You cannot decide to ignore an avatar you have already fixated on. You can decide the order in which you look at things, which is why the procedure below records label exposure as an explicit step and runs scoring first.

Sourcing enough qualified pull requests to read is its own job, and the pool is larger than most reviewers assume. In Refolk's index of professional profiles there are 178,995 TypeScript-skilled professionals in the US alone, which is why a specific, signal-based query beats scrolling a contribution graph. I built the index so you can ask for the exact profile in plain English and spend your time inside the diffs that matter.

How this goes wrong: the documented failure modes

These are the ways a competent reviewer still reaches a wrong read. Each has a check you can run.

Where a PR signal can fool you

Proves a lotProves little
Noisy commits, sprawling files
Read as weak design, usually correct
Deep core-file judgment, scoped diff
The signal you came for, trust it
Rough WIP on a hard problem
Look closer before rejecting
Clean squash, polished clone, dense types
Verify before you reward it
Looks weakLooks strong
The dangerous cell is top-right: a signal that looks strong but proves little, which is where most bad hires hide.

Squashed-history false positive. A single clean merge commit reads as senior hygiene but can hide messy work-in-progress commits and even record the merger, not the author, as the commit author. The check: read the source branch or fork and the pull request timeline, not the mainline commit. Where you have the repository locally, git reflog plus git reset --hard on a recovery branch restores pre-rebase history, and git show on the squashed SHA exposes the individual file blobs.

Seniority-label halo. Identical code scores more competent when the author looks senior. The check: score the diff before you open the profile, and write the score down so you cannot revise it silently.

AI-disclosure penalty. A visible AI badge can drag your read of competent code downward, hardest on women. The check: hold your reaction against the measured baseline of 9% overall, 13% for women, 6% for men, and judge the product rather than the method.

Social-signal leakage. You believe you ignore avatars and follower counts, but you fixate on them more than you recall. The check: re-review with the profile hidden and see whether your verdict moves.

Tutorial or clone repositories. A polished repository can be a followed-along tutorial, not original judgment; the diff looks competent and proves little. This is a heuristic, not a measured rate. The check: commit cadence over time, issue engagement, and whether any design decision was contested in a thread.

"Mechanical" file hiding logic. A rename-only or format-only file can carry a buried behavior change. The check: open every mechanical-tagged file and confirm it truly changes nothing that runs.

Many-files-touched misread. A 20-plus file pull request can look ambitious but usually signals weak scoping, while some cross-cutting changes legitimately span files. The check: whether migrations and config were split into separate dedicated pull requests.

Over-annotation as fake rigor. Dense explicit types on inferable local variables read as thoroughness but are considered poor style. The check: whether annotations sit at boundaries or clutter function bodies.

The review procedure, in order

Run these eight steps per candidate. The order matters: file triage comes before deep reading so you spend your 30 to 60 minutes on the files that carry the decision, and label-recording runs before scoring so you can discount bias rather than absorb it.

Reading a candidate's public pull requests

  1. Read description and commits first
    Pull the public pull requests and read the description and commit list before opening a file. You should be able to state the claimed intent before you see any diff.
  2. Triage files into core, supporting, mechanical
    Sort changed files into core (logic, state, money, auth, concurrency), supporting (call-sites, types, wiring), and mechanical (lockfiles, generated code). Isolate the handful that carry the decision.
  3. Review core files in passes
    Read the core files in three passes, correctness then design then style. You should be able to name one real design decision and say whether it is sound.
  4. Inspect commit granularity and messages
    Check each commit for atomicity and whether the message explains intent. Classify every commit as scoped-with-intent or noise.
  5. Recover evidence if history is squashed
    If the merge is one squashed commit, recover per-commit evidence from the branch, fork, or timeline. Produce authorship and granularity, or mark the history unrecoverable.
  6. Read the review thread
    Read how the candidate responded to challenge, revised, and explained tradeoffs. Cite one exchange that shows judgment under pushback.
  7. Record author-label exposure
    Note which labels you saw: seniority, an AI badge, the name and avatar. Write them down so you can discount known bias before scoring.
  8. Verify mechanical files are mechanical
    Open each mechanical-tagged file and confirm no logic edit hides inside a rename-only or format-only change.

From public PRs to a defensible verdict

  1. Candidate public PRs
    57%

    Where your attention naturally lands, on the code

  2. Technical signals read
    32%

    Commits, history, prior contributions

  3. Social signals read
    10%

    Profile, label, avatar, to be discounted

Most of the pull requests you open will not reach a scored verdict, and that is the point of triage.

The funnel percentages are the attention split from the eye-tracking study, reused here as a reminder of where your time goes by default. Your job in the procedure is to make each layer deliberate rather than habitual.

Calibrating against pool scale before you set the bar

Before you decide what a strong pull request looks like, know how many candidates can clear whatever bar you set, because an unrealistic bar is a sourcing failure disguised as a standards failure. The pool differs sharply by stack and geography.

SegmentCountDerived ratio
Senior Software Engineer, US181,65311.9x Germany
Senior Software Engineer, Germany15,272baseline
TypeScript skill, US178,99510.6x Go
Go skill, US16,935baseline

Column source: counts from Refolk's index; ratios derived from those counts.

10.6x
US TypeScript-skilled professionals versus Go-skilled
178,995 TypeScript to 16,935 Go in Refolk's index, which is why a Go-contribution requirement narrows the pool far faster than a TypeScript one.

Read this as calibration, not as a quota. If you require Go contributions with database migration commits, you are fishing in a pool roughly a tenth the size of the TypeScript one, so your per-signal bar should bend toward evidence of judgment rather than volume of output. Where the stack is deep and plentiful, you can afford to weight scoping and review-thread discipline heavily; where it is thin, a single well-handled core file may be the strongest signal available.

What to verify before you call the read done

Run this before the verdict goes into the loop. A read that skips these is a read you cannot defend when someone asks why.

Pre-verdict checklist

  • I read the description and commit list before I opened any file.
  • I triaged files into core, supporting, and mechanical, and I reviewed the core deeply.
  • I can name one real design decision in the change and say whether it is sound.
  • I classified every commit as scoped-with-intent or noise.
  • If the history was squashed, I recovered granularity and authorship from the branch or timeline, or I marked it unrecoverable.
  • I can cite one review-thread exchange showing judgment under pushback.
  • I wrote down which author labels were visible before I scored the diff.
  • I opened every mechanical-tagged file and confirmed no logic hides inside.
  • I checked whether migrations and config were split out before reading file count as weak scoping.

Keeping this reference current

Treat the signal rows as stable and the effect sizes as living. The procedure, the core-versus-mechanical triage, and the failure modes are grounded in how review works and will not drift. The bias magnitudes will: the AI-disclosure penalty is already context-dependent, producing a measurable 9% hit in one org and no detected effect in an AI-normalized one. Re-check the direction of that effect in your own setting rather than importing a fixed number, because whether a visible AI badge biases your reviewers depends on how normal AI use already is around them.

The way to keep it current is to run the blind-score test on your own loop. Have two reviewers score the same diff, one with the profile visible and one with it hidden, and compare. If their competence ratings diverge while their quality ratings agree, you have reproduced the label halo locally and you know your procedure needs the hide-first discipline more than your reviewers believe. Do that once a quarter and this reference stays honest about the one thing it cannot measure for you: your own team.

When you are ready to source the pull requests worth reading rather than scrolling contribution graphs, ask Refolk for the exact signal, name a real repository, a real skill, and a real place, and spend your time inside the diffs.

Questions practitioners ask

How do I judge seniority from commits without being fooled by a squash?

Read commit granularity and message intent on the source branch or fork, not the squashed mainline commit. A squash merge collapses history to one commit and can even record the merger rather than the author. If you only have the merged commit, mark granularity and authorship as unrecoverable rather than reading the clean single commit as discipline, because it may hide messy work-in-progress commits underneath.

Should a visible AI-assistance badge lower my rating of a candidate's code?

No. In a pre-registered experiment with 1,026 engineers, identical code was rated 9% less competent when the author was labeled as using AI, with women penalized 13% versus 6% for men. The code quality was the same; only the competence judgment moved. Judge the product, not the method, and treat the badge as a known bias to discount, not evidence.

What in a pull request actually proves engineering level versus just looking good?

The core files prove the most: how the candidate handles state, auth, money, and concurrency, and whether one real design decision is sound. Commit atomicity and intent-bearing messages prove discipline. A polished repository, many files touched, or dense type annotations prove little on their own and often mislead. Score the diff before the profile to avoid the seniority halo.

Is touching many files a good or bad sign in a candidate's PR?

Usually bad. Review difficulty rises with file count, not line count, and a pull request touching 20 or more files typically signals weak scoping. The exception is a legitimate cross-cutting change, but strong engineers split migrations and config into separate dedicated pull requests before the application code that uses them. Check whether that split happened.

How can I tell a real contribution from a tutorial clone?

A polished repository can be a followed-along tutorial rather than original judgment, and the diff will look competent while proving little. Check commit cadence over time, whether the author engaged with issues, and whether any design decision was contested in a review thread. This is a heuristic, not a measured rate, so treat a clean clone as unproven rather than disqualifying.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next