The AI-Authorship Discount: Reading Whose Skill a Repo Proves
You will be able to grade a candidate's public repos on authorship dimensions and output own-skill, AI-assisted, or inflated, with a documented reason each.
You are about to spend an interview slot, and a candidate's GitHub looks strong. The job this guide covers is deciding, before that slot is booked, how much of their public repository work reflects their own engineering ability versus AI generation. It is written for engineering managers, technical founders, developer-relations leads, and technical sourcers who build shortlists from public signals. By the end you will be able to score a candidate's public code on authorship-confidence dimensions and reach a defensible verdict - own skill, AI-assisted, or inflated - with a documented reason for each.
The existing public answers do not do this. They are either live-interview behavior checklists, which help only after you have already committed the slot, or generic AI-code detectors that admit low reliability and output a true/false flag instead of a staffing decision. Neither helps a sourcer grade a static public repo at list-building time. This guide turns the scattered tells into a scored, repeatable read.
Why clean code stopped proving skill
The short answer: AI coding tools made junior and mid-level code syntactically indistinguishable from senior code, so a clean diff is now commodity and no longer discriminates. The signal moved to the conversation around the code.
This is the shift that breaks the old sourcing instinct. For years, a tidy repo with sensible naming and real tests was a reliable proxy for a competent engineer. That proxy is dead. GitHub's developer survey reports that 92% of U.S. developers now use AI assistants, and Copilot generates roughly 46% of code written by its users, rising to 61% for Java developers. When half the code in a repo can be generated and the output reads clean regardless of who prompted it, "the code looks good" tells you almost nothing about the author.
What it leaves behind is a question of authorship confidence: not "is this AI" as a binary, but "how much of this work demonstrates skill I can interview against". That reframing matters because the goal is a staffing decision, not a detector verdict. A heavily AI-assisted repo from an engineer who clearly reasons through trade-offs in their issue threads is a strong candidate. A pristine repo from someone who never explains a single decision is a risk, regardless of whether a detector lights up.
The discriminating evidence has moved to where AI is weakest: design reasoning. The comment that explains why a pattern is dangerous, and the question that forces the author to reason through failure modes they had not considered, is senior engineering that AI cannot fake. That is the spine of this read.
The authorship-confidence dimensions
There are five dimensions worth scoring. Each proves something specific, and each has a characteristic way of lying to you - a false positive you must guard against.
| Dimension | What a strong score looks like | What it proves |
|---|---|---|
| Commit history | Incremental commits, descriptive messages, dated iteration | The work was built, not pasted |
| Code-level tells | Varied comments, consistent naming, lean deps, real tests | Line-level authorship discipline |
| Architectural coherence | One problem solved one way across files | Sustained design thinking |
| Review conversation | Trade-off narration, failure-mode questions | Senior reasoning AI cannot fake |
| Refactor/reuse evidence | Moved and refactored lines, deduplication | Ownership of the codebase over time |
Score each from strong to weak, then read the pattern rather than any single number. The order matters: the review conversation and refactor evidence are the durable positive signals, so a weak score there pulls a verdict down harder than a weak score on code-level tells, which a disciplined junior can also trip.
Here is the trap in plain terms. Clean code used to be the headline dimension. It is now the least diagnostic one, because it is exactly what AI produces well. Weight your verdict toward the bottom two rows of that table.
Where authorship evidence lives, strongest at the base
- README and polishEasiest to generate, lowest signal value
- Code-level tellsComment style, naming, dependencies, tests
- Commit and refactor historyHow the code came to exist over time
- Review conversationTrade-off reasoning AI cannot fake
Commit history and the one-way trailer
Start with the git log, because history is harder to fake than a snapshot. Human developers usually work incrementally with a rich commit history and descriptive messages, whereas AI-authored repos often show fewer commits, large chunks committed at once, and vague messages like "Updated files."
The most tempting shortcut here is the attribution trailer, and it is the most dangerous. Claude Code appends Co-Authored-By: Claude <noreply@anthropic.com> to commits and a "Generated with Claude Code" line to PR bodies as a default. Running git log --grep="Co-Authored-By: Claude" surfaces every AI-assisted commit that still carries it. The problem is the word "still".
So treat the trailer as a positive confirmation only, never as an acquittal. When it is missing, your evidence is the shape of the history: squashed or rewritten commits, a cluster of giant commits, no visible iteration. Note that these history-pattern tells are documented but weaker than the trailer, and several of them have innocent explanations covered in the failure-modes section.
What you want out of this step is a timeline. Many strong candidates have a repository history that predates widespread AI adoption and shows incremental, debugging-style commits in that era. That early iteration is positive evidence of hand-built skill, and it is one of the few things a sourcer can read directly from a public profile.
Reading code-level tells across files
Scan for the recurring, documented tells, and count them. One tell is a hunch; four or five together is a diagnosis. No single tell is proof, which is exactly why this is a scored read and not a detector.
The named tells, from most to least robust:
- Architectural incoherence. The same problem solved several different ways in one repo. AI code is uniform line by line and wildly inconsistent in architecture, because every prompt session reinvented the approach from scratch. This is the most robust tell because it is a workflow artifact, not a writing style you can clean up.
- Unused dependencies. Imported but never called. A classic byproduct of pasted suggestions.
- Comments that explain "what" not "why". AI comments describe what the code does rather than why a decision was made, are evenly distributed rather than clustered around complex sections, and often use comment-as-section-header patterns that humans rarely write.
- Generic identifiers.
data,result,temp. One study found AI introduced nearly 2x more naming inconsistencies, with unclear naming and mismatched terminology appearing frequently in AI-generated changes. - Tests that assert nothing. Present for coverage theatre, verifying nothing real.
- Duplicated blocks. Copy-paste debt rather than reuse.
That last one is measurable at scale, and it tells you what inflated code actually costs. GitClear's analysis of code composition shows the debt signature clearly.
| Metric | Baseline | 2024 | Direction |
|---|---|---|---|
| Moved/refactored lines | 24.1% (2020) | 9.5% | down ~60% |
| Copy/pasted lines | 8.3% | 12.3% | up 48% |
| Churn (revised in <2 wks) | 3.1% | 5.7% | up ~84% |
| Duplicate blocks | 1x | ~8x | eightfold |
The eightfold jump in duplicated blocks and the collapse in refactored lines mean an inflated repo reads as "works" while carrying copy-paste debt. This is why refactoring and reuse survive as positive skill evidence: a candidate whose history shows them moving and consolidating code is demonstrating ownership that generation does not produce.
An inflated repo reads as it works while quietly carrying the copy-paste debt that grading has to weigh against it.
The procedure
Work the dimensions in order. The sequence is deliberate: human reading comes before any tool, because a detector flag should start a review, not end one. Budget about an hour per candidate for a serious read, less once you have the pattern.
Grading a candidate's public code for authorship confidence
- Pull the public surfaceGather the candidate's repos, pinned projects, pull requests, and issue comments. Discard forks and tutorials so you have two or three substantive repos to read.
- Scan commit history and trailersEyeball or grep the git log for commit size, cadence, message quality, and AI co-author trailers. Date the before/after of AI adoption and note whether trailers are present.
- Read code-level tells across filesCheck comment density and style, naming consistency, unused dependencies, test quality, and architectural coherence. Count the tells; four or five together is a diagnosis.
- Run a detector as triage onlyOptionally scan with a code detector to point you at files, never to deliver a verdict. Treat any score as a pointer to re-read.
- Read the conversationReview issue threads, PR descriptions, and review comments for trade-off reasoning and failure-mode thinking. Quote one comment that shows reasoning AI cannot fake, or confirm none exists.
- Score each dimension and write a verdictClassify as own-skill, AI-assisted, or inflated with a one-line reason per dimension, weighting conversation above clean diffs.
There is a genuine order disagreement in the field worth naming. Some playbooks run tools first, then manual review. The academic-integrity sources insist a flag starts a review rather than ending it, which is why I put detectors after human reading, not before. If you run a detector first, you anchor on its score and read the code to confirm it. Read first, and the detector becomes what it should be: a second opinion on files you have already formed a view about.
Why detectors go last and stay in triage
Detectors do not attribute code to a person, and their accuracy evidence does not survive scrutiny the way vendors claim. Use them to point at files, never to decide.
Vendor claims run high. One free tool advertises 90%+ accuracy across Python, JavaScript, PHP, C/C++, and Java. The independent evidence is thin, and the closest rigorous analog is damning. The most cited text-detector study flagged 61% of essays by non-native writers as AI, against about 5% for native speakers. The statistical distributions of human and AI output overlap in feature space, which makes perfect separation mathematically impossible. OpenAI shut down its own text classifier citing low accuracy.
Two consequences matter for sourcing specifically.
The second consequence is a fairness and quality problem at once. The 61%-versus-5% bias against non-native writers maps directly onto a global candidate pool, so a detector-led screen systematically down-ranks international engineers. That is the opposite of what a sourcer wants, and it quietly shrinks your pipeline in exactly the markets where talent is cheapest to reach. If you want to source the strong non-native engineers that a detector would wrongly flag, grade them on history and review conversation and skip the detector entirely.
Code-specific tools do exist and can help at the triage step. An open-source scanner like ai-gen-code-search highlights candidate files and line ranges, but it does not connect findings to a specific tool, model, or developer session. That is the honest ceiling: it tells you where to look, not whose skill you are looking at.
One more thing you cannot do from a public repo: read acceptance rates. GitHub's telemetry across 934,533 Copilot users shows suggestions accepted around 30% of the time, but that data comes from the Metrics API against seats an employer owns. A sourcer reading a public repo has none of it, and even if you did, acceptance percentage measures adoption, not authorship or quality. Rewriting code based on AI inspiration is not captured, because GitHub does not capture intent through telemetry.
| Source | Acceptance rate | Note |
|---|---|---|
| GitHub telemetry, 934k users | 30% | rises with use |
| GitHub/Accenture enterprise | ~30% | enterprise seats |
| ACM CACM study | 27% | peer-reviewed |
| Hivel sample payload | 28.27% | 1,640 of 5,800 lines |
The point of this table is not to help you estimate anything. It is to show that even the people who can measure acceptance land near 30%, and that number says nothing about skill. Do not let acceptance-rate intuition sneak into your read.
Reading the conversation, the signal that survives
This is the step that decides most verdicts. Written review comments from a candidate provide a strong signal of their insight and collaborative capabilities, and it is the one dimension AI cannot generate on their behalf.
Go to the candidate's issue threads, pull request descriptions, and review comments on other people's code. You are looking for three things: trade-off narration (why this approach over that one), failure-mode thinking (what breaks and when), and incremental debugging (the back-and-forth of actually fixing something). The comment that explains why a pattern is dangerous, and the question that forces reasoning through failure modes, is senior engineering.
There is a quantified reason this matters. AI-generated code takes 12% longer to review than human-written code and draws 41% more comments. A candidate whose public history shows them doing that reviewing - catching the subtle breakage, asking the hard question on someone else's PR - is demonstrating the scarce skill, not the commodity one. Review is now arguably the most important skill an engineer has, and it is visible on public profiles if you go looking.
The staffing verdict in two variables
The matrix captures the central judgement. Strong code plus thin conversation is the classic inflated profile and the one that burns interview slots. Strong conversation is what rescues even a heavily AI-assisted candidate, because the reasoning is theirs. When you can quote one comment that shows reasoning AI cannot fake, write it into the verdict verbatim - it is the most defensible line in the whole read.
How this read goes wrong
The failure modes here are false positives in both directions: scoring a capable human as inflated, and scoring an inflated repo as skill. Each has a specific check.
- Trailer absence read as "human". Teams strip trailers by default, and attribution appears intermittently. Check: look for squashed or rewritten history, not just a clean grep.
- Clean code scored as skill. A junior with Copilot ships staff-looking code. Check: read the conversation, not the diff.
- Over-commenting flagged on a disciplined junior. Some humans genuinely over-document. Check: is the comment style uniform across unrelated files and sessions? One tell alone is a hunch.
- Detector score treated as proof. The base-rate trap turns a rare true positive into a flood of false ones. Check: corroborate with history evidence before it counts.
- "Few giant commits" over-weighted. Squash-merge workflows and bulk imports produce the same shape legitimately. Check: inspect PR-level granularity, not just the main-branch log.
- README polish mistaken for repo quality. AI repos often have READMEs too clean for a personal project. Check: does the code and test depth match what the README claims?
- Acceptance-rate intuition misapplied. Acceptance percentage is adoption, not authorship, and you cannot derive it externally anyway. Check: ignore it.
The through-line in every one of these is corroboration. A single signal, in either direction, is never the verdict. The method works because the signals are cheap to gather and the mistakes are predictable, so you can design the check before you need it.
Before you commit the interview slot
- You read two or three substantive repos, not forks or tutorials.
- You dated the before/after of AI adoption from the commit history.
- You counted code-level tells rather than reacting to one.
- You confirmed architectural coherence across files, not just within one.
- You checked refactor and reuse evidence, not just that the code works.
- You can quote one review comment that shows reasoning, or confirmed none exists.
- Any detector score was treated as a pointer and corroborated, not trusted.
- You wrote a one-line reason for each dimension and a single verdict word.
Keeping the read current
The tells in this guide are stable in shape but not in detail, because the tools change underneath them. Re-check three things on a cadence rather than trusting a value you learned once.
First, the attribution defaults. The Claude Code trailer behavior already moved from a setting to an attribution object, and such defaults will keep shifting. Whenever a major tool changes how it marks commits, your grep-based confirmations need updating. Treat any trailer convention as a current fact with a short shelf life.
Second, the base rates. Adoption is near-universal already at 92% of U.S. developers, which means the useful question is no longer "did they use AI" but "how well do they reason", and that only gets truer. As generation improves, weight the conversation dimension even harder.
Third, your own source of candidates. The whole read assumes you are starting from people whose public history actually shows the durable signals - years of commits, real review threads - rather than from a snapshot. Building that starting list by hand is the slow part. Describing the engineer you want in plain language and getting back people with substantive review conversation and incremental histories is exactly the friction Refolk removes, so you spend your hour grading candidates worth grading. In Refolk's index, the US pool of engineers listing GitHub and Open Source as skills is 16.8x larger than Germany's 1,042, so the market you search in shapes the shortlist before any grading begins.
The verdict you write is only as good as its reasons. Own-skill, AI-assisted, and inflated are all acceptable outcomes to advance or pass on - what is not acceptable is a verdict without a one-line reason per dimension. Write those reasons down. They are what makes the call defensible when a hiring manager asks why you spent, or did not spend, the slot.
Questions practitioners ask
Is this GitHub code AI generated?
You can rarely prove it from a single file, and you should not try to. The honest read is probabilistic: count corroborating tells like uniform comment density, generic identifiers, unused imports, and architecture that solves the same problem several ways across files. One tell is a hunch; four or five together is a diagnosis. The most robust single tell is architectural incoherence, because it is a workflow artifact of per-prompt sessions rather than a writing style you can clean up.
Does the absence of a Co-Authored-By: Claude trailer mean the code is human?
No. The trailer's presence confirms AI involvement, but its absence proves nothing. The includeCoAuthoredBy setting is deprecated, teams strip the trailer by default, and attribution may still appear intermittently despite settings. Read absence as silence, not as evidence of human authorship, and corroborate with squashed or rewritten history instead of relying on grep.
Can I trust an AI code detector to screen candidates?
Use it as triage, never as a verdict. OpenAI shut down its own text classifier citing low accuracy, and the most cited detector study flagged 61% of non-native-writer essays as AI against about 5% for native speakers. That bias maps directly onto a global candidate pool, so a detector-led screen systematically down-ranks international engineers. In rare-usage contexts even a 99%-accurate tool produces more false positives than true positives.
If AI makes junior code look senior, what still proves real skill?
The review conversation. AI coding tools made junior and mid-level code syntactically indistinguishable from senior code, so clean diffs lost their signal value. What survives is the comment that explains why a pattern is dangerous and the question that forces reasoning through failure modes. Trade-off narration, incremental debugging commits, and substantive PR review comments are the durable positive evidence. A polished README is not.
Can I see a candidate's Copilot acceptance rate from their public profile?
No. Acceptance-rate data comes from the Copilot Metrics API against seats an employer owns, not from any public repo. GitHub does not capture intent through telemetry, so rewriting code based on AI inspiration is invisible even to the org that owns the seats. Acceptance percentage measures adoption, not authorship or quality, so even if you had it, it would not tell you whose skill the code proves.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.