The Agent-and-Bot Commit Signal Reference: What Each Marker Proves
You can scan a mined commit batch and tag each author as human, agent-assisted, or bot using the exact git fields, trailers, and noreply patterns.
Key takeaways
- Bot-account lookup alone recovered only 28,154 of 850,157 Claude Code commits in one 180M-repo census snapshot, a 30x relative-recall gap, because agents commit under the human operator's author identity and mark themselves only in a body trailer.
- A subject-only scan using git's %s format returns zero agent commits every time, because trailers live in the message body; the reliable read is %B or the %(trailers) accessor.
- A commit trailer proves tool-touch, not authorship: any human can paste a Co-authored-by line, and Cursor earned real GitHub contributor badges this way, so a trailer can inflate a bot into a graph-counted contributor.
- GitHub's [bot] and @users.noreply.github.com patterns silently miss GitLab identities like project{id}_bot@noreply.gitlab.com and service_account_ prefixes, recounting bots as people.
- Absence of a trailer is not absence of an agent: opt-out settings and CLAUDE.md-configured repos leave silent commits under the human's identity that no single signal catches.
Before you trust a mined contributor list, you need to know which commits a person actually wrote. This is a lookup reference for recruiting and revenue operations: scan a batch of mined commits and tag each author as human, agent-assisted, or bot, using the exact git fields, trailers, and noreply-email patterns that separate them. It also tells you which signals silently miss agent work, so you stop declaring a dataset clean when it is not.
The job here is sourcing hygiene, not code audit. You are not checking whether code is good. You are checking whether a name on a contributor graph represents a human whose skill you can source against, or a machine identity that will waste your outreach.
What each commit signal proves, and what it misses
A commit carries three independent layers of authorship signal, and each one proves something different. The author identity field answers "whose account is credited"; the body trailer answers "what tool was present"; and the repo config answers "could an agent have worked here silently." No single layer is sufficient, and the most common sourcing mistake is reading only one.
Here is the layered structure to hold in your head before you scan anything.
The three layers of commit authorship signal
- Repo config (CLAUDE.md, .cursorrules, AGENTS.md)Proves an agent could have worked here, even with no trailer on any commit
- Commit body trailer (Co-authored-by + noreply email)Proves a tool was present; does not prove who wrote the code
- Author/committer identity (%an, %ae)Proves which account is credited; bots use [bot] suffixes and host noreply forms
The identity field proves attribution but lies by omission: agents that commit under the human operator's account leave the author field looking fully human. The trailer proves tool-touch but lies by overclaiming: any human can paste a Co-authored-by line, and a trailer counts toward the GitHub contribution graph whether or not the named agent wrote anything. The config layer proves opportunity, not act. Read together, they let you assign a defensible confidence. Read alone, each one is a trap.
The agent commit-signature table
Four coding agents produce the signatures you will see most often, and each has a default trailer email that maps to a specific GitHub account. Match on the email, never the display name, because model-name variants like Claude Opus 5.5 <noreply@anthropic.com> will slip past a literal Claude <noreply match.
| Agent | Default trailer email | On by default? | Maps to GH account |
|---|---|---|---|
| Claude Code | noreply@anthropic.com | Yes (opt-out) | claude |
| Cursor | cursoragent@cursor.com | Yes (IDE toggle; CLI/cloud harder) | cursoragent |
| Copilot IDE | copilot@github.com | No (git.addAICoAuthor) | Copilot |
| Copilot SWE agent | <id>+Copilot[bot]@users.noreply.github.com | Yes (bot identity) | copilot[bot] |
Three details change how you treat each row. Claude Code's trailer is on unless the user set includeCoAuthoredBy: false or an empty attribution.commit string, so its presence is reliable but its absence proves nothing. Cursor adds Co-authored-by: Cursor <cursoragent@cursor.com> to every commit and can only be disabled in the IDE; there is no off switch in the CLI, in cloud agents, or behind the "Generate git commit message" button. Copilot's IDE trailer is governed by the VS Code git.addAICoAuthor setting, which has values off, chatAndAgent, and all, with off as the default. The Copilot SWE coding agent is different in kind: it commits under a fixed bot account, so it shows up in the author identity layer, not just the trailer.
One oddity worth logging: the Copilot IDE default is a moving target. The git.addAICoAuthor default was flipped to all in VS Code 1.117, then reverted to chatAndAgent in 1.118. Attribution rate in the same repo therefore depends on contributor tool versions, not on the code, so do not treat a repo's trailer density as a fixed property.
Why you must read the body, not the subject
Trailers live in the commit message body, never the subject line, so a scan that reads only the subject returns zero agent commits every time. This is the single cheapest way to produce a false-clean dataset, and it is worth internalising before anything else.
Git's %s format prints only the first line of a commit message. The Co-authored-by trailer sits several lines down in the body, so %s never sees it. The reliable reads are %b (body), %B (raw body including subject), or the %(trailers) accessor, which displays the trailers of the body as interpreted by git-interpret-trailers. Trailers landed as a first-class git feature in 2.32.0, so the accessor is available on any modern install.
# Full body per commit, pipe-delimited header plus trailers git log --format="%H|%an|%ae|%(trailers:key=Co-authored-by)" # Or raw body for grep across the whole message git log --format="%B" | grep -i "co-authored" # Strict trailer parse on a single commit git log -1 --pretty=format:%B | git interpret-trailers --parse
Run from the repo root. Swap the key filter if you only want co-author trailers.
Both approaches are valid. Some practitioners prefer piping %B into git interpret-trailers --parse because it applies git's own trailer grammar rather than a loose grep; others use %(trailers:key=...) for a one-pass extract. Pick one and use it consistently, because mixing them across a batch produces inconsistent tagging.
The 30x recall gap: why bot lookup alone fails
Looking up bot accounts is grossly incomplete, and the gap is structural rather than a sampling artefact. In a census across 180 million repositories, multi-method detection identified 850,157 Claude Code commits in one snapshot, of which bot-account lookup recovered only 28,154: a 30x relative-recall gap.
The mechanism is simple once you see it. Agents like Claude Code commit under the human operator's author identity and mark themselves only in a body trailer. Any detector that reads the author field alone sees a person. So the 30x gap is not noise you can average away; it is the direct consequence of where the agent writes its signature. If you want recall, you must read the body and the config, not just the account.
The channel split compounds this. A pull-request-based census misses 79% of commit-detected Claude Code adopters, and a commit-based census misses essentially all Codex adopters, because Codex is the largest agent by pull requests but near-absent from commits. Commit data and PR data see nearly disjoint agent populations. A sourcing dataset built from one channel structurally misses the other's people.
Claude Code commit recall by method, one census snapshot
- 850,157Multi-method detection
full detected set
- 28,154Bot-account lookup only
3.3% of the full set
For scale on how fast this population is growing: Claude Code led with 886,122 commits across 17,295 projects, followed by Jules with 215,804, and commit-attributed agents collectively generate over 320,000 commits per month by the V2604 snapshot. This is not a rounding error in a contributor list. It is a growing slice of every mined engineering pool.
Bot lookup finds the 3.3 percent that announce themselves and declares the rest human.
The silent-commit blind spot no single signal catches
The hardest case is the commit that carries no agent signal at all: a human-identity author, no trailer, written with an agent that was configured to stay quiet. No field on the commit will catch it, which is why you need the config layer and a known-unknown column.
The census taxonomy names these explicitly. Type C agents are distributed under a human identity; Type D agents are silent. Trailer suppression is the common cause: Claude Code's includeCoAuthoredBy: false or an empty attribution string drops the trailer, and Cursor's IDE toggle does nothing to its CLI and cloud runs, so a repo you believe is "disabled" still emits agent work from other surfaces.
What you can detect is opportunity. Config files are a documented adoption proxy: a repo containing CLAUDE.md, .cursorrules, AGENTS.md, or a .github/copilot path is a repo where silent agent work is plausible. You cannot prove any specific commit was agent-written, but you can refuse to call it clean.
When you source engineers by their public commit footprint, the silent-commit problem is exactly why a tidy raw list can mislead. Refolk returns the pool from public GitHub and web signals; this reference is what you apply to that pool before you treat any contributor name as a verified human. Refolk gives you the people; the config-and-trailer scan tells you which of them actually wrote the code you are crediting.
The procedure: tag a commit batch
Run these seven steps in order on a mined commit batch. The whole sequence is minutes per repo, and it moves you from a raw author list to a dataset where every name carries a human, agent-assisted, bot, or cannot-confirm tag.
From raw commit batch to tagged authorship
- Dump bodies, not subjectsRun git log with --format=%B or --format="%H|%an|%ae|%(trailers:key=Co-authored-by)" so every commit's full body and trailer block is captured, not just the subject line.
- Match known agent noreply emailsGrep the trailer block for noreply@anthropic.com, cursoragent@cursor.com, copilot@github.com, and Copilot[bot]@users.noreply.github.com, then tag each trailer author agent or human.
- Resolve author and committer identityCheck the %ae author field for [bot] suffixes and host noreply forms across GitHub, GitLab, and Gitea, separating bot-identity commits from human-identity commits.
- Scan corroborating footers and session linksLook for Generated with lines, Claude-Session links, and Copilot session-log URLs in the body or PR, raising or lowering per-commit confidence.
- Scan repo config artifactsFlag repos containing CLAUDE.md, .cursorrules, AGENTS.md, or .github/copilot so silent-agent risk is recorded even where commits carry no trailer.
- Apply the silent-commit caveatMark human-identity commits in agent-configured repos as cannot confirm human, so the dataset carries a known-unknown column rather than a false clean.
- Normalise non-GitHub hostsRe-run bot patterns with GitLab and Gitea noreply forms before counting contributors, so no project_*_bot@noreply address is recounted as a person.
Corroborating markers beyond the trailer
When a trailer is present but you want more confidence, three corroborating markers tell you whether an agent genuinely ran rather than someone pasting a line by hand. Each raises confidence; none is required.
- PR footer. Claude Code adds a
Generated with Claude Codeline to pull request bodies by default, separate from the commit trailer. - Session links. Copilot cloud commits include a link to the agent session logs, a permanent link from the commit to the full session. Claude Code adds a
Claude-Session:link on commits from Remote Control or cloud sessions, controllable bysessionUrl: false. - Config files. The presence of
CLAUDE.md,.cursorrules, orAGENTS.mdcorroborates agent use at the repo level even when a given commit is silent.
A session link is the strongest single corroborator because it points at a real run that produced the commit, which a hand-pasted trailer cannot fake. Treat a trailer plus a live session link as agent-assisted with high confidence, and a bare trailer with no corroboration as tool-touch only.
How these signals go wrong
Every signal in this reference lies under specific conditions, and the failures all point the same way: toward a false-clean dataset where machine authorship reads as human. Work this table before you ship a tagged list.
| Failure mode | What it does | The fix |
|---|---|---|
| Subject-only scan (%s) | Reports zero agent commits; false clean | Read %B or %(trailers) |
| Bot-lookup only | Tags 3.3%, calls the rest human | Add body and config scans |
| [bot] match on GitLab | Recounts project_*_bot as people | Add host-specific regex |
| Trailer treated as proof | Credits hand-pasted or unlinked trailers | Corroborate with session link |
| Config/opt-out suppression | Silent commits read as human | Carry a cannot-confirm column |
Four of these deserve more than a row. Trusting the co-author as proof of AI overclaims in both directions: a human can paste a Co-authored-by line, and the deprecated legacy trailer format does not link to a profile at all, so it shows up as a grey ghost contributor with no avatar and no graph credit. The claude-code-action tool shipped exactly this legacy format at one point, unlinking its own commits from the contribution graph.
Host monoculture is the quiet one. Detectors keyed to [bot] and @users.noreply.github.com encode GitHub assumptions. GitLab bot identities carry no [bot] suffix: project bots use project{id}_bot@noreply.gitlab.com, group tokens use group_{id}_bot_{random}@noreply.{host}, and service accounts use a service_account_ prefix and are always external users. Self-hosted instances can use example.com-style domains. Every one of these passes a GitHub-shaped detector as a human.
Squash-merge author rewrite corrupts blame directly: Copilot can become the squash author of a merged branch, inflating a bot as the primary contributor even where humans did the work. And model-name variants defeat name-based matching, which is why the rule above is to key on the email.
Before you trust the list
Run this checklist against your tagged batch before you treat any contributor as a human you can source. It is the condensed form of everything above.
Pre-trust verification
- Commit bodies were read with %B or %(trailers), not %s
- Every trailer matched on noreply email, not display name
- Author %ae field was checked for [bot] suffixes
- GitLab and Gitea noreply forms were run as separate regex
- Session links and Generated with footers were scanned as corroboration
- Repos with CLAUDE.md, .cursorrules, or AGENTS.md are flagged
- Human-identity commits in agent-configured repos carry a cannot-confirm tag
- Squash-merged branches were checked for author rewrite
Keeping the reference current and sizing the pool
Two of these signals move, so re-check the mechanism rather than trusting a cached value. The Copilot IDE default (git.addAICoAuthor) has flipped between releases, so verify the current default against the VS Code setting rather than assuming. Agent defaults change with releases too, so when a new agent appears in your batch, find its default trailer email and confirm whether attribution is on or off by omission before you write a rule for it.
The size of the pool this work cleans is worth keeping in view. In Refolk's index of professional profiles, 49,157 US "Software Engineer" profiles list Git as a skill, against 4,652 in Germany, so the US pool is roughly 10.6x larger.
| Market | Profiles | Share of US (derived) |
|---|---|---|
| United States | 49,157 | 1.00x |
| Germany | 4,652 | 0.095x |
Skill depth narrows fast inside that pool. Among US Software Engineers in Refolk's index, 2,866 list Go against 599 listing Rust, so the Go pool is about 4.8x larger.
| Skill | Profiles | Multiple of Rust (derived) |
|---|---|---|
| Rust | 599 | 1.0x |
| Go | 2,866 | 4.8x |
When a pool is this large and this concentrated, a single detector flaw scales. A subject-only scan that silently passes agent commits, or a GitHub-shaped regex that recounts GitLab bots as people, does not misclassify a handful of records; it misclassifies a fixed percentage of tens of thousands. That is the argument for running all seven steps rather than the one that is easy. Source the people with Refolk, then apply this reference so the names you hand to a recruiter are the ones who actually wrote the code.
Questions practitioners ask
How do I detect AI agent commits in git history?
Dump commit bodies with git log --format=%B rather than reading subjects, then match the four known agent noreply emails in the trailer block: noreply@anthropic.com, cursoragent@cursor.com, copilot@github.com, and Copilot[bot]@users.noreply.github.com. Match on the email, not the display name, because model-name variants like Claude Opus 5.5 break literal string matches. Then scan repo config files for silent agents that leave no trailer at all.
What is the Co-authored-by Claude noreply email?
Claude Code appends Co-Authored-By: Claude <noreply@anthropic.com> to commits by default. That email resolves to the claude GitHub account, verified through the GitHub API. The trailer is suppressed only if the user has set includeCoAuthoredBy: false or an empty attribution.commit string, so its absence does not prove a human wrote the commit.
Why does filtering on [bot] miss most agent commits?
Because it is a structural failure, not a sampling one. Agents like Claude Code commit under the human operator's author identity and mark themselves only in a body trailer, so a detector reading the author field alone sees a person. In one census, bot-account lookup recovered only 3.3% of Claude Code commits, a 30x relative-recall gap. You must also scan commit message bodies and repo config files.
How do I exclude GitHub Actions and Copilot bot commits from a contributor list?
Check the %ae author field for the [bot] suffix and the GitHub form <id>+Copilot[bot]@users.noreply.github.com, which the Copilot SWE agent uses as a fixed bot identity. Watch for squash-merge author rewrites, where Copilot can become the squash author and inflate itself as the primary contributor, corrupting blame.
Does a Co-authored-by trailer prove AI wrote the code?
No. The trailer proves tool-touch, not authorship. Any human can paste a Co-authored-by line by hand, and Cursor earned real GitHub Pair Extraordinaire badges purely from co-author credit. Treat the trailer as evidence an agent was present, then corroborate with session links and config files before you decide whose skill the commit proves.
How do agent signals break on GitLab and self-hosted repos?
GitHub detectors encode host assumptions. GitLab bot identities carry no [bot] suffix and do not use @users.noreply.github.com; they use project{id}_bot@noreply.gitlab.com, group_{id}_bot_{random}@noreply forms, and a service_account_ prefix. A detector keyed to GitHub patterns silently recounts these as people. Self-hosted instances can even use example.com-style domains, so re-run host-specific regex before counting.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.