The Capability Gap Map: Benchmarking a Competitor's Team Skill Depth
You will produce a capability-by-capability gap map scoring a named competitor deeper, at parity, or weaker than your team, each score backed by public evidence and a confidence tag.
Key takeaways
- Score capabilities as rows, not job titles, because titles are a weak proxy for real skill; in an audit of 60 LinkedIn profiles, only 19 had Skills sections matching their target role.
- Read depth against supply, not in a vacuum: in Refolk's index, 965 US Rust engineers versus 182 in Germany is a 5.3x supply gap that changes what a competitor's team size means.
- Weight behavioural evidence over credentials, because a merged PR is a peer-review pass from an experienced engineer, while a self-reported skill list is often a wish list.
- Senior depth in a niche stack concentrates in a few firms: only 13% of US Rust engineers sit at Staff or Principal, and Google plus Meta hold 14 of those 146.
- Coverage confidence is the real deliverable; label each row high or low confidence and never score a competitor weaker on low public activity alone.
- Unauthenticated GitHub crawling caps at 60 requests per hour, so large rosters silently truncate; verify every contributor was fetched before you trust a coverage number.
This is the working method for answering one question about a named competitor: skill by skill, is their team deeper, at parity, or weaker than mine, and how sure am I? It is written for strategy and research teams, talent-intelligence analysts, and operators sizing a competitive landscape. You will finish with a capability-by-capability gap map where every score is backed by a public profile or repository and carries a coverage-confidence tag, so a reader can trust the confident rows and discount the thin ones.
The job matters because skill gaps are now the constraint executives cite most. In the WEF Future of Jobs Report 2025, which draws on data from over 1,000 companies, 63% of employers identify skill gaps as the biggest barrier to business transformation over the 2025 to 2030 period. If your own gaps are that decisive, so are your competitor's, and knowing where they concentrate their scarce capability is a genuine planning input.
Why a gap map instead of a headcount comparison
A gap map scores capabilities against evidence and marks how sure you are; a headcount comparison counts people and stops. The difference is the whole point, because headcount lies in predictable ways and capability read against evidence does not.
Titles are the first thing that lie. A competitor with many "Senior" engineers can look deep while capability is thin, because titles are a convenient label for org charts and pay bands and a weak proxy for what people can actually do. Self-reported skills lie in the other direction: in an audit of 60 public profiles, only 19 had Skills sections matching their target role, and 41 listed what they had done rather than what they wanted. So a naive count of "engineers who list Kubernetes" mixes real operators with job seekers padding a profile.
The gap map fixes this by doing three things a headcount cannot. It maps capabilities as rows in a skill graph rather than boxes in an org chart. It corroborates each self-reported claim against behavioural evidence from public code. And it attaches a confidence band, so a row scored on merged pull requests to an established project reads differently from a row scored on a skill tag alone.
A headcount tells you how many people. A gap map tells you what they can do and how sure you are.
The two public strata you are reading
Every capability score is built from two kinds of public data: self-reported profile networks and behavioural code repositories, with job postings as a forward signal. Each proves something different and lies in a different way, so you use all three and never trust one alone.
Self-reported profiles are what a person chose to write about themselves. That makes them broad but soft. A profile can list up to 50 skills while only the top three show by default, and across the job market people list an average of 11 skills when asked. Presence of a skill proves intent to be found for it, not competence. Code repositories are the opposite: the public GitHub graph is a portfolio of verified work where every merged PR and code review is a real demonstration. A developer with an 80% merge rate across multiple projects has consistently written code that senior engineers approved, which is stronger than any take-home. Job postings point forward, because who a company is hiring is the first outward sign of an internal shift, while a profile reflects where a person has been.
| Source stratum | What it proves | How it lies |
|---|---|---|
| Self-reported profiles | Intent to be found for a skill; tenure and title | Wish lists; 41 of 60 audited profiles listed past work, not target role |
| Public code repositories | Verified, peer-reviewed capability via merged PRs | Green-square theatre; fork farms with no original work |
| Job postings | Where the competitor is investing next | Aspirational reqs; roles posted but never filled |
The practical rule follows directly: use profiles to build breadth, use repositories to confirm depth, and read the two together. Behavioural evidence beats credentials because when a maintainer merges a PR from an outside contributor, that is a peer-review pass from an experienced engineer. That is why merged PRs, not self-reported skills, carry the highest confidence weight in the scoring step.
Read depth against supply, not in a vacuum
A competitor's team size only means something once you compare it to the available talent pool. Twenty engineers in a scarce skill is a moat; the same twenty in an abundant skill is replaceable. Supply is the denominator that turns a raw count into a judgement.
The numbers make this concrete. In Refolk's index of professional profiles, the US Rust pool and the German Rust pool differ by more than five times, so the same competitor team reads as a defensible moat in one geography and an ordinary staffing choice in the other.
| Market | Rust SWE/Sr SWE count | Top employer signals |
|---|---|---|
| United States | 965 | Meta, Oxide Computer, Google |
| Germany | 182 | Google, Bosch |
| US:Germany ratio (derived) | 5.3x | - |
Supply also tells you how much corroboration each row needs. A Go-heavy competitor produces far more public profiles per capability than a Rust-heavy one, so profile-only scoring is riskier for the abundant skill and repository corroboration matters more. Niche-stack rows are small enough to verify almost exhaustively.
| Skill | US SWE/Sr SWE count | Go:Rust multiple (derived) |
|---|---|---|
| Go | 4,940 | 5.1x |
| Rust | 965 | 1.0x (baseline) |
Seniority thins the pool faster than the raw counts suggest, and that is where parity findings become meaningful. Only 13% of US Rust engineers sit at Staff or Principal level, and Google plus Meta hold 14 of those 146. Parity at senior level in a niche stack is therefore a much rarer and more consequential finding than parity overall, because senior depth concentrates in a handful of firms.
| Seniority band | US Rust count | Share of IC+Staff pool (derived) |
|---|---|---|
| Software/Senior Software Engineer | 965 | 87% |
| Staff/Principal Engineer | 146 | 13% |
Pulling these supply denominators by hand is the slow part; the counts above come from asking Refolk's index for total supply per skill, geography, and seniority band. That is the input that lets you read a competitor's depth against the pool rather than against nothing.
The procedure, start to finish
The method runs seven stages from a blank page to a traceable gap map, and takes roughly six to nine working days for a single competitor across 8 to 15 capabilities. The stages below match the steps block exactly; do them in order, because the corroboration and supply steps depend on the roster and taxonomy being fixed first.
Benchmark a competitor's team skill depth
- Scope the mapPick the named competitor and define 8 to 15 capabilities as rows, not job titles, plus your own team's baseline score for each. Map by skill graph, not org chart.
- Anchor to a taxonomyMap every capability to a recognised skills framework such as O*NET so both teams are compared on the same terms. Every row gets a taxonomy code.
- Pull the roster from profile dataEnumerate the competitor's engineers with stated skills, titles, and tenure, treating all of it as self-reported. Use full-text passes, not just tag filters, so skill-tag omissions do not undercount.
- Corroborate with repository evidenceFor each capability, find contributors to the competitor's public repos and the top 3 to 5 ecosystem repos in that stack. Weight merged PRs to established projects highest.
- Size the market with aggregate countsPull total supply per skill, geography, and seniority so competitor depth is read against the available pool. A supply denominator sits next to every row.
- Score deeper, parity, or weaker with confidenceCompare competitor evidence to yours, assign a three-band score, and attach a coverage-confidence tag. High where repositories corroborate profiles, low where only self-reported.
- Package the gap mapDeliver a capability-by-capability table with score, evidence links, and confidence. Every score is traceable to a specific public profile or repository.
A few notes on the harder stages.
Scope (0.5 to 1 day). The output is a living document typically covering 10 to 15 critical and scarce capabilities. Lock your own team's baseline first, because you cannot score "deeper or weaker" without a "than what". Map by skill graph, not org chart, so that a capability like "distributed systems" is one row even if it lives across three teams.
Anchor to a taxonomy (0.5 day). A structured skills inventory mapped against a recognised taxonomy such as O*NET gives you a consistent way to compare roles and functions. Skip this and "backend" on your side may quietly mean something different from "backend" on theirs.
Pull the roster (1 to 2 days). Enumerate the competitor's engineers and their stated skills, treating everything as self-reported. Watch the Boolean exclusion trap: the Skills section is often the only field a recruiter's Boolean search filters as a discrete checkbox, so an engineer who omits a skill is excluded from a filtered search rather than ranked lower. Run keyword and full-text passes as well, or you will deflate the competitor's true depth.
Corroborate (2 to 3 days). To find engineers in a specific stack, monitor the top 3 to 5 repositories in that technology ecosystem, then map their contributors back to the competitor. Weight merged PRs to established projects highest, because a merge is a peer-review pass.
Scoring deeper, parity, or weaker with confidence
Each row gets two marks: a three-band score comparing the competitor to you, and a coverage-confidence tag saying how much of that score rests on behavioural evidence. No published standard exists for either, so state your rule plainly in the deliverable and make each row auditable.
The score is a comparison, not an absolute. Deeper means the competitor's corroborated evidence clearly exceeds yours for that capability; parity means they are within noise of each other; weaker means your corroborated evidence clearly exceeds theirs. Read against supply, so a small absolute team in a scarce skill can still score deeper.
Confidence is the map's real product. Set it high when repository evidence corroborates the profile data for a row, and low when a row rests only on self-reported skills or when public activity was too thin to read. The two dimensions together produce the judgement calls below.
Score against confidence
There is a real disagreement in the frameworks about ordering. Some run plan, identify, measure, then act, measuring before prioritising. Decision-first practitioners insist gaps be connected to priorities first, because the analysis phase is where most efforts stall: teams produce detailed gap reports then struggle to translate them into a plan, usually because the gap data was never connected to business priorities. My position for a competitor benchmark is to fix the priorities during scope, so every row already points at a decision by the time you score it.
Here is a copy-pasteable scoring rubric to drop into the deliverable.
Capability: <capability name> (taxonomy code: <O*NET or framework code>) Our baseline: <count of corroborated practitioners on our side> Competitor stated-skill count: <from profile data> Competitor behavioural-evidence count: <merged PRs to established repos> Supply denominator: <total pool for this skill/geo/seniority> Score: Deeper | At parity | Weaker Coverage confidence: High (repos corroborate) | Low (self-reported only) Evidence links: <profile URL(s)>, <repo/PR URL(s)> Strategic priority this row informs: <the decision it feeds>
Fill one block per capability. Keep the evidence links; they are what make the score defensible.
How this goes wrong
Most bad gap maps fail in one of eight predictable ways, and every one of them is a false positive or a false negative you can check for before publishing. Treat this section as the QA pass.
- Title inflation masks depth. A competitor stacked with "Senior" titles looks deep when capability is thin. Titles are a weak proxy. Check against repository evidence, not headcount.
- Self-reported skill lists are wish lists. A skill is often present because it aids a job search; 41 of 60 audited profiles listed what people had done, not what they wanted. Check tenure and corroborating work before you count it.
- Green-square theatre. A profile full of activity squares can be inflated by automated scripts or minor README edits. Check merged PRs to established repos, not raw commit counts.
- Fork farms. A profile that is entirely forked repos with no original work suggests someone who copies rather than builds. Check for original repos and accepted external contributions.
- Private-work blind spot. Six months of silence can mean a shift to private work, not lost skill. Mark confidence low; do not score weaker.
- Boolean exclusion undercounts. Because skills match by tag, an engineer who omits a skill is invisible to a filtered search, deflating the competitor's true depth. Check with full-text passes, not just tag filters.
- Rate-limit truncation. Unauthenticated repo crawling caps at 60 requests per hour, so a large roster silently truncates and your coverage looks worse than it is. Confirm every contributor was actually fetched before scoring coverage.
- Gap map with no decision. A beautiful report that never becomes a plan is the most common failure. Check that each row maps to a strategic priority before you publish.
The rate limit deserves a number, because it is the failure that corrupts coverage confidence without any visible error. Plan your crawl against the tier you actually have.
| Access tier | Requests per hour | Use when |
|---|---|---|
| Unauthenticated | 60 | Never for a full roster; spot checks only |
| Authenticated (personal) | 5,000 | Standard for a single-competitor benchmark |
| GitHub App (Enterprise Cloud org) | 15,000 | Large rosters or several competitors at once |
What "done" looks like before you publish
The gap map is finished when every score is traceable, every confidence tag is justified, and every row points at a decision. Run this checklist before you circulate it, because a map that overclaims is worse than no map.
Pre-publish gap map checklist
- Every capability is a row in a skill graph, not a job title, and 8 to 15 rows total
- Each row is anchored to a recognised taxonomy code so both teams compare consistently
- The roster was gathered with full-text passes, not tag filters alone, to avoid boolean undercount
- Each capability carries both a stated-skill count and a merged-PR behavioural count
- A supply denominator sits next to every row so depth is read against the pool
- Every row has a three-band score and a high/low coverage-confidence tag
- No row is scored weaker on low public activity alone
- Every contributor was fetched under an authenticated rate limit, not truncated
- Each score links to a specific public profile or repository
- Every row maps to a strategic priority or decision
Keeping the map current
A gap map is a living document, not a one-time audit, because the underlying skills move. Roughly 39% of core workforce skills are expected to be transformed or obsolete by 2030, so a benchmark frozen today decays. Re-run the corroboration and supply steps on a fixed cadence rather than rebuilding from scratch.
Two mechanisms drive most of the change, and both are re-checkable. First, hiring: job postings are the first outward sign of an internal shift, so a new wave of reqs in a capability is your early warning that a competitor is moving to deepen it, often before any new profiles appear. Track their open roles per capability alongside the score. Second, talent flow: watching where engineers move to and from tells you which capabilities a competitor is losing or gaining. A cluster of platform-team departures is a leading indicator that a "deeper" row is about to become "parity".
When you re-run, the supply denominators are what change fastest and matter most, because a skill's scarcity is what turns a small team into a moat. Pulling fresh aggregate counts per skill, geography, and seniority is exactly the kind of ask Refolk answers in plain English, so you can refresh the denominators without rebuilding a crawl each quarter. Keep the evidence links from the last run, diff the counts, and only re-corroborate the rows where the score or the supply moved. That keeps a defensible benchmark alive for the cost of maintenance rather than a rebuild.
Questions practitioners ask
Can I benchmark a competitor's team skills legally using only public data?
Yes. This method uses only self-reported professional profiles, public code repositories, and job postings, all of which are published by the people and companies themselves. You are reading disclosures, not extracting anything private. The output is explicitly an estimate, not a claim of internal knowledge, which is why coverage confidence is attached to every score rather than presented as fact.
How many capabilities should a competitor gap map cover?
Between 8 and 15 rows works best, and practitioner guidance points to living documents covering 10 to 15 critical and scarce capabilities. Fewer than eight and the map is too coarse to guide a decision; more than fifteen and corroboration cost explodes, because each row needs both profile counts and repository evidence. Define them as capabilities in a skill graph, never as job titles.
Why is a merged pull request stronger evidence than a listed skill?
Because a merged PR to an established project is a peer-review pass from an experienced maintainer, while a listed skill is something a person chose to write about themselves. Self-reported skill sections appeared on only 19 of 60 audited profiles matching their target role, and 41 listed what people had done, not what they wanted. Behavioural evidence carries the highest confidence weight for this reason.
How do I avoid GitHub rate limits truncating my roster?
Unauthenticated REST requests cap at 60 per hour, which silently truncates any large roster. Use authenticated requests, which raise the limit to 5,000 per hour, or GitHub App requests in a GitHub Enterprise Cloud org, which reach 15,000 per hour. Before you trust a coverage figure, confirm every contributor on your list was actually fetched rather than dropped by a hit limit.
What does coverage confidence mean on a gap map?
Coverage confidence marks how much of a capability's score rests on behavioural evidence versus self-reported claims. High confidence means repository evidence corroborates the profile data. Low confidence means the score rests only on stated skills or that public activity was too thin to read. It is not a published standard, so state your banding rule plainly in the deliverable so a reader can audit it.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.