# The Public-Evidence Leveling Standard: Senior, Staff, or Insufficient

*You can assign a candidate senior, staff, or insufficient public evidence from the same artifacts an independent reviewer would, and defend the call.*

- Canonical URL: https://www.refolk.ai/guides/public-evidence-leveling-standard
- Pillar: Engineering and open source
- Format: Standard
- Published: 2026-10-10
- Last reviewed: 2026-10-10
- Reading time: 7 min

This standard fixes the evidence required to call a candidate senior or staff from their public work, and forces an explicit "insufficient public evidence" verdict when the public record is too thin to support any band. It is for sourcers, engineering managers, developer-relations leads, and technical founders who attach a comp band to a level before the hiring-manager screen. Read it and you can grade a profile the way an independent reviewer would from the same artifacts, and defend the call.

Most GitHub-seniority write-ups hand you a loose bag of signals - stars, green squares, deletion commits - with no bar for when they prove a level. The result is that two people grade the same profile differently and leveling collapses to taste. This document does the opposite: it names the evidence each band requires, names the signals that are inadmissible, and makes "do not level" a legitimate, common outcome.

## What Senior and Staff actually require as public evidence

The Senior-to-Staff jump is a change in scope, not a skill increment, and only some of it is publicly observable. Published ladders agree on this. A Staff engineer operates across teams; a Senior operates within one. So leveling from public artifacts is really an exercise in reading scope off verifiable work.

The named Staff criteria that leave a public trace are concrete. One ladder lists having led a cross-team initiative that required alignment from three or more teams, having influenced technical direction beyond one's own team through adopted RFCs or standards set, and being sought out by other teams for technical input. Of these, cross-team RFCs, adopted standards, and design docs are publicly observable. "Sought out for input" and org trust are not, so they cannot be part of a public-evidence standard.

Senior is lower, but not low. Glitch's public ladder defines Senior engineers as the first level of technical leadership, with expertise in at least one technical domain and depth of knowledge that lets them debug systems effectively. In public terms, that looks like sustained, reviewed, within-team contribution of real complexity - not necessarily anything cross-team.

#### What is publicly observable at each layer

1. **Cross-team and org scope** - Adopted RFCs, set standards, published strategy - observable Staff evidence
2. **Within-team technical depth** - Reviewed PRs, maintained modules, debugging trail - observable Senior evidence
3. **Organizational trust** - Sought out for input, influence in meetings - real but not public
4. **Private delivery** - ~80% of all activity, invisible by default - route to insufficient, not low

*Only the outer layers leave a public trail you can verify; the inner ones are interview territory.*

The population itself backs the scope argument. In Refolk's index of professional profiles, there are 4.27 US Senior engineers for every US Staff engineer. Staff is rare because it is a step-change, not the top of an output curve. Counting more commits never crosses that gap.

**4.27:1 - US Senior to Staff engineers in Refolk's index**

214,983 Senior against 50,385 Staff; the ratio tells you Staff is a scope jump, not an output tier.

### Table A - Staff vs Senior supply in one market

| Band | Population | Ratio vs Staff |
| --- | --- | --- |
| Senior / Senior Software Engineer | 214,983 | 4.27x (derived) |
| Staff / Staff Software Engineer | 50,385 | 1.00x baseline |

Populations are from Refolk's index; the ratio is derived by dividing Senior by Staff. The reading is simple: if your rubric would promote one in four Seniors to Staff on output alone, it is not measuring the thing that makes Staff rare.

## The private-work cap that bounds every public verdict

A public profile exposes roughly one-fifth of a typical engineer's recorded activity, so absence of public work is never evidence of low level. This is the most load-bearing fact in the standard, and it is the reason the insufficiency verdict has to exist.

In 2025, 81.5 percent of contributions happened in private repositories, while 63 percent of all repositories were public. The pattern holds across years: more than 82 percent of contributions were private in 2024, across 4.3 billion contributions in more than 181 million private repositories, and private projects were more than 80 percent of all activity in 2023.

### Table C - Private share of activity by year

| Year | Private share of contributions | Source |
| --- | --- | --- |
| 2023 | 80%+ | github.blog state-of-open-source |
| 2024 | 82%+ | github.blog octoverse-2024 |
| 2025 | 81.5% | github.blog octoverse 2025 |

Figures are as reported in each year's Octoverse post. The implication for leveling is direct: a thin public profile is statistically normal, not a weakness. A standard that treats silence as low-level will mislevel most strong engineers, especially those who have spent their careers inside private codebases.

> **Rule:** Silence routes to insufficient, never to a low band
>
> Because roughly 80% of all activity is private every year, a thin public profile means you lack evidence, not that the candidate lacks level. Assign insufficient public evidence and let the screen gather the private signal.

Market size also changes what evidence is even available. In Refolk's index, the US Staff pool is 15.2 times the UK's. A reviewer leveling UK candidates works from a far smaller comparison set, so relative benchmarking is weaker there and verified artifacts should carry more of the load.

### Table B - Staff Engineer across two markets

| Country | Staff population | Share of US (derived) |
| --- | --- | --- |
| United States | 50,385 | 100% |
| United Kingdom | 3,319 | 6.6% |

Populations are from Refolk's index; the share is derived by dividing UK by US. The point is not that UK Staff are scarcer in absolute terms. It is that you cannot lean on "how this profile compares to the field" when the field is small.

## Signals that are inadmissible as leveling evidence

Stars, followers, raw commit volume, and contribution-graph density are context, not proof, and must be stripped from the evidence pile before you grade. Each one is either adversarially gamed or measures activity rather than value, so a band resting on them is indefensible.

Stars are a reputation market. A study using the StarScout tool analyzed GitHub event data from July 2019 to December 2024 and identified six million suspected fake stars across 15,835 repositories. One venture firm calls stars "vanity metrics" outright and tracks unique monthly contributor activity instead. A flashy repo with few real contributors is the false positive to watch for.

The contribution graph is just as soft. Bots push tiny meaningless changes to private repos daily - an empty commit, a whitespace edit, a single-character README change - and tools automate this, backdating commits to fill gaps. A dense green graph can be entirely scripted.

Commit and line-of-code volume measure the wrong thing. A developer who refactors a 400-line module down to 120 cleaner lines registers as negative productivity under a lines-of-code metric. Two engineers can ship the same feature with wildly different code volumes, and the one with less code often did the better job.

> **Watch out:** The green-graph halo and the star-count authority
>
> A dense contribution graph reads as "works daily" but can be scripted empty commits - open the actual diffs. A high star count reads as "important project" but can be purchased in bulk - cross-check unique monthly contributors.

> Stars and graphs carry a seller, not a skill, so no defensible band can rest on them.
> </pull>
>
> What survives the strip is the admissible pile: specific technical contributions you can point a reviewer at. The job from here is to verify who actually made them.
>
> ## Verifying authorship: the real bottleneck
>
> Confirming the candidate actually did the work is where leveling succeeds or fails, because claimed authorship is unverifiable by default. Collection is easy; verification is the hard part, and most rubrics skip it.
>
> Git's author fields are free-form text any client can set to anything. Cryptographic signing binds a commit to a key pair only the owner controls, but a measurement study found contribution workflows can be abused in 85.9 percent of 50,328 critical projects, with 573,043 email addresses a malicious actor could claim to hijack historic contributions. The same study found 95.4 percent of users never signed a commit and 72.1 percent of projects had no signed commit at all. So "contributor to X" is a claim, not a fact, until you verify it.

stat
number: 85.9%
label: Critical projects with abusable contribution workflows
note: Across 50,328 critical projects, with 573,043 email addresses a malicious actor could claim.

## Frequently asked questions

### Can I level someone as senior or staff from their GitHub profile alone?

Rarely, and never from profile metrics. Roughly 80 percent of GitHub activity happens in private repositories every year since 2023, so a public profile exposes about one-fifth of a typical engineer's recorded work. You can level from public artifacts only when you have verified-authorship items whose scope maps to a band. Otherwise the correct call is insufficient public evidence, not a low band.

### What is the single hardest part of leveling from public work?

Verifying authorship, not finding artifacts. Git author fields are free-form text any client can set, contribution workflows were abusable in 85.9 percent of 50,328 critical projects, and 95.4 percent of users never signed a commit. So a claim of contributing to a famous repo is unverifiable by default. Require signed commits, named doc authorship, a CFP listing, or a merge-review trail before an artifact counts.

### Why not just count stars, commits, and green squares?

Because all three are adversarially gamed or measure activity rather than value. One study found six million suspected fake stars across 15,835 repositories, tools backdate empty commits to fill the graph, and a 400-line to 120-line refactor reads as negative productivity under a lines-of-code metric. These are inadmissible as leveling evidence. Treat them as context only.

### How do I keep two reviewers from grading the same profile differently?

Use a shared rubric with behavioral anchors and have each reviewer score independently before discussing. Structured rubrics raise inter-rater reliability from about 0.37 to 0.67, and panels reach about 0.74 versus 0.44 for separate reviews. The key mechanism is independent evidence-cited scoring first, because the person who speaks first otherwise pulls the others toward their call.

### Is a Staff title at a small company the same as Staff elsewhere?

No. Ladders are not a universal ranking, and a Staff title at a small shop may equal Senior at a larger one. Grade the scope of the verified artifacts, not the word on the resume. The Senior-to-Staff jump is a qualitative shift to cross-team scope, so look for a cross-team initiative, adopted RFCs, or standards set beyond one team.

### Should I penalize candidates with thin public profiles?

No. Absence of public work is the default, not a red flag, because the large majority of all GitHub activity is private. Penalizing a thin profile mislevels strong engineers from private-shop backgrounds. Route them to insufficient public evidence and let the hiring manager gather the private signal through the interview, rather than assigning a low band you cannot defend.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/public-evidence-leveling-standard*
