Refolk
FrameworkInvesting and deal sourcing

The Early-Hire Team-Quality Score for Seed and Series A Deals

You can turn a company's early roster into a defensible team-quality score that flags both de-risking talent and warning signs before a partner meeting.

17 min readLast reviewed August 13, 2026Read as Markdown

Every published team guide is founder-facing and stops at the co-founders. This one scores the people the founders managed to attract: the first ten to twenty non-founder hires. That roster is separate evidence of talent magnetism, org coverage, and retention, and it deserves its own defensible score. This guide is for seed and Series A investors, platform and talent partners, and angels who need to walk into a partner meeting able to say whether a company's early roster makes the deal safer or riskier, using only public profiles.

The job is narrow and repeatable. You take a company's early roster from public evidence and produce a single team-quality score plus the two or three facts that decided it. The score flags de-risking talent and warning signs before you spend partner time. Below is the model: the four dimensions that matter, how to score each, what each score means, and the failure modes that quietly wreck the read.

Why the non-founder roster is its own diligence axis

The non-founder roster is direct evidence on the most heavily weighted diligence axis you have, not a tiebreaker. In a VC survey, 95% of investors called the management team important and 47% called it the single most important factor. That weight is the whole reason to score early hires separately: second-order team evidence compounds on the axis you already care about most.

Founders can rehearse a narrative. The roster cannot. The first five to ten hires are a documented signal in their own right - impressive people willing to take the bet de-risk the investment, because someone who could earn far more elsewhere chose the vision instead. That is why investors look for evidence a founder can attract talent above their weight. Landing one genuinely excellent early hire is a strong signal a founder can sell a vision, and a roster full of them is that signal repeated.

47%
of VCs who called the management team the single most important investment factor
95% called it important, which makes the roster direct evidence on your top-weighted axis, not a nice-to-have.

The founders are already graded elsewhere. Refolk's investing guides cover founder claim verification and reference calls. This score covers the people the founders hired, which is a different question with different evidence. A checkable track record on the founders tells you they can build. A high-tier, well-retained early roster tells you they can pull others into the build. Both matter, and this guide only addresses the second.

The four dimensions of the score

Score four dimensions and weight them into one number: talent tier, retention, coverage, and org proof. Each answers a distinct question, each has a public-evidence signal, and each has a way it lies. Grade all four before you compose a score, because any one of them can carry a false positive on its own.

DimensionWhat it provesPublic signalHow it lies
Talent tierThe founder can attract talent above their weightPrior employer tier and prior seniority per hireTitle inflation and pedigree pattern-matching
RetentionHires stay past the honeymoon phaseStart and end dates, GitHub commit recencyHoneymoon mirage at 12 to 18 months
CoverageThe right functions exist for the stage and modelFilled vs missing roles against the benchmarkFlagging gaps the benchmark says are normal
Org proofThe team is starting to layer, not just stackFirst manager and lead hires, reporting depthFlat org read as a fault at seed

The order matters when you compose. Talent tier and retention carry the most weight because they are the hardest to fake and the closest to the magnetism question. Coverage and org proof are calibration dimensions: they tell you whether the roster is the right shape for the stage, and they mostly protect you from false alarms rather than generating positive signal.

What the roster proves, outermost first

  1. Talent tier
    Can this founder pull people above their weight into the bet
  2. Retention
    Do those people stay past the honeymoon window
  3. Coverage
    Are the right functions filled for the stage and business model
  4. Org proof
    Is the team beginning to layer with managers and leads
Each layer answers a distinct diligence question, and you grade all four before scoring.

Set the stage benchmark before you grade anything

Fix the expected headcount and function order for the round before you look at a single hire, because every coverage judgement is relative to that benchmark. A missing role is only a red flag if the benchmark says it should be filled by now. Grade against a full org chart and you manufacture false negatives on every seed deal you touch.

Start with headcount. Seed startups averaged 5.3 employees in H1 2024, down from 6.9 in 2021, and Series A averaged 15.6, down from 17.6. The Angel Capital Association reports median headcount rising in orderly steps from 5 FTEs at pre-seed to 20 by Series C. Use the number for the round in front of you.

StageAvg or median headcountSource
Pre-seed5 FTE (median)Angel Capital Association
Seed5.3 (avg)Carta H1 2024
Series A15.6 (avg)Carta H1 2024
Series C20 FTE (median)Angel Capital Association

Then set function order, and here the sources disagree in a way you must resolve per deal. Carta and the ACA imply near-uniform headcount ramps, but function order varies sharply by business model. The first sales hire averages employee 9 overall, but 6 for SaaS and 15 for API companies. The first designer averages employee 19 and the first product hire averages 25, rising to 37 at Series A firms.

First functional hireAvg hire number (all)SaaSAPI
Sales9615
Designer1918 (Series A)later
Product25-37 (Series A)

Score talent tier and above-your-weight hires

Score talent tier by the delta between each hire's prior employer tier and prior seniority and the startup's current stage, not by their current title. There is no published numeric formula for "talent above your weight," so this delta is the defensible proxy. Flag anyone who plausibly left a materially bigger role to take the bet.

Read prior employer and prior seniority, never the current title. A "Founding Engineer" or "Head of X" at a four-person company signals nothing about tier on its own. The signal lives in what the person left behind: a senior role at a strong prior employer, walked away from for a seed company, is the magnetism evidence. Refolk's index shows how many people now carry that founding-engineer title - 4,434 in the US, 613 in the UK, 375 in Canada, and 235 in Germany - which is exactly why the title alone means little and the prior role means everything.

The repeat-teammate variant is the strongest single signal in this dimension. If former teammates line up to join a founder again, that is evidence you cannot buy: people who saw the founder at their worst chose to come back. Look for hires who followed the founder from a previous company, and weight them heavily.

18.9x
how much rarer the founding-engineer pool is in Germany than the US in Refolk's index
4,434 US to 235 Germany, so an above-weight technical hire in Berlin is a mechanically stronger magnetism signal.

Geography constrains the pool, so calibrate against it. With US founding-engineer supply roughly 7.2x the UK's and 18.9x Germany's in Refolk's index, an above-weight early technical hire in Berlin or Toronto is mechanically rarer than the same hire in San Francisco or New York, where the talent concentrates. A founder who landed one in a thin market did something harder. In Refolk's index this pool concentrates in San Francisco and New York in the US, almost entirely London in the UK, Berlin and Munich in Germany, and Toronto and British Columbia in Canada.

An above-weight hire in a thin talent market is harder to land and therefore a louder magnetism signal.

Finding the above-weight hires by hand is the slow part: manual profile research runs 20 to 30 minutes per person and the data goes stale fast. This is where a plain-English roster query removes the friction the paragraph just described.

Score retention against the year-two cliff

Score retention by flagging sub-12-month exits and year-two departure clusters, and weight departures only after the honeymoon window. Median startup job tenure was 2.2 years as of Q2 2024 against 4.1 across all industries, turnover is highest in the first two years, and there is roughly a 51% chance an employee leaves within three years. Short tenure is normal for startups. The question is where the departures fall.

Two windows carry the signal. A departure inside the first 12 months is the clearest warning, because first-year employees churn at 2 to 3x the rate of employees with five or more years. A cluster of exits in year two is the second, subtler flag.

The year-two window is the real retention test. With 95%+ retention normal in the first 12 to 18 months and a 51% three-year departure probability, the strongest single piece of evidence is a hire who has crossed roughly the 18-month mark and stayed. Weight those hires. Discount a perfect record on a young roster to roughly nothing.

Reading a departure by when it happened

Ordinary hireAbove-weight hire
Above-weight, early exit
Loudest warning; a strong hire left before the cliff, dig into why
Above-weight, late exit
Mild signal; a normal-tenure departure of a strong hire
Ordinary, early exit
Warning; a first-year churn flag against the 12-month cliff
Ordinary, late exit
Weakest signal; consistent with a 2.2-year median tenure
Early tenure at exitLate tenure at exit
The same exit means opposite things depending on tenure at departure, so plot every departure before you score.

Watch for the stale-roster trap here. A departed star who never updated their profile still lists the company, which makes the roster look stronger and the retention record look cleaner than it is. Cross-check start and end dates and GitHub commit recency against employer records before you credit anyone as a current hire.

Score coverage and org proof

Score coverage by mapping filled versus missing functions against the stage benchmark, then label each gap as expected or red flag for this business model. The framing investors reward is not a full org chart. It is a founder who names which functions are missing and which two or three hires or advisors this round adds. Every startup has gaps; the danger is ignoring them.

Separate honest gaps from red flags mechanically:

  • An absent designer before roughly hire 19 is an expected honest gap, not a red flag.
  • An absent dedicated product hire before roughly hire 25 is an expected honest gap.
  • An absent first sales hire well past hire 9 at Series A is a coverage red flag - past hire 6 for SaaS, but not until hire 15 for API firms.
  • A function the business model demands early but that is still empty is a red flag; a function the benchmark places later is not.

Org proof is the smaller, final read: note the first manager and lead hires and whether reporting layers exist yet. A team that has begun to layer - a first engineering lead, a first person who manages others - is different evidence from a flat roster of individual contributors. At seed, a flat org is expected and not a fault. By Series A, some layering is a sign the founder is building an organization rather than a group. Do not over-read this dimension; it is calibration, not a headline.

How this read goes wrong

The failure modes are the most valuable part of this standard, because a roster read that ignores them systematically over-scores teams. Every one below is a documented way the public evidence lies, and each has a specific check.

Stale-roster false positive. A departed star still lists the company because they never updated LinkedIn, so the roster looks stronger than it is. A lazy read over-scores every team this way. Check start and end dates and GitHub commit recency, and cross-check against employer records. This is where most of the defensible signal actually lives.

Title inflation. "Founding Engineer" or "Head of X" at a four-person company signals nothing about tier. Score prior employer and prior seniority, never the current title.

Honeymoon mirage. A 100% retention roster at 14 months is expected, not excellent, because the 95%+ first-year rate masks burnout and year-two exits. Weight departures only after the 12-month cliff and credit hires who crossed month 18.

Coverage false alarm. Flagging "no designer" or "no product hire" at seed when the benchmark places those around hire 19 and 25. Grade gaps against the stage benchmark, not against a full org chart.

Business-model mismatch. Applying SaaS function order, with sales at hire 6, to an API or infra company where sales arrives around hire 15 produces a false red flag. Set the benchmark per business type.

Pedigree over-weighting. A roster of ex-FAANG names is not proof of retention or coverage, and pattern-matching on pedigree and network access carries known bias. Require a departure read and a coverage read alongside tier before you let a shiny roster carry the score.

GitHub blind spot. Absence of public repos is not absence of talent, because only 18% of GitHub activity is public. Do not down-score infra or backend hires for a quiet GitHub, and remember Stack Overflow found 74% of developers are not actively job-hunting, so public profiles support pre-outreach evaluation rather than a live availability read.

The scoring procedure

Run these seven steps in order to move from a public roster to a one-page memo. The first six are analyst work; the last is partner-facing. Budget roughly two and a half hours for a clean deal, most of it in Step 1.

From public roster to defensible score

  1. Assemble the roster
    Pull all non-founder names from LinkedIn current-and-past-employees plus GitHub contributors and cross-match handles. Done means 10 to 20 named hires with title, prior employer, seniority, and start date. Manual checks run 20 to 30 minutes per person and go stale fast.
  2. Set the stage benchmark
    Fix expected headcount and function order for the round, roughly 5 at seed and 16 at Series A, adjusted for the business model. Done means a target roster shape to grade against.
  3. Score talent tier and above-weight hires
    Record prior employer tier and prior seniority for each hire and flag anyone who plausibly left a materially bigger role or followed the founder in. Done means a count of above-weight and repeat-teammate hires.
  4. Score retention
    Compute start dates and any departures, then flag sub-12-month exits and year-two clusters against the 2.2-year median. Done means a retention flag list that weights departures only after the 12-month cliff.
  5. Score coverage
    Map filled versus missing functions against the benchmark and separate honest gaps from red flags. Done means a gap list where each entry is labelled expected or red flag for this model.
  6. Score org proof
    Note the first manager and lead hires and whether reporting layers exist yet. Done means a seniority-mix read that shows whether the org is starting to layer.
  7. Compose the score and write-up
    Weight the four dimensions, produce a single defensible score, and name the two or three deciding facts. Done means a one-page memo ready for the partner meeting.

Use a fixed rubric so the score is comparable across deals and defensible in the room. Weight talent tier and retention highest, and use coverage and org proof to calibrate rather than drive.

Early-hire team-quality rubric
Company: __________  Stage: seed / A  Business model: SaaS / API / other
Roster size (non-founder, dated-current): ___ vs benchmark ___
Talent tier (0-40): above-weight hires ___  repeat-teammates ___  thin-market bonus? Y/N
Retention (0-30): sub-12mo exits ___  year-two cluster? Y/N  hires past 18mo ___
Coverage (0-20): red-flag gaps ___  honest gaps ___ (labelled per model)
Org proof (0-10): first manager/lead present? Y/N  layering started? Y/N
Total (0-100): ___
Deciding facts (2-3): __________
Verdict: de-risks / neutral / adds risk

Adjust the weights to your fund's priors, but keep talent tier and retention dominant and set the coverage benchmark per business model.

Before you call the score done

Verify the read against this checklist before it goes in front of a partner. A roster score is only defensible if every current hire is dated, every gap is graded against the right benchmark, and the two or three deciding facts are named rather than implied.

Ready-for-partner checklist

  • Every current hire is dated against start-date, end-date, or commit recency, not taken at face value from an un-updated profile.
  • Talent tier is scored on prior employer and prior seniority, not on current titles like Founding Engineer or Head of X.
  • Retention flags weight departures only after the 12-month cliff and credit hires who crossed roughly month 18.
  • Every coverage gap is labelled expected or red flag against the stage benchmark for this specific business model.
  • The roster is not carried by pedigree alone; a departure read and a coverage read stand alongside the tier read.
  • Infra and backend hires are not down-scored for a quiet GitHub, since only 18% of activity is public.
  • The memo names the two or three facts that decided the score, so a partner can challenge them directly.

To keep the read current, re-date the roster before any decision that depends on it, because manual profile checks go stale fast and the honeymoon window moves under you as the company ages. A roster you scored six months ago has crossed part of the year-two cliff since, which can turn a neutral retention read into a strong one or a warning. The mechanism to re-check is the same every time: pull the current-and-past roster again, re-date each hire, and re-grade the gaps against the stage the company has now reached. A tool like Refolk that returns the roster with prior employers and start dates in one pass removes the per-person research cost that otherwise makes re-checking too expensive to do often.

Questions practitioners ask

How many early hires should a seed or Series A company have?

Use the stage benchmark as your baseline. Seed startups averaged 5.3 employees in H1 2024 and Series A averaged 15.6, down from 6.9 and 17.6 in 2021. The Angel Capital Association reports median headcount rising in orderly steps from 5 FTEs at pre-seed to 20 by Series C. Judge a roster against the number for its round, not against a full org chart, and treat significant overstaffing for stage as its own question.

What tenure counts as a retention red flag for an early hire?

A departure inside the first 12 months is the clearest warning, since first-year employees churn at 2 to 3x the rate of five-year employees. A cluster of exits in year two is the second flag, because early-stage retention is often artificially 95%+ for the first 12 to 18 months before a year-two cliff. Median startup tenure is 2.2 years, so a hire past roughly month 18 who stayed is strong evidence.

How do I judge whether an early hire is talent above the company's weight?

No published numeric formula exists, so use a documented proxy: compute the delta between the hire's prior employer tier and prior seniority and the startup's current stage. Flag anyone who plausibly left a materially bigger role for the bet. The repeat-teammate variant is also strong evidence: former colleagues who chose to join the founder again saw them at their worst and came back anyway.

Is a missing sales, design, or product hire a red flag at seed?

Usually not, and grading it as one is a common false alarm. The first designer averages employee 19 and the first product hire averages 25, so their absence at seed is an expected honest gap. Sales is model-dependent: the first sales hire averages employee 9 overall but 6 for SaaS and 15 for API firms, so an absent sales hire past headcount 10 at a Series A SaaS company is a real coverage flag while the same gap at an API company is normal.

Can I trust a roster built only from public profiles?

With care. LinkedIn carries self-reported titles, prior employers, seniority, and start dates but misses people who never updated a departure, hides private profiles, and inflates titles. GitHub covers engineers but only 18% of activity is public and it has no start dates or open-to-work signal. Cross-check start and end dates and commit recency, and treat an un-updated star listing as a stale-roster false positive rather than a live hire.

Try it on your own search

Stop building boolean strings. Just describe the person.

Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.

  • One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
  • Read live at search time, not from a database that went stale last quarter.
  • Watch every step as it runs, and see why each name made the list.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next