Refolk
FrameworkInvesting and deal sourcing

Grading a Founding Team on Public Evidence

You will score any founding team across six weighted dimensions from public sources, reach a defensible go/hold/pass call, and know exactly what to verify later.

14 min readLast reviewed July 31, 2026Read as Markdown

Key takeaways

  • Venture investors cite the management team as important 95% of the time and most important 47% of the time, yet team predicts initial funding more than it predicts long-term outcome, so cap its weight in your rubric.
  • Prior specific-industry experience is the strongest documented predictor of high-growth founding, out-predicting age and pedigree, but most rubrics score an impressive resume instead of a domain match.
  • Cofounder conflict drives 65% of high-potential startup failures but leaves almost no public trace, so alignment must be scored as low-confidence and pushed to references rather than scored high because nothing bad appeared.
  • Serial-founder credit is non-linear: second-time founders succeed at 34%, third-time at 28%, and fourth-time-plus at 22%, so weight domain fit over raw repetition.
  • Only 2.15% of US founders in Refolk's index list machine learning as a skill, which is why technical-founder signals inflate and why authorship verification beats a raw count of commits.
  • With 70% of US VCs holding elite-institution degrees and favoring similar founders, stripping pedigree credit and re-scoring on verified output surfaces underpriced non-pedigree teams.

This guide is for early-stage investors, platform and talent partners, and angels who need to form a defensible view of a founding team before they spend partner time or ask for references. It isolates one repeated judgement call: scoring a founding team from public evidence alone, arriving at a go/hold/pass decision, and knowing exactly which claims to verify later. It is not a full diligence checklist and not a reference-check how-to. It is the scoring model for the moment before those.

The existing public material splits into two piles. One is process diligence that assumes you are already deep in a deal. The other is reference-check advice that assumes you are at term sheet. Neither gives you a repeatable way to grade two teams the same way at the earliest gate, using only what the open web will tell you. That gap is where conference-room gut feel takes over, and gut feel is exactly where pedigree bias lives. This framework replaces it with named, weighted dimensions.

Why score the team at all, and how much weight it deserves

The team is the single most cited factor in venture decisions, but it is a predictor of funding more than of outcome, so you score it for the meeting and verify it for the check. In the most-cited survey of venture investors, the management team was named an important factor by 95% of firms and the most important factor by 47%, ahead of business model at 83%, product at 74%, market at 68%, and industry at 31%.

That would argue for weighting team above everything else. A newer study of over 8,000 sourced deals complicates it: team predicts which companies raise initial funding, while market and product better explain larger financings and long-term success. In that dataset roughly 30% of sourced deals raise at least $1M. The lesson is not to demote the team. It is to cap its weight so that a strong founding narrative cannot drag a weak market through your gate. Practitioner framing rarely lets team fall below 30%, on the logic that strong founders course-correct when data changes. Keep it heavy, keep it bounded.

47%
Venture firms that name the management team their single most important factor
The team is also cited as an important factor by 95% of firms, ahead of business model, product, and market.

There is a cheaper reason to score the team early. Cofounder conflict is the load-bearing human failure: in Wasserman's work, 65% of high-potential startups fail from conflict among cofounders. That is a team-scoped statistic, not a general failure ranking. The general ranking looks different, and you should hold both in view.

The six dimensions and what weight each carries

Score every team on the same six dimensions, fixed in writing before you search any names. Freezing the rubric first is the control that stops a Stanford line or a famous logo from quietly reweighting your judgement mid-read.

DimensionWhat a strong score provesHow it lies
Domain / industry fitFounders have prior specific-industry experience, the strongest documented predictorAn impressive but adjacent resume reads as fit when it is not
Founding-team track recordPrior building experience, weighted for outcome not count"Third-time founder" scored top when outcomes decline past venture two
Technical build evidenceReal shipped work and verifiable authorshipHigh commit volume mistaken for capability
Cofounder alignmentTenure overlap, prior shared work, role clarityAbsence of public friction read as harmony
Integrity / verificationClaims cross-checked against public registriesRegistry silence read as a clean record
Market pullExternal signal of demand: customers, patents, usageFounder assertion accepted as validation

Two dimensions deserve a specific caution before you weight them. Industry fit is the strongest predictor in the census-scale data yet is the one most rubrics never score explicitly, defaulting to "impressive resume" instead of "domain match." And cofounder alignment is the most predictive dimension that is the least publicly visible, which is why it must carry a low confidence flag rather than a high score.

What each public signal proves, and where it misleads

Every public signal is a proxy, and every proxy has a failure mode. Score the artifact, name it, and record what it would take to make it lie.

Domain and track record

Prior experience in the specific industry predicts much greater rates of entrepreneurial success, and the mean age at founding for the top 1-in-1,000 fastest-growing ventures is 45.0, not the mythical dropout. Serial-founder credit is real but non-linear. A Danish panel found serial entrepreneurs post 67% higher sales than novices, yet outcome-scoped figures show success peaks at the second venture and declines after: second-time founders at 34%, third-time at 28%, fourth-time and beyond at 22%. One US study even found that owners with a prior business had a 7% lower probability of exit. Read this as: weight domain fit over raw repetition, and stop giving linear credit for the count of past startups.

Domain fit against founding repetition

Serial founderFirst-time founder
Serial, wrong domain
Discount the pattern-match; past wins may not transfer
Serial, right domain
Highest score, but confirm the prior outcome type
First-timer, wrong domain
Lowest confidence; lean on market pull and references
First-timer, right domain
Underrated; domain fit is the strongest single predictor
Low domain fitHigh domain fit
Domain fit is the stronger axis; repetition helps most at venture two and fades after.

Technical build evidence

Build evidence is the one dimension the open web renders honestly, if you inspect substance rather than volume. High commit counts are theater when they are forks and trivial commits, so read authorship and the substance of what shipped. Patents are a genuine external signal: patent-owning startups are 6.4 times more likely to attract investment. Scarcity explains why this signal inflates. In Refolk's index, only 2.15% of US founders list machine learning as a skill, so credential and commit theater pays off precisely because real technical founders are rare.

SegmentCountShare of US founders
All US founders656,139100%
US founders listing Machine Learning14,1172.15%
2.15%
Share of US founders in Refolk's index who list machine learning as a skill
Technical-founder scarcity is why authorship verification beats a raw count of commits.

Finding the small population that actually has the build history, rather than the population that claims it, is the friction. Describing the profile you want in plain English and pulling it straight from the open graph removes a day of manual disambiguation.

The verification pass: registries that prove existence, and only that

Verification means cross-checking every claim against a public record and marking it verified, unverifiable, or contradicted. These sources prove the existence or absence of a filed record. None of them capture private disputes, unfiled conflicts, foreign jurisdictions, or how someone actually performed inside a prior role.

SourceWhat it verifiesWhat it cannot see
PACER / CourtListenerFederal bankruptcies, fraud, civil docketsState courts, settled or unfiled disputes
SEC EDGARCompany filings and enforcement actionsPrivate-company conduct
FINRA BrokerCheckSecurities licensing and disciplineNon-securities history
USPTO patent searchGranted patent ownershipWhether the patent has commercial value
Secretary of State registriesPrior-company existenceOperating history or performance

Resume inflation is why this pass is not optional: roughly 33% of applicants admit lying about their education, and only 34% of employers verify the credentials on a resume. A LinkedIn title is a claim, not a fact. Route every unverifiable or contradicted claim to the verify-later list rather than scoring it as if it were confirmed.

Scoring a team from public evidence, step by step

Run the same procedure on every team so two deals are graded identically. The sequence below is one to two hours of analyst work per team plus review, and it ends in a one-page memo. Sources disagree on whether to verify before scoring capability or the reverse; I verify first because a contradicted claim changes what the capability score is even measuring.

The pre-partner-meeting grading procedure

  1. Frame the deal and freeze the dimensions
    Fix the six dimensions and their weights in writing before searching any founder name. This blunts pedigree anchoring by committing you to the variables first.
  2. Identify and disambiguate the humans
    Resolve each founder to a canonical LinkedIn, GitHub, Google Scholar, and any prior-company registration. End with a person-by-person map and no name collisions.
  3. Verify the claimed record against public registries
    Cross-check employment and dates, degrees, prior companies, patents, publications, litigation, and regulatory history. Mark each claim verified, unverifiable, or contradicted.
  4. Score domain and execution evidence
    Grade industry-experience match and build evidence such as commits, shipped products, and publications. Cite the specific artifact behind each score.
  5. Score cofounder alignment from public traces
    Read tenure overlap, prior shared work, equity signals, and role clarity. Produce an alignment score flagged as the highest-uncertainty dimension.
  6. Apply the false-signal controls
    Re-run every score with pedigree and media-list credit stripped out. Output an adjusted score plus a list of load-bearing unverified claims.
  7. Write the go/hold/pass memo with a verify-later list
    Convert the weighted score to a call and list exactly which claims references and formal checks must confirm. End with a one-page memo defensible in IC.

From public evidence to a defensible call

  1. Disambiguate
    Resolve each founder to canonical public profiles
  2. Verify
    Mark each claim verified, unverifiable, or contradicted
  3. Score
    Grade the six dimensions with cited artifacts
  4. Debias
    Strip pedigree and media-list credit, re-score
  5. Decide
    Weighted score becomes go/hold/pass plus a verify-later list
Verification precedes scoring so a contradicted claim reshapes the score, not the memo.

How this goes wrong: the false-signal traps

Most bad team calls come from a short, well-documented list of traps. Each one is a false positive, so treat this section as the checklist you run against your own scores before the memo.

Pedigree substituted for capability. A Stanford or ex-FAANG line scores the team "strong" with no shipped artifact behind it. This is your default error: 70% of US VCs hold elite-institution degrees and demonstrably favor similar founders, and fund-level data shows one firm directing 67% of its seed investments to Ivy-League-Plus alumni and another 63%. Strip the school and employer credit, re-score on verifiable output only, and the bias becomes an edge, because it is a market-wide mispricing.

Media-list halo. "Forbes 30 Under 30" read as validation. Treat lists as unverified. Multiple honorees have faced fraud charges, including an AI startup founder, a Harvard spinout, charged with defrauding investors of $10M.

Unverifiable claim scored as verified. A title or date accepted because LinkedIn displays it. Mark it contradicted or unverifiable and route it to the verify-later list; roughly a third of applicants inflate their education.

Cofounder alignment inferred from absence of conflict. No public friction read as "aligned," when conflict is private by nature. Score alignment as low-confidence and require a downstream reference. This is the single largest failure cause at 65%.

GitHub or commit theater. High commit volume mistaken for capability. Inspect substance and authorship, not counts; forks and trivial commits inflate activity.

Serial-founder overcredit. "Third-time founder" scored top marks. Outcomes decline after the second venture from 34% to 28% to 22%, so weight domain fit over repetition.

Registry silence read as clean. No PACER or EDGAR hit taken as integrity, when absence of a federal record covers neither state courts, settled or unfiled disputes, nor foreign jurisdictions.

Pedigree bias is not just your error; it is a market-wide mispricing you can arbitrage with a rubric that scores verified output.

Turning the score into a call you can defend

The output is a weighted score converted into go, hold, or pass, paired with a list of exactly which claims references and formal checks must confirm downstream. Reference calls and formal background checks are not part of this gate. Formal diligence typically begins after investors express serious interest and runs two to six weeks, and reference calls are usually a term-sheet-stage activity. Your job here is to decide whether the deal earns that time, and to hand the downstream checkers a precise list rather than a vague "kick the tyres."

A defensible memo does three things: it states the weighted score and the artifact behind each dimension, it names the load-bearing unverified claims, and it says what the go/hold/pass call would flip on. Use this skeleton verbatim.

One-page team grade memo
TEAM: <company> / founders <names>
CALL: Go / Hold / Pass

SCORES (weight x score):
- Domain / industry fit: __  (artifact: ____________)
- Founding-team track record: __  (artifact: ____________)
- Technical build evidence: __  (artifact: ____________)
- Cofounder alignment: __  [LOW CONFIDENCE]  (artifact: ____________)
- Integrity / verification: __  (registries checked: ____________)
- Market pull: __  (artifact: ____________)

WEIGHTED SCORE: __ / 100
DEBIASED SCORE (pedigree + media credit stripped): __ / 100

LOAD-BEARING UNVERIFIED CLAIMS:
1. ____________  (route to: reference / formal check)
2. ____________  (route to: reference / formal check)

THIS CALL FLIPS IF: ____________

Fill each dimension with a score and the cited artifact; the verify-later list is the handoff to references.

Two markets frame how much supply you are choosing from, which matters when a partner asks whether this team is the best available rather than merely acceptable.

MarketFounder-titled profilesRatio vs UK
United States656,1394.08x
United Kingdom160,6561.00x

Against that supply, the framework is a filter, not a verdict. It advances or stops a deal and tells the next stage what to prove.

Keeping the framework honest over time

Re-run the debiasing control on your own past calls, not just the live deal, because that is how you learn whether your rubric is measuring capability or prestige. Before you sign off any grade, work the checklist below.

Before you call it graded

  • The rubric and weights were written before any founder name was searched
  • Each founder is resolved to canonical profiles with no name collisions
  • Every resume claim is marked verified, unverifiable, or contradicted
  • Each dimension score cites a specific public artifact
  • Cofounder alignment is scored and explicitly flagged low-confidence
  • Scores were re-run with pedigree and media-list credit stripped
  • Load-bearing unverified claims are on the verify-later list with a route
  • The memo states what would flip the go/hold/pass call

To keep the model current, track two things rather than any single number. First, whether your go calls survive the downstream reference and formal checks; a pattern of contradictions there means your verification pass is too shallow. Second, whether your debiased scores diverge from your raw scores in a consistent direction; if stripping pedigree keeps dropping the same teams, your rubric is still leaking prestige credit despite the frozen weights. The evidence base itself shifts, so re-check the mechanism, not the figure: when a new dataset revises how much team predicts outcome, adjust the cap on team weight, not the whole framework. The dimensions are stable. The weights are the thing you keep honest.

Refolk removes the two most manual parts of this work: disambiguating founders to canonical public profiles and finding the narrow, verifiable populations, such as founders with granted patents or prior senior roles at a named lab, that your build-evidence and domain scores depend on. The judgement stays yours; the sourcing and disambiguation stop eating the hour before the partner meeting.

Questions practitioners ask

How much should the team dimension weigh versus market and product?

Weight the team heavily but cap it. Venture investors cite the team as most important 47% of the time, yet a study of over 8,000 sourced deals found team predicts initial funding while market and product better explain long-term success. Practitioner framing rarely drops team below 30%. I keep the composite team-related weight meaningful but never let it exceed the combined market and execution weight, so a charming team cannot carry a weak market past your gate.

Which public sources actually verify a founder's background?

Use PACER and CourtListener for federal court records including bankruptcies and fraud, SEC EDGAR for company filings and enforcement, FINRA BrokerCheck for securities licensing and discipline, USPTO for patent ownership, Google Scholar for publications, and Secretary of State registries for prior-company existence. Each proves only the existence or absence of a public record. None capture private disputes, unfiled conflicts, foreign jurisdictions, or performance inside a prior role.

Where do reference and background checks belong relative to this scoring?

Downstream. Formal diligence typically begins after investors express serious interest and runs two to six weeks, and reference calls are usually a term-sheet-stage activity. This framework covers the earlier pre-partner-meeting moment when you have only public evidence. Your output is a scored call plus a verify-later list that tells references and formal checks exactly what to confirm.

How do I score cofounder fit when conflict is private?

Score it as low-confidence by design. Cofounder conflict drives 65% of high-potential startup failures but leaves almost no public trace, so the absence of visible friction is not evidence of alignment. Read tenure overlap, prior shared work, role clarity, and equity signals, then flag alignment as the dimension most in need of a downstream reference call rather than scoring it high because nothing bad appeared.

Is a Forbes 30 Under 30 listing or an elite degree a reliable signal?

No. Treat media lists and pedigree as unverified. Multiple honorees have been charged with fraud, including an AI startup founder accused of defrauding investors of $10M. About 33% of applicants admit lying about education and only 34% of employers verify credentials. Since 70% of US VCs favor founders with similar elite backgrounds, stripping that credit and re-scoring on verified output is both a bias control and an edge.

Read next