The Technical-Founder Build Read: Advance, Probe, or Pass
You will score a founder's public repos, papers, and shipped products across five dimensions and reach a defensible advance, probe, or pass call before requesting repo access.
You are diligencing a technical or AI founder and need to decide, from public work alone, whether they can build the thing they are pitching, before you spend goodwill requesting repo or data-room access. This guide is for early-stage investors, platform and talent partners, and angels who want to triage which founders to advance. It scores only what is visible before any private access and turns that into an advance, probe, or pass verdict.
Most writing on this topic sits on the other side of the table. It is either founder-side prep for the data-room audit or a description of the full review a fund's technical advisor runs once a repo invite exists. This guide covers the gap: the pre-access read, where you have public repos, papers, prior shipped products, and talks, and nothing private yet.
What public artifacts actually predict build ability
The strongest public signals are proof of work, not claims. A merged pull request into a serious third-party project, a maintained repo with real docs, consistent recent commit activity, and collaboration conduct in issues and reviews predict build ability far better than anything a founder writes about themselves.
The reason is simple and worth stating as a rule. A GitHub repository is evidence; a PDF is a claim. A paper or a talk proves the founder can frame a research problem and present it. Neither proves they can ship a system that survives contact with production.
Signals that are hard to fake share a property: they require sustained effort over time inside a system the founder does not control. Recent activity shows a current habit of shipping rather than one burst a year ago. Project ownership and merged PRs into serious projects show others accept the work. Collaboration history in comments and reviews shows the founder can operate in a moving codebase.
Common thresholds get cited as shorthand: 100-plus stars, 20-plus forks, 50-plus followers, and 50-plus PRs per year. Treat these as context, not proof. Every one of them misleads in a predictable way. Stars and forks are vanity-gameable. A flashy personal repo can mask an inability to work in a system someone else is also changing. Commit history can be backdated or squashed to look busier than it was.
The five dimensions of the build read
Score the founder on five dimensions, each read from a specific artifact, each with a known way it lies. The dimensions below are the model; the rest of the guide tells you how to score and combine them.
| Dimension | Public artifact that proves it | What it looks like when it lies |
|---|---|---|
| Contribution depth | Merged PRs into third-party projects | 2,000-star demo repo, no sustained commits |
| Product defensibility | Shipped product, eval suites, data | Thin wrapper the prompt-paste test exposes |
| Claim reproducibility | Benchmark card, public demo | Score with no N-shot, temperature, or test-set |
| Governance hygiene | Git history, dependency licenses | Secrets in history, copyleft core dependency |
| Pedigree vs shipping | Prior shipped product, lab history | Frontier-lab logo, nothing personally shipped |
Each dimension answers a different question. Contribution depth asks whether the founder can write code others accept. Defensibility asks whether the product is more than a prompt. Reproducibility asks whether the headline number is real. Governance asks whether the thing is safe to buy. Pedigree versus shipping asks how much of the signal is borrowed from a brand name.
The load-bearing insight across all five: the repo rarely kills the deal, the unanswered question does. Technical diligence rarely blocks a round outright. It routinely re-prices one after the term sheet has removed the founder's leverage. Unanswered questions read as unmanaged risk, and unmanaged risk prices deals. So your pre-access read is really hunting for the questions you will later need the founder to answer.
Genuine depth versus infrastructure theater
The durable public test for an AI founder is defensibility beyond the prompt. Ask whether a technical user could get roughly 80% of the product's output by pasting the core prompt into ChatGPT or Claude. If yes, it is a thin wrapper. If no, something proprietary sits in the way.
The something that sits in the way is what you are scoring. Depth markers are proprietary data, workflow lock-in, action outside the chat box, and shipped eval suites. Real products change systems of record. They file tickets, update claims, route approvals. Action depth separates a tool from a talking screen. Building production-grade LLM features requires extensive evaluation, monitoring, and quality control, so a founder who ships a serious eval suite is doing real engineering whether or not they trained a model.
The cautionary tale sits at the other extreme. Builder.ai's "AI" was 700 human engineers in India, and "Natasha" was a brand name for a staffing company. That is infrastructure theater in its purest form: the pitch describes a system, the system is people. The public tell is almost always absence. No repos, no eval suites, no action depth, no reproducible demo, just marketing. Y Combinator's own Request for Startups tells founders to stop bolting a chatbot onto a 2010 UI and to rethink the workflow instead. The founders worth advancing are already doing that, and it shows in what they have shipped.
Reproduce one claim before you trust it
Reproducing a single public benchmark is the cheapest kill-test you have, and most founders have never been asked to pass it. Historical irreproducibility in machine learning is high enough that you should assume a headline number is wrong until you can reconstruct it.
| Study or source | Irreproducible or divergent rate |
|---|---|
| Raff 2019 (255 ML papers) | 36.5% not reproduced |
| Gundersen and Kjensmo 2018 | 74% could not be reproduced |
| OLMES vendor-score replication gap | typically 1 to 5% |
Read the table as two different failures. The first two rows are papers that could not be reproduced at all. The third row is the gap you see even when reproduction succeeds: independently re-running vendor-reported scores usually yields different numbers, more often lower, by 1 to 5%. A survey of questionable ML practices catalogs 43 distinct ways reported results get undermined, so there is a lot of room between an honest number and a misleading one.
What makes a benchmark unverifiable is missing methodology. A score without N-shot setting, chain-of-thought flag, temperature, max output tokens, and test-set version cannot be checked. When those are absent, you do not need to reproduce anything to move the founder from advance to probe. The absence is the finding.
A benchmark score with no methodology card is not a result, it is a hope with a decimal point.
When you can reproduce from public artifacts, do it. A gap inside the 1 to 5% band is normal and fine. A gap much larger than that, or a score you cannot get near, is a load-bearing question for the founder.
Where this read lives in the diligence timeline
Everything in this guide is pre-access and belongs before the formal technical review. Technical review is no longer a late event; it shows up at seed and Series A, and AI-assisted codebases are pulling it earlier still.
| Stage | Overall window | Technical form |
|---|---|---|
| Pre-seed | 1 to 2 weeks | conviction, prototype review |
| Seed | 2 to 4 weeks | roughly 90-minute CTO call |
| Series A | 4 to 6 weeks | formal audit, 30 to 50 page report |
At seed the technical slice is light, often a single 90-minute CTO call. At Series A and beyond it becomes a formal audit by an outside firm, frequently a 30 to 50 page report, and cybersecurity and AI or ML startups hit that rigor earlier. The whole diligence process runs 2 to 6 weeks, with seed averaging 2 to 3 weeks and Series A-plus 4 to 6 weeks. The public build read happens in the first days of that window, before you spend relationship capital asking for a repo invite.
There is a real disagreement about sequencing worth knowing. Interview-led diligence talks to the team first; code-led diligence reads the repository first and talks second. Only one of those finds the secret in the git history. For the pre-access read, you are code-led by necessity: you have the public artifacts and not yet the team's time.
Sourcing the pool you read from is its own job, and Refolk turns the plain-English version of that query into a ranked list across GitHub, LinkedIn, and the open web. The point is to start your read on founders whose public footprint already clears the first dimension, rather than reverse-engineering it one profile at a time.
The eight-step pre-access procedure
Run these eight steps in order. The first two are investor work, the middle are reviewer work, and the last two bring it back to the investment call.
Score the public footprint, pre-access
- Scope the hardest claimWrite down what the founder says they built and the single capability you are underwriting. Done when you have one sentence naming the hardest technical claim.
- Enumerate the public footprintPull GitHub org and personal profiles, published papers, conference talks, prior shipped products, and patents. Done when you have a list of artifacts with URLs.
- Read repos for behavior, not metricsFocus on patterns and habits, not every line: merged PRs into third-party projects, commit message quality, and bus factor. Done when you have a contribution map by author.
- Run the wrapper and depth testApply the prompt-paste test and look for action depth, proprietary data, and shipped eval suites. Done when you have a thin or thick classification with evidence.
- Reproduce one claimAttempt to reproduce the headline benchmark or demo from public artifacts, noting missing N-shot, temperature, or test-set version. Done when the result is reproduced, partial, or not.
- Scan for kill-signalsScan git history for secrets, map the dependency license graph, and check for admin endpoints without auth. Done when you have a pass or flag list.
- Weight pedigree, domain, and shippingScore each sub-dimension explicitly and note the substitution logic for the stage. Done when you have three sub-scores.
- Reach the verdictCall advance, probe, or pass with the single load-bearing reason named. Done when you have a written rationale someone else could challenge.
Experienced reviewers rarely examine every line of code at step three. They focus on patterns, habits, and evidence of how someone works with modern tooling. You are building a contribution map, not auditing a codebase.
The pre-access build read
- ScopeName the one hard claim being underwritten
- EnumerateList every public artifact with its URL
- ReadMap contributions to authors, not metrics
- Test and reproduceClassify depth, reproduce one claim
- VerdictAdvance, probe, or pass with one reason
Scoring the verdict: advance, probe, or pass
Turn the five dimensions into one of three calls. The verdict is a function of how many dimensions clear and whether any single dimension produced a kill-signal.
- Advance when contribution depth and defensibility both clear, the headline claim reproduces or has a complete methodology card, and no kill-signal is open. You are confident enough to request repo access.
- Probe when the footprint is promising but one dimension is unresolved: a claim you could not reproduce, a bus-factor concentration, a pedigree-heavy profile with thin shipping. You advance only after the founder answers the specific open question.
- Pass when a kill-signal is confirmed and material, or when the product fails the wrapper test with no depth markers, or when the entire read rests on a logo with nothing personally shipped behind it.
Depth against defensibility
On the pedigree sub-score, there is no publicly established weighting. In deep tech, pedigree is heavily indexed, especially in AI, and founder credibility can substitute for ordinary revenue proof. Five frontier-AI seed deals raised $9.11B at a $1.03B median, driven by credibility rather than traction. That is exactly the highest-risk place to lean on a logo, because public shipping evidence is thinnest there. A Lunar Ventures view cuts the other way: deep domain expertise can beat a polished entrepreneurial track record. Score pedigree, domain, and shipping separately and write down which one you are letting carry the verdict.
How the build read goes wrong
The read fails in predictable ways, and most of them are false positives that advance a founder you should have probed or passed. Each failure mode below pairs the trap with the check that defeats it.
| Failure mode | False positive it creates | The check that defeats it |
|---|---|---|
| Stars and forks as proxy | Viral demo repo, no sustained commits | Contribution graph recency, merged third-party PRs |
| Pass on clean code | Governance gap hidden behind tidy code | Can the founder explain architecture in decision terms |
| Benchmark at face value | "94.2% on MMLU" with no settings | Reproduce it or demand the methodology card |
| Bus-factor misread | Solo repo looks productive, one departure kills it | Modules touched by exactly one contributor |
| Pedigree halo | Frontier-lab logo, no shipped product | A prior thing they personally shipped and maintained |
| Secrets look historical | "That key was rotated" | Scan git history, not just current HEAD |
Two of these deserve extra weight. First, most founders do not fail because the code is bad. They fail because they cannot answer the questions, and "our developer handled that" shifts the reviewer to evaluating the governance gap instead of the code. Clean code is not a pass; it is the start of the conversation.
Second, secrets in git are a live and worsening base rate. Developers pushed 28.65 million new secrets to public GitHub in 2025, up 34% year over year. A published technical due diligence report flagged a GPL v3 core PDF library and AWS root keys that lived solely on the CTO's local machine as high-risk findings, and separately 68% of audited codebases contain open source license conflicts. Absence of any public secret hygiene is itself a signal, and presence is confirmed only by scanning history, not HEAD.
The departed-contributor trap is the quiet one. An impressive core module may have been written by someone no longer on the cap table. Check commit authorship against the current team and against IP assignments before you credit the founder with work another person did. This is the inverse of the pedigree halo and just as easy to miss.
The deep-engineering founder pool is thin and concentrated
The founders who clear this bar are geographically scarce, which raises the premium on verifying the few who exist. In Refolk's index of professional profiles, founder and owner-level people with both PyTorch and CUDA skills number 3,136 in the United States and 524 in the United Kingdom.
| Market | Founder/owners with PyTorch + CUDA | Share of US |
|---|---|---|
| United States | 3,136 | 1.00x |
| United Kingdom | 524 | 0.17x |
The US pool is roughly 6.0x the UK pool. For a UK or European deep-tech thesis, that concentration means you are competing over a thin set of genuinely qualified technical founders, so a disciplined public read is worth more, not less. In Refolk's index, the US cohort clusters at employers including StarTree, Axon, and Gecko Robotics, and the UK cohort at Greyparrot, TurinTech AI, Tenyks, and G-Research. Those are where the shippers come from, which is useful both for sourcing and for sanity-checking a claimed pedigree.
One caution on headline matching. A Refolk query filtering ML founders by frontier-lab keywords returned zero headline matches, because headline keyword matching is strict. Absence there is not evidence of absence. Verify pedigree through artifacts the founder produced, not through the brand name in their title line.
Before you call it: the pre-access checklist
Run this before you write advance, probe, or pass. Every item is a thing you can verify from public artifacts alone.
Pre-access build read
- The single hardest technical claim is written in one sentence.
- Every public artifact is listed with a URL: repos, papers, talks, prior products.
- Merged PRs into third-party projects are counted, not just stars and forks.
- Contribution authorship is mapped, so bus factor and departed-contributor risk are visible.
- The prompt-paste wrapper test has a thin or thick result with evidence.
- One headline claim is reproduced, partial, or not, with missing methodology noted.
- Git history, dependency licenses, and admin endpoints are scanned, not just current HEAD.
- Pedigree, domain, and shipping are scored separately, with the load-bearing one named.
- The verdict names one reason that someone else could challenge.
Keeping the read current
The method is evergreen, but three inputs drift and need re-checking. Re-run the secret base rate when you refresh your priors: the figure moves year over year and the direction has been up. Re-check the stage-timing table against your own recent deals, because AI-assisted codebases keep pulling technical review earlier. And re-pull the founder pool counts when you size a market, because the US-to-UK ratio is a snapshot of Refolk's index, not a fixed constant.
The discipline to keep is the sequence, not the numbers. Scope the claim, read the artifacts for behavior, test for depth, reproduce one thing, scan for kill-signals, then weight and call it. Do that before you request access and you will spend your private diligence hours on the founders who have already earned them, and you will walk into the CTO call already holding the one question that prices the deal.
Questions practitioners ask
Can I really judge a technical founder before getting repo access?
Yes, for a triage decision. Merged pull requests into serious third-party projects, commit recency, prior shipped products, published papers, and a prompt-paste wrapper test are all visible publicly. These let you reach an advance, probe, or pass call. What you cannot see publicly is the private codebase's governance, IP assignments, and architecture decisions, which is why the verdict is a triage, not a final underwriting.
What is the single fastest public kill-test for an AI founder?
Reproduce one headline claim. Historical irreproducibility runs 36.5% in Raff's 2019 study of 255 ML papers and 74% in Gundersen and Kjensmo's 2018 survey, and most founders have never been asked. A benchmark score without N-shot setting, chain-of-thought flag, temperature, max output tokens, and test-set version is unverifiable. If you cannot reproduce it from public artifacts and the methodology is omitted, that alone moves a founder from advance to probe.
How do I tell a real thick wrapper from a thin one?
Apply the prompt-paste test: if a technical user can get roughly 80% of the product's output by pasting the core prompt into ChatGPT or Claude, it is thin. Depth markers are proprietary data, workflow lock-in, action outside the chat box, and shipped eval suites. Real products change systems of record. Cursor and Harvey are built on foundation models and are not thin, so judge the data, lock-in, or distribution moat, not model ownership.
Do stars and forks mean anything?
They signal community attention but are vanity-gameable and routinely mislead. A viral demo repo can carry 2,000 stars with no sustained commits and mask an inability to work in a moving system. Cited thresholds like 100-plus stars and 50-plus PRs per year are context, not proof. Read the contribution graph for recency and merged PRs into third-party code instead, which are far harder to fake.
Should pedigree outweigh shipping history for an AI founder?
No fixed weighting is publicly established. In deep tech, pedigree is heavily indexed, and in frontier AI, founder credibility can substitute for traction entirely, with five seed deals raising $9.11B at a $1.03B median. That is precisely the highest-risk place to lean on a logo, because public shipping evidence is thinnest there. Always look for a prior thing the founder personally shipped and maintained before crediting the brand name.
When does the formal technical audit happen, and what runs before it?
Technical review now shows up at seed and Series A, with AI-assisted codebases pulling it earlier. Seed is usually a 90-minute CTO call; Series A-plus is a formal outside audit, often a 30-50 page report, within a 4-6 week window. Everything in this guide is pre-access: public repo secret and license scans, commit concentration, prior shipped products, benchmark reproducibility, and the wrapper test, all done before you request the data room.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.