Building a Living Sector Map for Deal Sourcing
You will build a thesis-driven sector map, prove its coverage with a recall test, keep it current on a fixed cadence, and rank who to contact first.
A sector map is the artifact a thesis-driven investor keeps open all week: a maintained list of every company in a defined space, the founders behind each one, how each fits the thesis, and who to call first. This guide is for early-stage investors, platform and talent partners at funds, and angels who source against a stated thesis. It gives you the queries, the fields, the coverage test, and the maintenance loop that turn a one-time landscape slide into a living sourcing tool.
Most published advice on sector mapping stops at "define your thesis and use your network." That produces a static PowerPoint. What you need is a map that provably covers the thesis, tells you honestly where the whitespace is, stays current as founders move, and hands you a ranked contact list on any given Tuesday. That is the job here.
What a complete sector map actually requires
A complete market map is defined by how it was built, not by how many companies it contains. It requires a business-description-based search rather than a list of known competitors, coverage of long-tail private operators, verified ownership data, scale proxies, timing signals, and a live connection to your sourcing workflow.
Row count is a trap. A 2,000-row map can still miss an entire sub-segment if it was built from headlines. The discipline is to assemble from many sources, then test coverage against something you know is complete.
The layers of a living sector map
- Contact short listTop-ranked rows with a warm path and a next step
- Thesis-fit scoreFixed rubric ranking every row against the boundary
- Founder linksNamed, verified people behind each company
- Company setDeduplicated, source-tagged operators from business-description search
- Thesis boundaryOne-line filter on stage, sector, geography, check, model
The sources you assemble from are wide: referrals, founder networks, accelerators, universities, communities, market maps, databases, inbound pitches, and targeted outbound. The mistake is to lean on generic large language models, which return the well-known names that show up in prominent industry articles. That gives you a press-ranked map, and press-ranked maps systematically under-cover the exact targets an investor wants, because nearly half of companies that get acquired are bootstrapped at the time, with no VC backing and little coverage.
Assembling the company set from business descriptions
Assemble the company set by searching on what companies do, not on names you already know. Business-description search surfaces the long-tail private operators that named-competitor lists never reach, and every row lands with a source tag so you can audit where coverage is thin.
The named-competitor approach fails on its own terms: it can only find companies adjacent to ones you already named. Business-description search asks the opposite question, "who does this thing," and returns operators you have never heard of. That is where whitespace lives.
Do this across the full source stack in parallel:
- Directories and databases for structured firmographics and stage.
- Accelerator and community rosters for early operators before press exists.
- Registries for freshly incorporated companies inside the boundary.
- Open build signals on GitHub and patent filings for technical founders who have not launched.
- Inbound and referrals, tagged by channel so you can measure source quality later.
Deduplicate on the way in and keep the source tag on every row. When you later find a coverage hole, the source tag tells you which channel to widen. A tool like Refolk does the business-description search directly - you describe the operator profile in plain English and it returns matching people and companies across GitHub, LinkedIn, and the open web, which collapses the manual cross-referencing this step otherwise demands.
Measuring coverage against a known-complete benchmark
Coverage is measured, not asserted, by rebuilding a universe you know is complete and computing recall. Rebuild one Y Combinator batch inside your thesis boundary using only your own queries, then divide the rows you found by the official batch size. That fraction is your recall, and it is the single most honest number your map produces.
The YC directory works as a benchmark because it is known-complete and free. It lists all Y Combinator companies since 2005, filterable by batch and industry, and updates with every new cohort. If your thesis touches a sector YC has funded, you have a ground-truth list to test against.
| Metric | Value | Source |
|---|---|---|
| Known-complete universe | 5,668 YC companies | flowjam.com |
| Still operating | ~87% | flowjam.com |
| Unicorns / IPOs | 82 / 17 | flowjam.com |
The critical move is to compute recall per sub-segment, not overall. A map can score 90% recall on funded, famous names and 30% on the unfunded long tail, and the average hides the failure. Segment your recall by stage and by press profile. If recall is high only on the companies that make headlines, your map is press-biased and you need to widen your registry and community sources, not your database subscriptions.
The YC-batch recall test
- 5,668Official batch
Full YC universe, ground truth
- subsetIn your boundary
Rows matching your thesis filter
- your rowsRebuilt by your queries
What your sourcing actually found
- your rows / boundary subsetRecall
The number you report and defend
Linking every company to its founders
Every active row needs at least one named, verified founder, because you source people, not logos. Attach founder names and profile links, and for rows you pulled from a registry, cross-reference the officer name to a real profile before you trust it.
Directory-style sources ship founder identity: each company profile carries its founders with LinkedIn and Twitter/X links, plus company social links including GitHub and Crunchbase. Use those as the backbone. For registry-derived rows there is no built-in profile, so you must do the cross-reference yourself: match the officer name on the filing to a person online and confirm they are the operating founder, not a formation agent.
This is also where geography matters more than a map's aesthetics suggest. Founder pools are wildly asymmetric across markets, and a thesis that splits sourcing effort evenly across regions will over-cover thin ones.
| Market | Financial-services founders / co-founders | Ratio vs Germany |
|---|---|---|
| United States | 15,663 | 15.0x |
| United Kingdom | 4,270 | 4.1x |
| Germany | 1,043 | 1.0x |
Counts from Refolk's index of professional profiles; the ratio column is derived by dividing each count by the Germany count. In Refolk's index, the US financial-services founder pool is 15 times the size of Germany's and about four times the UK's. A "US plus Europe" thesis that allocates equal sourcing hours to each geography is mispriced: effort should track the size of the founder pool, not the symmetry of a slide.
Watching for new company formation inside the boundary
A new-formation watch catches companies before press coverage exists, which is where the timing edge lives. Monitor registries, domains, GitHub, patents, and hiring inside your thesis boundary, and route every hit into the map with a dated first-signal field.
The mechanism is simple and public. Any company that intends to employ people, sign contracts, or accept investment must incorporate, and that incorporation is recorded in a public government registry accessible within days of filing. In the UK, Companies House publishes SH01 share-allotment forms that can reveal stealth-stage funding before any announcement. A researcher who stops publishing papers and files a provisional patent signals a research-to-IP transition that often precedes company formation by three to six months.
Watch formation signals, not funding announcements, and you buy a three-to-six-month head start on the round.
Not every signal is a company. A lateral job move is not a startup, and a sabbatical is not a stealth launch. Read a registration or a network signal to confirm before you add a row. A first-time incorporation from someone who held a VP role at a funded startup is a far stronger signal than a generic LLC registration, so weight the person's history, not just the filing.
The step-by-step build
Run the build in order the first time, then let the refresh loop take over. The whole first pass is roughly a working week for an analyst plus a day of partner and reviewer time; the artifact it produces then persists.
Build the map end to end
- Define the thesis boundaryWrite explicit inclusion rules on stage, sector, geography, check size, and business model. Done means a one-line filter a colleague can use to return the same universe.
- Assemble the company setPull from referrals, accelerators, communities, databases, inbound, and outbound using business-description search, not named competitors. Done means a deduplicated master list with a source tag on every row.
- Measure coverage against a benchmarkRebuild one YC batch inside your boundary from your own queries and compute recall per segment. Done means a stated coverage percentage and a list of misses.
- Link each company to its foundersAttach founder names and profile links; cross-reference registry officer names to profiles. Done means every active row has at least one named, verified founder.
- Add a new-formation watchMonitor registries, domains, GitHub, patents, and hiring inside the boundary. Done means alerts route in with a dated first-signal field.
- Score thesis fit and rankApply a fixed rubric plus disqualifiers; two reviewers score a 20-row calibration sample, resolve disagreements, then split the rest. Done means a ranked short list and a recorded agreement score.
- Set the refresh loopRun a monthly people and status refresh, a quarterly full re-verify and thesis-mix review, and an annual retrospective. Done means a calendared cadence with an owner per field.
- Produce the contact short listTake top-ranked rows and attach the warmest path plus a next step. Done means a contact-this-week list with owner and action.
There is a legitimate disagreement about order. Thesis-driven funds often front-load a deep white paper before broad list-building: they choose a sub-sector per quarter, pull together a landscape map, run industry calls, and reach out to the ten to fifteen top players. That is a valid variant when the thesis itself is still forming. The version above assumes the thesis is set and the job is coverage.
Scoring thesis fit and ranking who to contact
Thesis fit is scored on fixed dimensions so two people rank the same company the same way. Score alignment across stage, sector, check size, geography, and business model, stack the weighted dimensions into a 100-point rubric, and apply hard disqualifiers on top. High scorers close faster and expand more, so the rank is worth getting right.
The columns you carry make the score auditable. Documented VC deal-flow practice records the company, source, sector, stage, owner, thesis fit, next step, diligence notes, and final decision. Add a last-verified date and a first-signal date, and you can measure pipeline health with qualified rate, first-meeting conversion, diligence conversion, source quality by channel, time in stage, and pass reasons.
Stage (20): Company is at [seed/Series A] per last funding or headcount signal. Yes = 20, adjacent stage = 10, else 0. Sector (20): Business description matches [defined sub-sector]. Yes = 20, tangential = 10, else 0. Check size (15): Round size fits [check range]. In range = 15, one band off = 8, else 0. Geography (15): HQ or founder base inside [geographies]. Yes = 15, else 0. Business model (15): Model is [B2B SaaS / marketplace / etc]. Yes = 15, hybrid = 8, else 0. Founder signal (15): At least one founder has [prior exit / senior role at funded startup / patent history]. Yes = 15, else 0. Disqualifiers (auto-zero): Outside legal/ethical mandate, direct portfolio conflict, or confirmed dead.
Phrase every item as an observable yes/no. Replace the bracketed thesis values with your own boundary before use.
Consistency between reviewers is measurable, not a matter of taste. Have two reviewers score a 20-row calibration sample, then compute the inter-rater reliability as an ICC. A documented rubric reached an overall ICC of 0.7, which counts as good agreement on the standard scale: below 0.5 is poor, 0.5 to 0.75 is moderate, 0.75 to 0.9 is good, and above 0.9 is excellent. That result came from concrete, behaviorally anchored criteria. Broad global ratings reproduce poorly, so a yes/no test will out-agree a 1-to-5 quality slider every time.
Keeping the map current: cadence tied to field decay
Refresh cadence is driven by how fast fields decay, and the fastest-decaying fields are the ones about people. Refresh people and status fields monthly, re-verify firmographics quarterly, and run an annual retrospective that connects sourcing data to fund outcomes. This split is the difference between a living map and a stale one.
The decay is uneven, which is the whole point. Company addresses barely move; job titles and roles move constantly. Set your cadence to the fastest-mobility field, not the average.
| Field / measure | Decay rate | Source |
|---|---|---|
| Overall B2B contact record | 22.5% / year (2.1% / month) | HubSpot via datamagnet.co |
| Contact record (ZoomInfo) | 25-30% / year | datamagnet.co |
| Email addresses | 3.6% / month | landbase.com |
| Role / job change | 30-40% / year | landbase.com |
Job changes drive the most decay. Average job tenure is 2.8 years, so 30 to 40% of contacts change roles annually. A row can look completely live while the founder left months ago, which is why every row carries a last-verified date: anything older than the cadence is presumed wrong until re-checked. This is where the tiered review structure earns its keep - monthly pipeline reviews catch operational drift, quarterly strategy reviews reveal whether the sourcing mix still matches the thesis, and the annual retrospective ties it all to outcomes.
The refresh loop
- MonthlyRe-verify founder and status fields; re-stamp last-verified dates
- QuarterlyFull re-verify of firmographics plus a thesis-mix review
- AnnuallyRetrospective connecting sourcing data to fund outcomes
- ContinuousNew-formation alerts flow in with dated first-signal fields
For hot sectors, quarterly-only cleanup is not fast enough: by the time you clean, 15 to 20% has already decayed again. Track the decayed-row rate between refreshes, and if it exceeds your tolerance, shorten the cadence. Manual verification is not a substitute for cadence, either - one exercise reached 91% accuracy but took 143 hours on 10,000 contacts, so verification effort has to be scoped, not applied everywhere at once. Refolk's re-run-the-query model is the cheap version of a monthly refresh: the same plain-English search that built a segment can rebuild it, and the diff shows you who moved.
How this goes wrong
The map fails in predictable ways, and the checks for each are cheap. The weakest point is almost never company coverage - it is the people fields and the method behind the counts.
- LLM-built maps look complete but miss the long tail. Generic models return the names that show up in prominent industry articles. Check: run the YC-batch recall test; if recall is high only on funded or famous names, the map is press-biased.
- Coverage measured by count, not method. A 2,000-row map can miss a whole sub-segment. Check: report segment-level recall, never total count.
- Stale people fields masquerading as current. A row looks live while the founder left months ago. Check: a last-verified date per row; anything older than the cadence is presumed wrong.
- Employment signal mistaken for a new company. A lateral shift or sabbatical read as a stealth launch. Check: require a business registration or network signal before adding a formation row.
- Registry noise. A generic LLC registration scored the same as a founded-startup VP's first incorporation. Check: cross-reference the officer name to a profile before adding.
- Rubric scores that don't agree. Subjective items reproduce poorly. Check: compute ICC on a calibration sample; below 0.5 means rewrite the rubric.
- Quarterly-only cleanup too slow for hot sectors. By cleanup time, 15 to 20% has decayed again. Check: track the decayed-row rate between refreshes and shorten the cadence.
The through-line: companies persist, founders move, and method beats volume. A map that is big and current-looking can be wrong in exactly the places an investor cares about most.
Keeping the artifact alive
The map is only worth building if it stays usable, so the last thing you set up is the discipline that keeps it honest. Before you call the build done, verify the map against the checks that separate a living artifact from a slide.
Before you call the map done
- The thesis boundary is a one-line filter a colleague can reproduce.
- Every row carries a source tag showing where it came from.
- You have a stated recall percentage from a YC-batch test, computed per segment.
- Every active row has at least one named, verified founder.
- New-formation alerts route in with a dated first-signal field.
- Reviewers hit at least 0.5 ICC on the 20-row calibration sample.
- Every row has a last-verified date and the monthly refresh is calendared with an owner.
- A contact-this-week list exists with owner, warmest path, and next step per row.
To keep it current, run the refresh loop on the calendar you set and re-run the recall test each quarter against a fresh benchmark batch, because a map that covered the thesis last quarter can drift as the sector grows. When a founder-status field decays past your tolerance between refreshes, that is the signal to shorten the cadence for that segment, not to clean once and hope. The map earns its keep on the Tuesday you open it, filter to the top-ranked rows inside the boundary, and know exactly who to reach first and why.
Questions practitioners ask
How do I know my market map is actually complete?
You cannot know from the row count. Rebuild one known-complete universe inside your boundary and compute recall. A Y Combinator batch works because the official directory lists all 5,668 companies since 2005, filterable by batch and industry. Run your own queries against that boundary, match rows to the batch, and report recall by sub-segment. High total recall that collapses on unfunded or non-press names means your map is press-biased.
How often should I refresh a deal sourcing watchlist?
Refresh people and status fields monthly and re-verify firmographics quarterly. Cadence is set by decay: B2B contact records lose about 22.5% of accuracy a year, and 30 to 40% of contacts change roles annually because average job tenure is 2.8 years. In hot sectors, quarterly-only cleanup is too slow, since 15 to 20% can decay again before you finish, so track your decayed-row rate and shorten the cycle if it exceeds tolerance.
What columns belong in a sector map for VC?
Documented deal-flow practice records the company, source, sector, stage, owner, thesis fit, next step, diligence notes, and final decision. Add a last-verified date per row and a dated first-signal field for new-formation rows. These two dates are what let you separate live records from stale ones and measure your timing edge, which the standard nine-column layout omits.
How do I catch stealth or brand-new companies inside my thesis?
Watch formation signals, not funding announcements. Incorporation is recorded in a public government registry within days of filing, and in the UK, Companies House SH01 share-allotment forms can reveal stealth funding before any press. A research-to-patent transition often precedes company formation by three to six months. Cross-reference the officer name to a real profile before adding a row, since a founded startup VP is a far stronger signal than a generic LLC.
How do I make two reviewers score thesis fit the same way?
Use behaviorally anchored yes/no items, not a broad 1 to 5 quality slider. A documented rubric reached an overall inter-rater reliability of 0.7 (ICC) using concrete criteria, which counts as good agreement. Have both reviewers score a 20-row calibration sample, compute the ICC, and if it falls below 0.5 rewrite the vague items before splitting the remaining rows.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.