The 160-Startup Repo Every AI Engineer Just Starred. Now Reverse It.
The awesome-ai-startups-hiring repo turned 160+ AI companies into a candidate wishlist. Every star is a public job-search signal. Here's how to source it.
You are trying to source AI engineers who actually want to move, and every channel you have (LinkedIn InMail, cold email, referral asks) treats intent as a guess. Meanwhile, thousands of engineers are publicly bookmarking a list of AI startups they'd rather work at, and almost no sourcing team is reading that list.
The repo is awesome-ai-startups-hiring, maintained by Vinit Shahdeo (engineering lead at zzazz, ex-Postman). It curates 160+ funded AI startups grouped by category and latest round, aimed at engineers who want to job-hop into AI. Every star, fork, and watch on that repo is a candidate raising their hand in public.
Why this repo is different from every other "awesome list"
Stars on a hiring-list repo are not a skill claim. They are a job-search claim. That inversion is the entire opportunity.
The standard guidance on GitHub sourcing (see Nexus IT Group's writeup) is correct but narrow: "stars and followers can add context, but they are supporting signals; they should never carry the evaluation on their own." That warning is about stars on someone's own projects, where a star means "I saw this and liked it."
Stars on awesome-ai-startups-hiring mean something else entirely. The repo's own instructions to visitors are explicit: browse the list, find teams building tech that excites you, head to their /careers pages, and cold-email the founders if you want to stand out. A star here is a bookmark on a job-search action. A watch is a subscription to "tell me when new AI startups are hiring."
That's not noise. That's intent, timestamped and public.
The category coverage is broad enough to matter
The 160+ startups on the list span:
- Agent infrastructure
- LLM inference and serving
- AI developer tools
- Data and retrieval infrastructure
- AI security
- Voice AI
- AI fintech
That is roughly the same taxonomy every AI-first founder is hiring against right now. If you're sourcing for any of those categories, the stargazer list is a pre-filtered pool of engineers who have already decided to leave their current job for something in your exact space.
The math on 5,000 stars
At 5,000 stars, and roughly 60% of stargazers with a public GitHub bio matchable to a LinkedIn, this one repo could surface about 3,000 warm, intent-tagged AI candidates. That is roughly the entire US Forward Deployed Engineer pool.
Here's the comparison against Refolk's index of professional profiles:
| Segment | Count | Top employers | Note |
|---|---|---|---|
| Forward Deployed Engineers, US | 3,068 | Palantir, Modal, Amp Code, Roboflow, Cresta | The archetypal AI-startup generalist role |
| AI / Applied AI / LLM / Agent Engineers, US | 5,026 | Distyl, Sandgarden, Develop Health, Hedgineer | ~1.6x the FDE pool |
| AI Engineers globally with LangChain, LlamaIndex, or RAG skills | 3,449 | Composio, PostHog, Procure AI, Thomson Reuters | Majority non-US |
| Startups on the hook list | 160+ | Repo maintainer's own count | Curated, funded only |
| New engineers on GitHub in 2025 | 36.2M | Octoverse via Pin | Fastest-growing dev community |
The scarcity is the story. Every AI startup on Shahdeo's list wants forward-deployed engineers who can talk to customers and ship. The entire US pool is 3,068 people, concentrated in NYC and SF, with Palantir still the top employer by a wide margin. If this repo hits 5,000 stars and even 5% of stargazers hold an FDE-adjacent title, that's 150 pre-qualified, intent-tagged candidates from one URL. You cannot buy that list.
The star-as-intent signal, decoded
A star on awesome-ai-startups-hiring is one of the highest signal-to-noise public actions an engineer can take on GitHub, because it has zero code overhead and binary intent meaning.
Compare it to the usual sourcing signals:
- Contribution graph activity. Tells you they code. Tells you nothing about whether they'd take a call.
- A star on a random project. "I found this interesting." No commercial meaning.
- A LinkedIn "open to work" banner. Public, but broad; not tied to a category.
- A star on
awesome-ai-startups-hiring. "I want to leave my current job for an AI-native startup." Category, intent, and timing in one click.
There is a reason B2B sales tools like GitLeads already treat stars as buying signals. Their pitch is that GitHub stars, forks, issues, and keyword signals turn into qualified developer leads. Sales figured this out years ago. Recruiting has barely caught up, partly because the classic advice ("stars aren't skill signals") kept people from noticing that stars can be intent signals in a different context.
Watchers are the smaller, better tier
A star is a bookmark. A watch subscribes the engineer to notifications every time a new AI startup gets added. Watchers are the tighter cohort:
- Star = "I might job-hunt eventually."
- Fork = "I want my own copy, maybe to add to it."
- Watch = "Ping me every time a new hiring AI startup shows up."
Almost no sourcing tool differentiates these three tiers. If you're building an outreach sequence, watchers deserve the top slot in your queue.
The 82% dark problem cuts the other way here
Only 18% of GitHub activity is public, per Pin's GitHub recruiting analysis, which normally makes GitHub a frustrating sourcing surface. Stargazer lists are the exception: they are 100% public by default.
That is why this angle works. Most of what you'd want to know about a GitHub engineer (private repo work, private org membership, DMs, "open to work" state) is hidden. But the stargazer endpoint (/repos/{owner}/{repo}/stargazers) is fully public and paginated, returning username, timestamp, and enough to link out. It is the rare GitHub interaction where the visibility default flips in the recruiter's favor.
That matters more as the surface grows. GitHub added 36.2 million new engineers in 2025, roughly one every second according to the Octoverse report. As the platform scales, the ratio of public-to-private activity gets worse, not better. Signals that stay public by default become disproportionately valuable.
Stars on a hiring-list repo are not a skill claim. They are a job-search claim. That inversion is the entire opportunity.
The workflow, then, is not "scrape GitHub and hope." It's:
- Pull the stargazer list for
vinitshahdeo/awesome-ai-startups-hiring. - Filter to accounts with a public bio, location, or linked personal site.
- Match those handles to LinkedIn, Twitter/X, or a personal domain.
- Cross-reference against title, seniority, and current employer.
- Rank by how recently they starred (fresh stars beat stale ones).
Step 4 is where most in-house recruiting teams stall. Manually stitching a GitHub bio to a LinkedIn profile takes 3 to 6 minutes per person, which is fine for 20 candidates and unworkable at 3,000. This is the exact gap Refolk closes: you describe the pool in plain English ("engineers who starred awesome-ai-startups-hiring in the last 30 days, currently at Series B+ US startups, titled Applied AI or FDE") and get a ranked shortlist with the identity resolution already done.
The RAG geography is inverted from the money
Recruiters who treat this repo as a US-only signal will miss the majority of the qualified RAG-skilled pool. That pool sits mostly outside the US.
Refolk's index shows 3,449 AI engineers globally list LangChain, LlamaIndex, or RAG as skills. The top regions are:
- Cape Town
- Kraków
- Poznań
- Bengaluru
- São Paulo
None of those are San Francisco. Which is a problem if your outreach template assumes the candidate is in Pacific time and expects a US-dollar offer. It's also an opportunity: the 160+ startups on Shahdeo's list skew seed-to-Series-B and are already remote-friendly, and the intent signal travels across geographies just fine. A star from a Kraków-based LangChain engineer on this repo is as legible as a star from an SF one.
There's a second cut worth making. The same index shows 5,026 AI, Applied AI, LLM, and Agent Engineers in the US, which means only about 69% of US AI engineers explicitly list LangChain, LlamaIndex, or RAG as a skill. Filtering the stargazer list on skill keywords alone will cut your pool by a third for no good reason. Filter on title and intent first; treat skill tags as tiebreakers.
The caveats a serious recruiter should flag
Two caveats will kill this play if you ignore them.
The funding data is AI-mined, not verified
Shahdeo's repo says so directly: funding data is mined using AI, spot-check before relying on it. If you're using the list to prioritize outreach ("only chase Series B and up"), verify each row against Crunchbase or PitchBook before you build a target account list. Do not send an email that opens with "congrats on your Series C" if the round was Series A eighteen months ago.
Star velocity, not star count
A two-year-old star means nothing. A star from last week is a live signal. Sort the stargazer list by timestamp descending and put anything older than 90 days at the bottom of the queue. For sourcing forward deployed engineers, the freshness window is even tighter: FDE-titled engineers job-hop on 12 to 18 month cycles, and stars go stale fast.
What to send when you reach out
The message that works acknowledges the signal without being creepy about it. You do not open with "I saw you starred a repo." You open with the category fit the star implied.
Concrete template shape:
- Line 1: Specific reference tied to their public profile (a repo they own, a talk they gave, a company they work for).
- Line 2: A precise role description that matches the category they signaled interest in (agent infra, voice AI, retrieval, etc.).
- Line 3: Comp band and location, up front.
- Line 4: One question that's easy to answer in a single line.
Refolk surfaces the profile detail that makes line 1 possible, which is usually the block between "I know the intent exists" and "I have something specific to say." Sourcing AI engineers at scale is mostly an identity-resolution problem, not an intent problem, once you have a repo like this one in play.
Why this window closes fast
The awesome-list format has a predictable arc: obscure, then trending, then noisy. Right now, awesome-ai-startups-hiring is in the middle phase, gaining stars, and not yet on every sourcing team's radar. That is the sweet spot.
Two things will end this window:
- Sales tools index it. GitLeads and its peers will start pulling this repo's stargazers for outbound sales into the 160 startups themselves. Once that traffic starts, some of the stars become noise from vendors, not candidates.
- The maintainer adds a "please don't scrape stargazers" note. Shahdeo has been generous so far. That could change.
Until then, the play is straightforward: pull the stargazer list, resolve identities, filter by title and geography, and reach out with a message that respects the signal.
FAQ
Is scraping GitHub stargazers allowed?
The stargazer endpoint is part of GitHub's public REST API and returns paginated data for any public repo. Using it within GitHub's rate limits and terms of service is standard practice. What you cannot do is republish stargazer lists as a product or use them to spam. Individual recruiter outreach based on a public star is squarely inside normal sourcing practice, the same way a LinkedIn "open to work" banner is.
How is a star different from a LinkedIn "open to work" signal?
Both are public intent signals, but the star is category-specific and the LinkedIn banner is not. A LinkedIn "open to work" says "I'll consider offers." A star on awesome-ai-startups-hiring says "I want to work at a funded, AI-native startup in one of these seven categories." The star also does not blow up the candidate's relationship with their current employer, which is why many senior engineers use it instead of the LinkedIn banner.
Can Refolk actually filter by "starred this specific repo"?
Refolk's model is that you describe the person in plain English, including GitHub activity signals like starring, forking, or watching a specific repo, and get a ranked shortlist with LinkedIn, current title, and employer resolved. If the identity is publicly resolvable across GitHub and LinkedIn, it shows up. The value is not the scrape itself; it is the identity resolution and the filter stack (title, seniority, geography, employer stage) applied on top of the intent signal.
What if the repo goes viral and the stargazer list becomes noise?
Shift from stars to watchers, and shift from all-time to last-30-days. Watchers subscribe to future updates, which is a much higher intent bar than a one-time bookmark. And if the repo hits 20,000+ stars, the recency filter (stars in the last 30 to 60 days) becomes the sharpest cut. The signal degrades gracefully as long as you keep tightening the window.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.