Refolk
PlaybookEngineering and open source

The Developer-Advocate Sourcing Playbook: Teaching to Shortlist

You will take an open developer-advocate req to a ranked shortlist of 15 to 25 named candidates, each scored on credibility and communication from their own artifacts.

16 min readLast reviewed October 9, 2026Read as Markdown

Filling a developer-advocate seat is not a recruiting problem, it is a reading problem. The engineers you want already teach in public, ship sample code, and answer strangers' questions, but almost none of them carry the title. This playbook is for engineering managers, technical founders, developer-relations leads, and technical sourcers who need to go from an open req to a ranked shortlist of 15 to 25 named candidates, each scored on demonstrated technical credibility and public communication read from their own artifacts rather than their follower count.

The published advice stops at "look in your community" and "don't trust follower count." That is correct and useless. What follows is the stage-by-stage method: how to define the archetype, build a two-axis rubric, source from artifacts instead of applicants, and resolve a messy longlist into a ranked, evidenced shortlist you can defend to the hiring manager.

Why the titled pool is the wrong pool

The engineers who would make great advocates mostly do not hold the title, so sourcing by title throws away the field. The function is young - developer relations as a craft is not more than 15 to 20 years old, and "DevRel" only appeared in titles and conference names around 2015. That youth means supply still lives in adjacent engineering roles.

Refolk's index makes the gap concrete. In the United States, 363 profiles match Developer Advocate, DevRel Engineer, or Developer Evangelist. But 23,997 US engineers already carry both technical writing and public speaking as skills - the latent advocate pool. That is 66 times larger than the titled pool.

66.1x
US latent advocate pool vs titled pool, in Refolk's index
23,997 engineers with technical writing plus public speaking against 363 titled developer advocates.

This reframes the whole job. You are not competing for a scarce bench of titled advocates; you are identifying engineers who already do the work and have not been labeled for it. Amplify's advice captures the mechanism: a great advocate might never have been one.

SegmentProfilesMultiple vs titled
Titled DevRel (Advocate, DevRel Eng, Evangelist)3631.0x
Engineers with technical writing + public speaking23,99766.1x

Geography bends this further. The US titled pool is 10.1x Germany's 36 profiles - steeper than the roughly 4x company-count gap reported in 2020, where 56% of companies practicing DevRel originated from the USA and Europe had almost four times fewer. Title adoption lags company presence outside the US, so a European search has to lean even harder on the latent pool.

CountryTitled profilesMultiple vs Germany
United States36310.1x
Germany361.0x

The practical consequence: if your pipeline tool filters on the DevRel title, it is showing you a hundredth of the field in the US and far less in Europe.

Which archetype you are actually hiring

Before you source anyone, decide which of two jobs you are filling, because they are different jobs and one person rarely does both well. Hiring usually boils down to either growing top of funnel or making the existing developer audience more successful - completely different goals, and it is unlikely you will find someone good at both.

The cleanest framing maps the hire to a four-stage funnel. DevRel program metrics align with the four objectives of awareness, activation, engagement, and retention. A content-first advocate owns the top of that funnel; a community-first advocate owns the bottom. The taxonomy extends further in practice - Developer Advocate or Evangelist, developer-experience practitioner, and Technical Community Manager - but for a single req, the content-versus-community split is the decision that matters.

The four-objective DevRel funnel

  1. Awareness
    wide

    top-of-funnel content, talks, tutorials

  2. Activation
    narrower

    first successful integration, quickstarts

  3. Engagement
    narrower still

    ongoing community, OSS answering

  4. Retention
    narrowest

    deep success, advocacy, NPS

Decide which stage your hire owns before you write the rubric, because it changes what counts as a strong artifact.

One more variable hides inside the req: the reporting line. Two advocates with the same title can have a significant comp gap because of where the role sits. In 2021 data, reporting lines split across marketing (26.2%), product (17.4%), engineering (15.9%), CEO (11.8%), and CTO (10.8%). Senior and head-of-DevRel bands run $170k to $260k, against $90k to $120k entry-level. Build the req around the line, not the title, because the line sets both comp and scope.

The two-axis rubric that separates credibility from communication

Score every candidate on two independent axes - technical credibility and public communication - because strength on one does not imply the other. Collapsing them into a single score is how you hire a polished writer who cannot read a stack trace, or a brilliant maintainer who cannot explain a concept to a junior.

Credibility is evidenced by engineering artifacts: documentation, client libraries, code labs, sample code, and video tutorials that a candidate built to help developers succeed with a product. Communication is evidenced by published content you can read or watch. The sourcing heuristic from the field is direct: pay attention to who wrote it when you read interesting technical writing or watch a quality video.

Credibility by communication

Strong credibilityWeak credibility
Strong maintainer, cannot explain
Probe with the explain-to-a-junior exercise before passing
Advocate hire
Advance to friction log and work-sample
Neither
Rule out
Polished but unproven
Open a PR trail and run their code before trusting
Weak communicationStrong communication
Plot every shortlisted candidate on both axes; only the top-right quadrant is a clean advocate hire.

Build the rubric as a scoring sheet where each candidate earns a credibility score and a communication score separately. Count user contributions whether pull requests, questions, answers, blog posts, or meetup talks. The quantifiable credibility signals the field names are Stack Overflow mentions, GitHub project stars, and community-driven responses to questions.

Each signal should also tell you how it lies. A criterion you cannot falsify is not a criterion.

SignalWhat it provesWhat it looks like when it lies
GitHub stars and forksShipped code others use500 stars on one template repo, no recent commits
Stack Overflow answersCan teach under pressureHigh rep from years ago, nothing in 12 months
Published tutorialCommunication abilityGhostwritten or marketing-authored; code does not run
Conference talkOn-camera communicationOne talk, no supporting artifacts or OSS trail
Follower countReach, not resultsLarge audience, zero shipped code

The last row is the one the market gets wrong. Follower count measures reach, not results, and it is the easiest metric on a profile to inflate. It is not even a top-four program metric: 2023 programs weight active users (45.1%) and content engagement (39.6%) far above developer NPS (22.2%) and site visits (15.3%).

MetricShare of programs
Active Users45.1%
Content Engagement39.6%
Developer NPS22.2%
Site Visits15.3%

If you must use an audience signal at all, use the ratio instead of the raw number: average views more than five times the following count is a valuable signal that actual reach exceeds following.

Source from artifacts, not from applicants

The best candidates have demonstrated skills publicly through blogs, open source, and talks; if they have not created anything, they are unproven - look for a portfolio, not just potential. So you source backwards: start from the artifact and find its author, rather than posting a req and waiting.

Three veins produce the longlist. First, your own community and power users - the people already filing good issues and writing about your product. Second, the open-source graph: search GitHub and the projects your company depends on to find and approach top contributors. Third, the public teaching surface: tutorial authors found via HackerNews, Stack Overflow answerers in your stack, and conference speakers.

Start from the artifact and find its author, not from the req and wait for one.

This is where the latent-pool math pays off, and where a plain-English search beats a title filter. You are describing a behavior - teaches in public, ships code in your stack, speaks on camera - not a job title.

Refolk is built for this exact move: ask for the behavior in plain English and get named people back across GitHub, LinkedIn, and the open web, instead of hand-assembling a list from a dozen tabs. Aim for a raw longlist of 60 to 100 names, each attached to one linked artifact, before you score anything.

The procedure, end to end

Run these eight stages in order. The first four produce the ranked shortlist; the last four convert it to a hire. Each stage names its owner, its rough duration, and what done looks like.

Developer-advocate sourcing procedure

  1. Define the archetype and funnel goal
    Decide content-versus-community and which funnel stage the hire owns. Done when a one-page req names the archetype, the target developer ICP, and a single North Star metric.
  2. Build the artifact rubric
    Separate credibility signals (OSS issue answering, code samples, docs and SDK work) from communication signals (blogs, talks, tutorials, video). Done when the scoring sheet gives each candidate independent credibility and communication scores.
  3. Source from artifacts, not applicants
    Work backwards from content and your own power users; search GitHub contributors and Stack Overflow answerers in your stack. Done when you have a raw longlist of 60 to 100 names with one linked artifact each.
  4. Score and rank the longlist
    Apply the rubric and compute the views-to-following ratio where available instead of raw follower count. Done when 15 to 25 candidates each carry two scores and at least two linked artifacts.
  5. Run the friction-log async screen
    Ask shortlisted candidates to produce a short friction log of your onboarding. Done when you have evidence of real product engagement and unassisted written communication.
  6. Run the technical and communication interview
    Probe API, SDK, and docs judgment plus an explain-this-concept exercise and business-model awareness. Done when you have a hire or no-hire with notes on funnel understanding.
  7. Run the paid work-sample day
    Give every finalist the same paid, representative task - quickstart, demo build, or bug triage. Done when you hold comparable deliverables scored on one rubric.
  8. Reference and offer
    Check references and confirm the reporting line (engineering vs marketing) because it drives comp. Done when the offer is signed and the reporting line is agreed in writing.

A note on pace: steps 1 and 2 are a few hours of work by the hiring manager and DevRel lead. Step 3 is the long pole at one to two weeks. Steps 4 through 8 run three to five days each, with the work-sample day being a single paid day scheduled around the candidate.

Score and rank, concretely

When you compute the two scores, resolve ties with recency and with authorship you can verify. A friction log is the cheapest authorship test you have.

Candidate scoring row
Name | Primary artifact (link) | Second artifact (link) | Credibility /5 | Communication /5 | Last activity (date) | Views:following ratio | Funnel stage fit | Notes

One row per candidate; keep the two scores separate and never average them into one number.

The async screen that exposes real engagement

A candidate-initiated work sample - the friction log - tells you two things at once: whether they actually used the product, and how they write when nobody edited it.

Friction-log async screen brief
Spend up to two hours onboarding to [product] as a new developer would.
Write a short friction log: where you got stuck, what confused you,
what you would change, and one thing that worked well.
Include any commands you ran and errors you hit.
Return as a Markdown doc or public gist.

Send to shortlisted candidates; cap at two to four hours of their time and review in about 30 minutes.

Interview to separate adoption-drivers from output-producers

A published list of questions that cleanly separate activation-ownership from retention-ownership is not established in the field, so you convert the funnel language into questions yourself. Ask a channel-triage question - how do you handle triaging and responding to developer questions across GitHub, Slack, and forums without burning out - to read engagement and retention instinct. Ask a business-literacy probe, because a zero-to-one developer advocate should always be able to explain to the community how the business makes money or intends to make money. And run an explain-this-concept exercise to the level of a junior engineer, which is the fastest way to catch strong credibility paired with weak communication.

The paid work-sample day

Work-samples beat portfolios because authorship is verifiable. PostHog publishes the clearest model: a paid full day of work, where the task is shared at the start of the day, is representative of the real role, and is always the same for each candidate so you can make clear comparisons. Use one of three documented prompts and score every finalist on it.

Prompt typeWhat candidates doScored on
Technical contentWrite a quickstart for integrating an API in a given language, with auth, error handling, troubleshootingClarity, correctness, security, completeness, developer empathy
Demo buildBuild a sample app demonstrating webhooks, retries, idempotency; present in 10 minutesCode quality, presentation, judgment
Bug triageDebug a real GitHub issue in a simulationDebugging approach, communication, escalation judgment

How this goes wrong

Most failed advocate hires trace to one of a small set of predictable errors. Each has a false positive and a check you can run before it costs you a hire.

  • Follower count masquerading as credibility. The tell is 100k followers from one viral post and zero shipped code. Check the views-to-following ratio and require two linked engineering artifacts before shortlisting.
  • Stale audience. A once-active blog or channel that went quiet still reads as influential. Check artifact dates and require activity in the last 12 months.
  • Content without credibility. Polished tutorials that were ghostwritten or marketing-authored. Open a PR or issue trail on GitHub and run a code sample; a sample that does not run is the tell.
  • Credibility without communication. A strong maintainer who cannot explain a concept. Catch it with the explain-to-a-junior exercise and the quickstart work-sample.
  • Archetype mismatch. Hiring a community builder for a top-of-funnel content goal, or the reverse. Re-read the req against the funnel stage before advancing anyone, because it is unlikely you will find someone good at both.
  • Unfair work-samples. Different tasks per candidate break comparability. Fix it with one fixed, paid task, as PostHog documents.
  • Reporting-line surprise. Comp and scope swing on engineering-versus-marketing placement; confirm the line before the offer.
  • Pool-size denial. Expecting a deep bench of titled advocates. The titled US pool is 363 profiles; plan to convert the 66x-larger latent pool instead.

One honest limit: the field does not publish a reasoned, DevRel-specific teardown of follower count as a hiring filter, and it does not publish a named interview-question set that separates activation-ownership from retention-ownership. Both are load-bearing gaps here. The anti-following case is borrowed from adjacent influencer-marketing literature, and the activation-versus-retention questions are converted from the metrics framework rather than quoted. If you want harder evidence than this guide can cite, run a small internal calibration: score five of your own best past advocates on this rubric and see whether the two axes actually predicted their performance.

Before you call the shortlist done

Run this checklist against the shortlist before you hand it to the hiring manager. It is the difference between a list of names and a defensible, evidenced pool.

Shortlist readiness

  • The req names the archetype (content vs community) and a single North Star funnel stage.
  • The reporting line (engineering vs marketing) is decided and reflected in the comp band.
  • Every candidate has an independent credibility score and communication score, not an average.
  • Every candidate has at least two linked engineering artifacts.
  • Every candidate shows activity within the last 12 months.
  • Where an audience number appears, the views-to-following ratio is recorded, not the raw follower count.
  • The shortlist holds 15 to 25 named candidates drawn mostly from the latent pool, not the titled pool.
  • A single fixed, paid work-sample task is chosen and ready for all finalists.

Keeping the pool current

The advocate pool is not static, so treat it as a standing search rather than a one-time pull. The addressable base is enormous - SlashData estimates more than 40 million active developers worldwide - and the titled slice keeps shifting as the function matures. Re-run your artifact searches on a cadence tied to your release calendar, because every new SDK, API surface, or docs push creates fresh tutorials and OSS activity you can read as capability evidence.

Two things to re-check rather than assume. First, the market's demand signal moves: a 2021 LinkedIn search returned more than 800 Developer Relations opportunities worldwide, and a later US snapshot showed 603 listings, of which 158 were entry-level and 337 mid-senior. Those numbers date quickly, so when you need a current read, re-run the search rather than quoting a figure. Second, comp bands drift with the reporting line; confirm the current band for the line you chose before every offer, not once per year.

The durable part of this playbook is the mechanism, not the counts: engineers who teach in public leave artifacts, artifacts prove credibility and communication separately, and a fixed paid work-sample verifies authorship that a portfolio hides. Build the pipeline around those three facts and the shortlist will hold up no matter how the title or the market shifts around it.

Questions practitioners ask

How do I source developer advocates when almost nobody holds the title?

Stop searching for the title and search for the behavior. In Refolk's index, the United States has 363 titled advocates but 23,997 engineers who already carry technical writing and public speaking, a 66x larger pool. Source from public teaching and shipped code: GitHub contributors in your stack, Stack Overflow answerers, tutorial authors, and conference speakers. Then score them on artifacts, not on whether their profile says DevRel.

Why shouldn't I filter developer relations candidates by follower count?

Follower count measures reach, not results, and it is the easiest metric on a profile to inflate. A single viral post can produce a large audience with zero shipped code behind it. If you must use an audience signal, use the ratio: average views more than five times following count indicates real reach. Better still, require two linked engineering artifacts before anyone reaches the shortlist.

How do I evaluate a devrel candidate's technical credibility?

Read their engineering artifacts directly. Open a pull-request or issue trail on GitHub, run a code sample, and check that a published quickstart actually works. Polished tutorials can be ghostwritten or marketing-authored, so a code sample that does not run is the tell. Credibility lives in the open-source trail and in a paid work-sample day, not in the prose.

What does a fair developer-advocate work-sample look like?

One fixed, paid, representative task given to every finalist so you can compare deliverables directly. PostHog publishes the clearest model: a paid full day of work with the same task for each candidate. Documented prompts include a quickstart for integrating an API with auth and error handling, a sample app demonstrating webhooks and idempotency presented in ten minutes, or a bug-triage simulation on a real GitHub issue.

How big is the developer-advocate candidate pool?

The titled pool is thin because the function is only about 15 to 20 years old. In Refolk's index, the United States holds 363 titled advocates and Germany holds 36. The addressable audience is huge, with more than 40 million active developers worldwide, but the titled fraction is tiny. Plan to convert the latent pool of engineers who already teach, which is roughly 66 times larger in the US.

Does the reporting line really affect what I should offer?

Yes. Two advocates with the same title can have a significant comp gap because of the reporting line. In 2021 reporting data, marketing accounted for 26.2% of lines and engineering 15.9%, and senior bands run $170k to $260k. Confirm whether the role sits under engineering or marketing before you make an offer, because it sets both comp logic and scope.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next