# From Conference Agenda to a Ranked Candidate Shortlist

*You can take one published speaker agenda and produce a deduplicated, seniority-inferred, ranked shortlist with a resolved contact path per person in one session.*

- Canonical URL: https://www.refolk.ai/guides/conference-agenda-to-shortlist
- Pillar: Recruiting and sourcing
- Format: Playbook
- Published: 2026-10-04
- Last reviewed: 2026-10-04
- Reading time: 16 min

A conference agenda is a pre-qualified candidate list that nobody has worked yet. Someone chose each speaker, each speaker chose a topic deep enough to stand in front of a room, and every name comes with a public bio and a reachable social handle. This guide is for in-house recruiters, sourcers, and founders hiring their own teams, and it carries one published speaker agenda all the way to a deduplicated, seniority-inferred, ranked shortlist with a resolved contact path per person, inside a single working session.

Most sources list conference speaker sourcing as one tactic and stop at "speaker lists are published online." That is the easy 20%. The hard parts are reading seniority out of a talk abstract, deduplicating one speaker across several events, and landing outreach inside the post-event warm window before the reference goes cold. This playbook handles those three.

## Why a conference agenda beats a cold list

A speaker agenda is a filtered, self-selected pool with a built-in personalization hook. Each name was vetted by a program committee, carries a public topic you can quote, and usually links to at least one social profile. That combination is what makes event sourcing compound.

The reason to prefer this over a scraped title search is reply rate. On the same 165,000-plus candidates contacted on both channels, 16.6% replied via LinkedIn versus 4.4% via email, and a personalized message outperforms a generic one by roughly 3 to 4x. A speaker reference is exactly that personalization lever. You are not inventing a reason to reach out; the person stood on a stage and told you what they care about.

**16.6% - LinkedIn reply rate on candidates contacted on both channels**

The same people replied at 4.4% by email, so channel choice alone is worth roughly 4x before you add a speaker reference.

Scarcity decides whether one agenda matters. In Refolk's index, Senior-band US people with Kubernetes skill number 50,654, while Senior-band Rust holders number 1,266 - a 40x gap. For a common skill, a single agenda is noise against the pool. For a rare one, a single niche conference can represent a meaningful fraction of everyone reachable. Check the size of your pool first; it tells you whether to work the list exhaustively or rank hard.

## What a speaker directory reliably gives you

Speaker directories reliably expose a name, a session title and description, a tagline, and a bio, plus optional social links. They do not reliably expose a structured job title or current employer, which is the field recruiters most want.

On the most common CFP platform, the mandatory, always-present fields are session title and description, speaker name, email address, tagline, and biography. Public profiles additionally expose location, photo, LinkedIn and Twitter/X URLs, company and personal website, and topic tags. Two caveats matter and both bite later:

- **Employer is not guaranteed.** Current job title and employer live inside the free-text bio or tagline, not a structured column. Parsing them is a step, not an assumption.
- **Social links are optional.** LinkedIn and Twitter/X URLs are captured only when speakers list them. Absence means missing data, not a dead end.
- **Email is held, not published.** The system stores a speaker email, but the public directory does not expose it.

The platform scale is large enough that most agendas you want are indexed: it hosts 313,000 speakers and 86,000 public speaker profiles. For events that have ended or moved, the Wayback Machine is the fallback. It has saved over 916 billion web pages, stores them indefinitely, and runs a 3 to 10 hour lag between a crawl and appearance. Archived pages carry the employer at event time, which is a historical signal you must re-verify.

#### What you can extract from a speaker record

1. **Always present** - Session title, description, speaker name, tagline, bio
2. **Usually present** - Location, photo, topic tags, personal or company website
3. **Optional** - LinkedIn and Twitter/X URLs, only when the speaker listed them
4. **Buried or absent** - Current employer and job title, hidden in free-text bio

*The reliable layers sit on top; the fields you most want sit underneath and need parsing.*

## The end-to-end procedure

Run these seven steps in order. The whole pass fits in one working session; the time estimates assume one agenda of a few dozen speakers.

#### Agenda to ranked shortlist, one session

1. **Pull the agenda** - Export speaker and session rows from the event feed or scrape the program page; if the site is gone, retrieve the archived page via the Wayback Machine. Done when you have one row per talk with name, title, affiliation or tagline, bio, and links.
2. **Normalize fields** - Split combined names, standardize affiliation strings, and flatten social URLs into consistent columns. Done when every row has the same clean schema and no combined cells.
3. **Deduplicate within and across events** - Block on email domain or company plus surname, fuzzy-match names with Levenshtein distance, and merge above-threshold records into clusters with an aliases field. Done when there is one record per real person, flagged if they appear at more than one event.
4. **Infer seniority** - Read the CFP audience-level tag where present; otherwise apply title and abstract heuristics where 101 signals teaching and a deep-dive signals senior hands-on depth. Done when each record is tagged practitioner versus manager with a seniority band.
5. **Rank and qualify** - Score each record on role fit, inferred seniority, repeat-speaker status, and contactability. Done when you have an ordered shortlist with a one-line reason per rank.
6. **Resolve a contact path** - For each person pick the best channel, starting with LinkedIn, then work email or GitHub. Done when every record carries one resolved channel.
7. **Sequence outreach to the warm window** - Send the first speaker-referencing touch inside 48 hours of the event when possible, then plan three touches over about 30 days. Done when messages are drafted and scheduled.

The one timing constraint that overrides everything: pull the agenda before or during the event, not weeks after. The peak outreach lift is the first 48 to 72 hours, but deduplication and enrichment eat a full session. If you start the week after the event closes, the window is already half gone.

> **Watch out:** The warm window and your timeline collide
>
> A six-month-old agenda with a "great to see you at X" note is a false warm signal. After a month you have missed the moment, and the reference stops lifting. Pull and process the agenda while the event is still live so your first touch lands inside 48 hours.

## Deduplicating a speaker across events

Deduplication is entity resolution: matching records on attributes, then merging above-threshold records into clusters while recording an aliases field with name variations. The discipline is what stops two different people collapsing into one record and one person splitting into three.

Name alone is never enough. A match on only the name is not dispositive because many names are shared; resolution depends on using information beyond the name. The practical tooling uses blocking to stay cheap: comparing every pair of 50,000 records is 1.25 billion comparisons, so records are grouped by cheap keys - email domain, company prefix, surname Soundex, a name initial - and only compared within a group. Quality depends on having at least one blockable field. For westernized names, edit distances such as Levenshtein handle deletions, insertions, and substitutions.

Here is the discipline in order:

1. **Pick a block key.** Email domain or company prefix plus surname. If a record has neither, it cannot be safely merged; park it rather than guess.
2. **Fuzzy-match names within the block.** Levenshtein distance catches "Jon" versus "John" and transposed middle initials without matching unrelated people.
3. **Require a second agreeing field before merging.** Same name plus same company domain is a merge; same name alone is not.
4. **Record the aliases.** Keep every name variation and every event in one cluster.

That last point pays a bonus. The merge step surfaces people invited to speak at more than one event - a repeat-speaker flag you get for free once clustering is done. Being asked back repeatedly is a credibility proxy, so promote repeat speakers in your ranking.

> **Rule:** Two fields or no merge
>
> Never merge two speaker records on name alone. Require agreement on a second blockable field - email domain, company, or a verified social URL - before you collapse them into one person. Name-only merges produce false positives that poison the whole shortlist.

I ran this search: `Developer advocates in Germany who spoke at a Kubernetes or cloud-native conference in the last year, with a LinkedIn profile` - [see the full result list](https://www.refolk.ai/s/rrdt8s82f8).

*Returns a deduplicated, contact-resolved set of speaker-candidates in one pass, so you skip the manual export, dedup, and link-finding steps.*

Deduplication and link resolution are the two steps that burn the most minutes by hand. When the job is "speakers at event X who match role Y and have a reachable profile," [Refolk](/) runs the match, the merge, and the contact resolution together, which is the difference between a same-day shortlist and a two-day one.

## Reading seniority out of a talk

Seniority in a conference program lives in two places: an explicit audience-level tag, where the platform forces one, and the language of the title and abstract, where it does not. Read both, and trust the abstract over the tag.

Most CFP systems force an audience level with options for Beginner, Intermediate, Advanced, and All. A common numeric scale maps to years: 100 is novice under one year, 200 intermediate at one to three years, 300 advanced at four to six, 400 expert at six-plus. Session tags can also name the audience type directly, listing technical practitioners such as front-end, back-end, platform architects, DevOps managers, and data scientists.

When the tag is missing, the title pattern carries the signal. A "101" or "What is X" talk teaches novices. A deep-dive title like "A Deep Dive Into Garbage Collection and the GIL" signals senior hands-on depth. Words like "leading" or "managing" in the abstract point to a people manager rather than an individual contributor.

> **Watch out:** Self-assigned levels lie in both directions
>
> A "300/advanced" tag can mean a manager's overview, not hands-on depth - some conferences define level 100 as a management-level overview. Grade against the abstract verbs and the speaker's current title, never the tag in isolation. Mis-targeting the audience level is the single most common talk-rejection reason, which tells you how loosely these tags are applied.

The useful output of this step is two tags per person: practitioner versus manager, and a seniority band. Those two tags drive the ranking.

## Ranking, and how much ranking you actually need

Rank on four inputs: role fit, inferred seniority, repeat-speaker status, and contactability. How hard you rank depends on pool size - a small market you work exhaustively, a large one you rank and cut.

Geography concentrates speakers far more sharply than raw headcount suggests. In Refolk's index, Developer Advocates run 8.8x higher in the US than in Germany (327 versus 37), and the German pool clusters in Berlin at a handful of employers. A German-market shortlist is small enough to work top to bottom; the US one needs real ranking to be usable.

| Market | Matching people | Top employer signal |
|---|---|---|
| United States | 327 | Stripe, Google, Roblox |
| Germany | 37 | JetBrains, SAP |
| US:Germany ratio (derived) | 8.8x | - |

The same logic applies to skills. The scarcer the skill, the more one agenda is worth, and the less you need to rank because the whole list is precious.

| Skill | Senior-band people (US) | Share of Kubernetes pool (derived) |
|---|---|---|
| Kubernetes | 50,654 | 100% (baseline) |
| Rust | 1,266 | 2.5% |
| Kubernetes:Rust multiple (derived) | 40.0x | - |

> For a rare skill, a single niche conference agenda can be a meaningful slice of everyone reachable.

#### How hard to rank, by pool and fit

Horizontal axis runs from Large reachable pool to Small reachable pool. Vertical axis runs from Weak role fit to Strong role fit.

| Quadrant | What it means |
| --- | --- |
| Rank hard, cut deep | Common skill, loose fit - score aggressively and keep only the top band |
| Work exhaustively | Rare skill, loose fit - small list, qualify every name by hand |
| Rank, keep the top slice | Common skill, strong fit - rank and take the best contactable rows |
| Contact everyone | Rare skill, strong fit - the whole agenda is your shortlist |

*Pool scarcity and role fit decide whether you work the list exhaustively or cut it hard.*

## Resolving a contact path and sequencing the touch

Pick one channel per person, defaulting to LinkedIn because it wins on reply rate, and send the first speaker-referencing touch inside 48 hours of the event. Keep the note short, both to lift replies and to protect your account.

The channel data is decisive. Across 4 million-plus messages, LinkedIn messages replied at 17.08% versus 6.31% for recruiter-written email and 4.96% for automated email, and all-industry cold email sits at 3.43%.

| Channel | Reply rate |
|---|---|
| LinkedIn message | 17.08% |
| Recruiter-written email | 6.31% |
| Automated recruiter email | 4.96% |
| All-industry cold email | 3.43% |

There is a live disagreement on order. Email reaches a yes faster; LinkedIn earns more replies. One reasonable sequence: LinkedIn connection first within 24 hours, work email as the second touch. Reverse it if speed-to-yes matters more than reply volume for your role.

#### Reply yield by channel per 100 recruiting messages

| Stage | Figure | Note |
| --- | --- | --- |
| LinkedIn message | 17 | ~17 replies per 100 |
| Recruiter-written email | 6 | ~6 replies per 100 |
| Automated recruiter email | 5 | ~5 replies per 100 |
| All-industry cold email | 3 | ~3 replies per 100 |

*Expected replies fall roughly fourfold from LinkedIn to all-industry cold email.*

Brevity is both a reply lever and a compliance tool. InMails under 400 characters earn about 22% higher response, while ones over 1,200 characters run 11% below average. Short notes also keep you above LinkedIn Recruiter's 13% response floor over 14 days; fall below it and you can trigger an improvement period that restricts sending. A personalized speaker reference plus a tight message is how you clear both bars at once.

**First touch, LinkedIn, inside 48 hours**

```
Hi {first name} - caught your {talk title} session at {event}. The point about {one concrete detail} stuck with me. I run hiring for {role} at {company} and your work maps closely. Open to a short chat this week?
```

*Swap the talk title and the one concrete detail. Keep the whole thing under 400 characters.*

On timing, the defensible claim is a 48 to 72 hour peak with a usable-but-decaying window out to about two weeks. One guide cites HBR that new-contact conversion probability drops 50% after 48 hours without a touchpoint; another reports a 50% drop beyond 72 hours. The exact decay curve for speaker-reference outreach, as opposed to face-to-face, is not publicly established, so treat the numbers as direction, not precision. Plan three touches over about 30 days, and after a month drop the event reference entirely.

## How this goes wrong

The failure modes below are where a clean-looking shortlist turns out wrong. Each one has a concrete check.

- **Self-assigned level is unreliable.** A 300/advanced tag may mean a management overview, not hands-on depth, because some conferences define level 100 as a management-level overview. Check the abstract verbs and the speaker's current title, not the tag.
- **Name-only dedup creates false merges.** Two "Michael Chen" speakers collapse into one person. Require a second blockable field such as company or email domain before merging.
- **Missing employer field.** Affiliation often lives in free-text bio, so "qualified by company" fails silently when you assume a company column exists. Parse the tagline and bio.
- **Stale affiliation on archived pages.** A Wayback program shows the employer at event time, not now. Treat it as a historical signal and re-verify the current role.
- **Warm window already closed.** Sending a "great to see you at X" note on a six-month-old agenda is a false warm signal. Past a month, drop the event reference.
- **Templated LinkedIn blast triggers the floor.** Bulk generic InMails depress reply rate; falling below the 13% floor over 100 messages in 14 days can trigger a sending restriction. Personalize each.
- **Optional social links read as no contact.** LinkedIn and Twitter/X URLs appear only when speakers list them, so absence is missing data. Search GitHub and the open web before discarding anyone.

> **Tip:** Turn the dedup step into a free signal
>
> The aliases field you build to merge duplicates also tells you who was invited to multiple events. Promote repeat speakers in your ranking - repeated invitations are a credibility proxy you got for nothing.

## Before you call it done

Run this checklist against your finished shortlist. If any item fails, the list is not ready to work.

#### Shortlist readiness

- [ ] Every row traces to a specific talk at a named event, with the session title captured.
- [ ] No two records share a name without also agreeing on a second field; aliases are recorded for every merge.
- [ ] Each record carries a practitioner-versus-manager tag and a seniority band, graded against the abstract and not the CFP tag alone.
- [ ] Employer is parsed from the bio or tagline where no structured field existed, and archived affiliations are re-verified as current.
- [ ] Each person has exactly one resolved contact channel, with LinkedIn preferred and GitHub or the open web checked where social links were absent.
- [ ] The shortlist is ordered, each rank has a one-line reason, and repeat speakers are promoted.
- [ ] First-touch messages are drafted under 400 characters, reference the talk, and are scheduled to land inside 48 hours of the event.

## Keeping the method current

The mechanics drift, so re-check three things rather than trusting a fixed number. First, platform credit pricing moves - additional InMail credits rose from $3 to $21 in late 2025 - so confirm the current cost before committing to a LinkedIn-heavy sequence. Second, the response-floor rule and the 30-day reply-counting window are platform policy that can change; verify the current threshold inside the tool before you blast. Third, the warm-window decay figures are loosely sourced practitioner claims, not controlled studies on speaker-reference outreach, so treat the 48 to 72 hour peak as a planning assumption and watch your own reply rates by days-since-event to calibrate.

The one thing that will not drift is the shape of the method: pull before the event ends, dedup with a second field, read seniority from the abstract, rank by pool scarcity, and lead with a short LinkedIn note that quotes the talk. Run it once on a single agenda and you will have a ranked, contact-resolved shortlist to show for the session.

## Frequently asked questions

### How do I deduplicate a speaker who appears at several conferences?

Never merge on name alone; a name match is not dispositive because many names are shared. Block records by a cheap key like email domain, company prefix, or surname, then fuzzy-match names within each block using edit distance. Merge only above-threshold pairs that also agree on a second blockable field, and keep an aliases field holding the name variations. A person who survives merging across two or more events gets a repeat-speaker flag, which is a free credibility signal.

### Can I read seniority from a talk title or abstract?

Partly. Many CFP systems force an audience-level field where 100 means novice under a year, 200 intermediate, 300 advanced, and 400 expert. A 101 or what-is-X talk teaches novices; a deep-dive abstract signals senior hands-on depth. But the tag is self-assigned and inconsistent: some conferences define level 100 as a management-level overview. Check the abstract verbs and the speaker's current title against the tag before you trust it.

### How late is too late to use the event reference?

The peak lift is roughly the first 48 to 72 hours; one guide cites HBR that new-contact conversion probability drops 50% after 48 hours without a touchpoint, and another reports a 50% drop beyond 72 hours. The window is usable but decaying out to about two weeks. After a month you have missed the moment and the reference stops lifting, so drop it or use it only as quiet context.

### Why prefer LinkedIn over email for speaker outreach?

On the same 165,000-plus candidates, 16.6% replied via LinkedIn versus 4.4% via email, so LinkedIn wins by roughly 4x on identical people. A speaker reference is exactly the personalization lever that compounds with LinkedIn. Keep notes under 400 characters, which earn about 22% higher response, and personalize each one to stay above LinkedIn Recruiter's 13% response floor.

### Where do I find agendas for events that have already ended?

Current and past public event feeds are often reachable by event ID with no login or API key. For pages that changed or went offline, the Wayback Machine is the fallback; it has saved over 916 billion pages, stores them indefinitely, with a 3 to 10 hour lag between a crawl and appearance. Treat the affiliation on an archived page as the speaker's employer at event time, not now, and re-verify the current role.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/conference-agenda-to-shortlist*
