Wispr's $280M Series B: 14 ASR Engineers Have the Exact Title
Wispr's $280M Series B squeezed a tiny US ASR pool. Refolk's index shows 14 exact-title matches vs 5,093 skill-based. How to source both.
On Aug 17, 2026, Wispr closed a $280M Series B at a $2B valuation led by Menlo Ventures, roughly tripling from November's $700M mark. If you run talent at Deepgram, Otter, Willow, Monologue, Aqua, or Superwhisper, the next quarter of your sourcing plan just got harder, because the strict-title US ASR bench is measured in dozens and every rival is now fishing the same shortlist.
Wispr's new Canto speech model claims to cut word error rates from over 30% to under 10% in noisy conditions, and the company is hiring against that roadmap across a research lab led by ex-Alexa founding member Ariya Rastrow. This piece walks through what the actual US pool looks like in Refolk's index, why title-based sourcing misses roughly 99.7% of it, and where the under-priced feeders sit.
The strict US "ASR engineer" pool is 14 people
If you filter US professionals by the strict titles "Speech Scientist," "Speech Recognition Engineer," "ASR Engineer," or "Lead IVR Engineer," you get 14 results. That is the entire on-the-nose pool most recruiters actually search for, and it is what Wispr's Series B is now fighting over.
In Refolk's index, those 14 people cluster at legacy voice employers, not modern generative-AI shops:
- Dialpad (2 of 14, roughly 14% of the pool)
- Apple
- Uber
- Cerence AI
- Microsoft
- Sorenson / CaptionCall
Two things jump out. First, the largest single employer of exact-title ASR talent is Dialpad, a business-communications company, not a foundation-model lab. Second, roughly 29% of the narrow pool sits in Sunnyvale and Seattle, mirroring where the Alexa, Siri, and Google Speech benches have historically lived. If you are running voice AI hiring out of San Francisco and only sourcing locally, you are missing more than half of the exact-title supply before you start.
The skill-based pool is 5,093, a 364x expansion
Broaden the same query to any US professional listing "Speech Recognition" as a skill and the pool jumps to 5,093 people, a roughly 364x expansion over the strict-title cut. That is the entire arbitrage in ASR engineer recruiting right now.
The mechanism is title inflation. During the LLM wave, speech people rebranded. A production ASR engineer at a big lab is more likely to carry "Applied Scientist," "Foundation Model Engineer," "Speech ML Engineer," or "Audio ML Engineer" on their profile than "Speech Recognition Engineer." Recruiters who filter by title miss roughly 99.7% of the actual pool, and the miss is systematic: it skews toward exactly the modern generative-speech people Wispr's Canto roadmap needs.
The skill-based pool concentrates in five metros in Refolk's index:
- San Francisco
- Seattle
- Sunnyvale
- Boston
- Los Angeles
Those are the same cities where Alexa, Siri, Google Speech, and Meta Reality Labs teams sit, which is not a coincidence. Production speech expertise still lives inside the incumbents, and it leaves through referral networks, not job boards.
Why title-based sourcing fails for Wispr Flow hiring
Title-based sourcing fails because "Speech Scientist" was an early-2020s convention that the LLM era erased. Search on skills, on prior employers, and on publication history, not on current title.
Consider Wispr's own chief scientist, Ariya Rastrow. He was a founding member of Amazon's Alexa team and previously led multimodal foundation-model work at Meta. On a strict title filter for "Speech Recognition Engineer," he would not appear. He is the archetype of the hire every voice-AI company wants, and he is invisible to the way most recruiters query.
This is the exact gap Refolk closes. You describe the person in plain English ("US-based engineer who shipped production streaming ASR at Alexa, Siri, or Meta, now at a non-speech role") and get a ranked shortlist that pulls from publication history and profile signals across the open web, not from a title dropdown.
| Cohort | US count | What it tells you |
|---|---|---|
| Strict title: Speech Scientist / ASR Engineer / Speech Recognition Engineer | 14 | The pool title-based sourcing sees |
| Skill: "Speech Recognition" (any title) | 5,093 | The real addressable pool |
| Ratio (broad / narrow) | ~364x | Cost of title-only filtering |
| Narrow pool at Dialpad | 2 of 14 (~14%) | One competitor holds the largest slice |
| Narrow pool in Sunnyvale + Seattle | 4 of 14 (~29%) | Regional over-index vs SF |
| Deepgram public developer footprint | 200,000 devs / 1,400 orgs | Demand-side proxy for the category |
Canto's WER claim is a hiring spec in disguise
Canto's pitch to cut word error rates from over 30% to roughly 5% to 10% in wind, background noise, and strong accents is a hiring specification, not just marketing. Getting there requires three specific sub-specialties that most "ML engineer" candidates cannot fake.
- Noise-robust training. Data augmentation pipelines, simulated room impulse responses, multi-condition training. This lives at Alexa Speech, Google Speech, and Apple's Siri org.
- Streaming inference. Low-latency, chunked encoder-decoder architectures, streaming transformer variants. Deepgram is a concentrated feeder.
- Accent and dialect modeling. Multilingual acoustic modeling, code-switch handling, sub-word tokenization for low-resource variants. Meta AI, Google, and Cerence AI have the deepest benches.
The intersection of those three, in the US, is a subset of the 5,093 skill-based pool, probably in the low three digits once you filter for production experience. That is the real Whisper competitor talent pool everyone is fighting over, and it is why the Menlo valuation math depends on Wispr closing a specific slice of hires within the next two quarters, not a generic ML org build-out.
Recruiters filtering by title miss 99.7% of the actual speech recognition engineer pool.
The legacy voice feeder nobody is mining
The most under-priced source of production ASR talent is not OpenAI or Anthropic. It is the legacy voice stack: Dialpad, Cerence AI, Sorenson/CaptionCall, and the call-center analytics vendors. These are the top employers of strict-title ASR engineers in Refolk's index, and almost nobody in the current wave of voice AI hiring is prospecting them systematically.
The mechanism is boring and durable. Call centers and IVR platforms have been shipping production speech recognition for years. Their engineers have handled the un-glamorous problems Canto now claims to solve: noisy phone lines, regional accents, real-time streaming, cost-per-minute economics. A senior engineer from Cerence who has spent years on in-car voice under tight latency budgets is closer to what Wispr needs than a fresh PhD who has only fine-tuned Whisper on clean audio.
The second under-priced feeder is the Amazon Alexa alumni network. Rastrow's move to Wispr is a leading indicator. Senior Alexa speech people have landed at generalist ML roles where they now show up under titles like "Principal Applied Scientist" or "Staff ML Engineer." Query that population by skill and prior employer, not title, and you surface a bench that could walk into a Canto-scale problem tomorrow.
Who else is fishing the same pond
The direct competitors squeezing this pool alongside Wispr are Willow, Monologue, Aqua, Superwhisper, Deepgram, and Otter, per TechCrunch's Aug 17 coverage. Each has a different countermove that will shape comp and offer velocity over the next quarter.
- Deepgram. Crossed $100M ARR, serves 200,000 developers across 1,400 organizations, and has processed over 50,000 years of audio and transcribed more than 1 trillion words. The incumbent. Expect retention grants for anyone with tenure on the core ASR team.
- Otter. Long-tenured team, historically less aggressive on foundation-model hires. Likely to counter on stability and equity refresh rather than headline comp.
- Willow, Monologue, Aqua, Superwhisper. All dictation-adjacent, all earlier stage. Their play is speed: get to offer within 10 business days of first contact, before Wispr's brand pull kicks in.
- Menlo portfolio pull. Menlo led the round and backs adjacent AI infra companies. Expect portfolio-wide talent sharing that routes candidates toward Wispr first.
If you are running speech recognition engineer sourcing for one of the smaller players, your best defense is a mapped list of the 5,093-person skill pool, segmented by whether each person's current employer is a likely counter-offer risk. Ask Refolk for "US speech ML engineers currently at Alexa, Siri, or Google, who joined before 2022," and get a shortlist that reflects the actual poachability signal, not the title.
FAQ
How many US engineers actually match the exact "ASR engineer" title Wispr is hiring for?
Fourteen, in Refolk's index, across the strict titles "Speech Scientist," "Speech Recognition Engineer," "ASR Engineer," and "Lead IVR Engineer." Two of the fourteen sit at Dialpad, with the rest spread across Apple, Uber, Cerence AI, Microsoft, and Sorenson/CaptionCall. That is the pool every named Wispr competitor is now fighting over on a title-based search.
Why is the skill-based pool so much larger than the title-based pool?
Title inflation during the LLM wave. Speech engineers rebranded as "Applied Scientist," "Foundation Model Engineer," or "Audio ML Engineer" to align with how their employers pitched them internally. The underlying skill did not go anywhere, but the label did. 5,093 US professionals still list "Speech Recognition" as a skill even though only 14 carry a strict ASR title.
Which employers should I actually prospect for production ASR talent?
Start with the legacy voice stack: Dialpad, Cerence AI, Sorenson/CaptionCall, and the call-center analytics vendors. Then layer the Alexa alumni network, Apple Siri, and Google Speech. Deepgram is the concentrated modern feeder for streaming inference specifically. Skip generic "AI startup" prospecting for this role; the density is not there.
What does Wispr's Canto WER claim tell me about the hiring spec?
That the roadmap needs three specific sub-specialties: noise-robust training, streaming inference, and accent/dialect modeling. A candidate who has only fine-tuned Whisper on clean audio will not move Canto's numbers. Filter the 5,093-person skill pool for production experience in at least two of those three areas and the real target list is in the low three digits.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.