# Reconstructing a Company's Interview Loop From Public Reviews

*You will turn public interview reviews into a predicted round-by-round loop for your exact role, and know how far to trust each number.*

- Canonical URL: https://www.refolk.ai/candidates/guides/reconstruct-interview-loop-reviews
- Pillar: Reading the market
- Format: Teardown
- Published: 2026-08-17
- Last reviewed: 2026-08-17
- Reading time: 18 min

You have an interview scheduled and you want to know what this company's loop actually looks like for your exact role before you plan your prep. This guide is a procedure for turning scattered public interview-experience data into a predicted round-by-round loop: the number of stages, the format of each, the difficulty, the likely duration, and the recurring question themes. It carries one worked example all the way through, including the wrong reads and the cross-check against the recruiter's own schedule, so you can follow along on your own case.

Most "research the company" advice stops at mission, culture, and stalking your interviewers on LinkedIn. That is not the job here. The job is to reconstruct the loop shape, and to attach a trust level to every number so you know which ones to build prep on and which ones to ignore.

## What public interview pages actually report

Public company interview pages report four headline numbers computed from candidate self-report surveys: a difficulty score, a percent-positive-experience figure, an average days-to-hire, and a count of user-submitted interviews, plus an application-source breakdown. Those four numbers are your company-level card. They are the coarsest read you will get, and on their own they will mislead you, which is why the rest of this guide exists.

The difficulty score sits on a five-point scale: 1.0 is very easy, 3.0 is average, 5.0 is very difficult. The average interview difficulty rating across the platform is 2.8. Percent-positive is the share of respondents who called their experience positive. Average days-to-hire is a mean of self-reported timelines. The interview count, the N, is the single most important number on the page, because it tells you how much weight the other three can bear.

Here is a live company-level card so the shapes are concrete. Read these as the blended, all-titles figures they are.

| Company | % positive | Difficulty /5 | Avg days | N interviews |
|---|---|---|---|---|
| Scale | 38 | 3.03 | 23 | 416 |
| Glassdoor | 74.8 | 3.09 | 24 | 1,141 |
| EMPLOYERS | 46.4 | 2.85 | 14.62 | 28 |
| Cro Metrics | 36 | 3.1 | 13 | 50 |

Two things jump out. First, difficulty barely moves: everything clusters between 2.85 and 3.1, within a fifth of a point. That is not a coincidence, it is the shape of the whole distribution, and I will come back to why. Second, percent-positive swings hard, from 36 to 74.8, while days-to-hire ranges from 13 to 24. The positive-percent and the timeline carry most of the company-to-company signal. Difficulty carries almost none at this level.

One more source note. Not every site uses the same scale. Indeed reports on 1-to-10 scales for poor-to-excellent and easy-to-difficult, so never mix an Indeed number into a five-point comparison without converting it. Log which platform each figure came from.

> **Note:** The N is the number that matters
>
> Difficulty, positive-percent, and days-to-hire are only as trustworthy as the count sitting next to them. A card with N=28 and a card with N=1,141 are not the same kind of evidence, even when the ratings look similar.

## Why the company average is the wrong number for you

The company average blends every title, and your role can sit far off it. This is the single most common way this research goes wrong: quoting a company-wide 3.03 difficulty or 38% positive as if it described your loop. Interview feedback varies widely across roles even inside the same company, because different roles have different pipelines, different expectations, and different processes.

The evidence is direct. Within one large company, only 25.4% of Data Scientist candidates found their interviews difficult, but that figure was 36.4% for Senior Software Engineer. Positive experience swung the same way: 51% for Senior Software Engineer versus 69.5% for Data Engineer. Same employer, same era, and a 13 to 18 point spread depending on which door you walked through.

| Role | % finding difficult | % positive experience |
|---|---|---|
| Data Scientist | 25.4 | n/a |
| Senior Software Engineer | 36.4 | 51 |
| Data Engineer | n/a | 69.5 |

Notice that the Senior Software Engineer difficulty is roughly 43% higher, in relative terms, than the Data Scientist figure inside the same walls. Much of that gap is seniority, not company. Senior candidates are harder to please, they are often interviewers themselves, and they expect more from the process. So filtering the page to your level matters more than choosing a "hard" or "easy" company in the first place.

**51% vs 69.5% - Positive-experience share for two roles at one company**

The company-wide average hides an 18-point role spread, which is why step 3 filters to your exact role before you trust any number.

The practical rule: never quote a company number without the role-level N beside it. If the page will not filter to your role, or the role-level N collapses to single digits, you do not have a role-level read. You have a company average wearing your role's name tag.

## Reading difficulty and positivity without being fooled

Difficulty scores are downward-biased by construction, and positive-experience percent is partly an offer-outcome proxy rather than a measure of how hard the loop is. Read both against their mechanisms, not their face value.

Start with difficulty. Only 10.5% of interviews in the underlying study were classified as difficult, and only 1.2% as very difficult. That is why every card in the first table clusters near 3. The consequence is that a role reading 3.4 is not "slightly above average," it is genuinely toward the top of the real distribution. Calibrate accordingly: treat 3.0 as ordinary, 3.3 and up as a hard loop, and anything at 3.5 or beyond as a signal to prepare seriously.

Two forces push a difficulty number away from the truth. Rejection inflates it: rejected candidates rate harder, and the correlation between onsite-to-offer ratio and the share finding interviews difficult is minus 0.49. A page skewed toward rejected candidates reads scarier than the loop is. Format inflates it too: adding a group panel interview adds a statistically significant 13% to difficulty ratings. So if the reviews mention a panel round, part of the reported difficulty is format stress, not the intellectual bar.

> **Rule:** Read the outcome tag on every body
>
> Weight offers and rejections separately when you estimate difficulty. A cluster of rejected-candidate reviews will overstate the bar, and a cluster of offer-holders will understate it. Balance the two before you commit a number.

Now positivity. There is a 0.75 correlation between onsite-to-offer ratio and positive experiences, and roughly 70% of candidates with a positive experience got offers versus 20 to 30% of neutral and negative ones. Getting an offer makes people think fondly of the process. So a low positive-percent often signals a selective funnel, not a bad loop. Do not read 38% positive as "this company treats candidates badly"; read it as "this company says no to most people, and the people it said no to filled in the survey."

> A low positive-experience percent usually means a selective funnel, not a broken interview process.

## The procedure: from scattered reviews to a predicted loop

This is the core of the guide. Seven steps take you from a company name to a round-by-round table with a confidence flag on every row. Budget about two hours the first time. The work is reading and tagging, not guessing.

#### Reconstruct the loop for your role

1. **Fix the target** - Pin the exact company plus exact role title and level into one company-role-level string, because role variation inside a company is large.
2. **Pull the company-level card** - Record difficulty, percent positive, average days-to-hire, and the interview count from the company's interview page, with the date you pulled them.
3. **Drill to the role** - Filter the page to your role, note the hardest and easiest-role callouts and the per-role N, or flag that there are too few reviews for your role.
4. **Read 15 to 30 free-text bodies** - Tag each review by round sequence, format, duration, application source, and question themes, stopping when a new review stops changing the modal round list.
5. **Reconstruct the loop** - Collapse tags into an ordered stage list, keep only stages that recur across independent reports, and mark team-match or hiring-manager stages that appear only at senior levels.
6. **Cross-check the recruiter schedule and posting** - Compare your reconstruction to what the recruiter told you and the live posting; the schedule wins for the stage list, reviews win for themes and difficulty.
7. **Stamp trust levels** - Attach N and recency to each number and flag anything resting on fewer than roughly 10 to 15 role-level reviews as directional only.

#### The seven-step reconstruction

1. **Fix the target** - One company-role-level string
2. **Pull the card** - Four headline numbers plus N
3. **Drill to the role** - Role-level difficulty and positivity, or a too-few flag
4. **Read bodies** - Tag round sequence, format, duration, themes
5. **Reconstruct** - Ordered stage list, confidence flag per row
6. **Cross-check and stamp** - Reconcile against the recruiter, mark trust

*Each step narrows a scattered set of reviews toward a single predicted loop you can prep against.*

### The worked example, including the wrong turn

Take a Software Engineer loop at a payments company as the carried example, because it exposes every trap at once. Triangulated across guides, the onsite is five rounds: coding, a debugging round some candidates call Bug Bash, system design, an API integration round, and behavioral. The debugging round is distinctive to this company, which is exactly the kind of stage you only recover by reading bodies, not by reading the headline numbers.

Now the wrong turn, which is the point of step 6. One published guide states the loop "consists of 3 rounds." Others describe a two-screen-plus-five-onsite loop. Both cannot be your loop. If you had trusted the first aggregator you opened, you would have prepped for three rounds and walked into seven or eight. The fix is the rule in step 6: require a stage to appear in two independent sources before you commit it, and let the recruiter's schedule break the tie.

The loop also varies by level in a way the average never shows. New-grad loops are often leaner, typically coding, integration, and debugging without a full system design round, while senior loops may add an API design round or a final hiring-manager conversation. So mark those senior-only stages separately when you build the table. Your reconstruction is not "the company's loop," it is "the loop for my role at my level."

### Tagging themes, not just stages

Round count and format are not the only things you can extract. Recurring question themes are recoverable too, if enough bodies mention them. In the carried example, one guide noted that incremental problem design is a recurring format across coding rounds, with problems presented in multiple parts, each stage adding a new constraint, a pattern synthesized from 14 candidate reports. That is a prep instruction hiding in the reviews: practice building solutions that extend cleanly, because you will be asked to add constraints mid-problem. Tag themes wherever three or more bodies name the same one.

**Per-review tagging row**

```
Review date | Role + level | App source | Outcome (offer/reject/withdraw) | Round sequence (ordered) | Format per round | Total duration | Question themes | Notes
2 months ago | SWE, mid | applied online | offer | phone screen > coding > debugging > system design > behavioral | 1 recruiter, 4 technical/virtual | 3 weeks | incremental coding, real-codebase debugging | panel on system design
```

*Copy this into a sheet, one row per review body. The "modal" of each column is your reconstruction.*

Once ten or so rows are tagged, the modal round sequence is your predicted loop. The columns that keep repeating are high confidence; the ones that appear once are outliers you note but do not commit.

> **Tip:** Stop when the list stabilises
>
> You do not need every review. Stop reading when adding another body no longer changes the modal round list. For most roles that happens between 15 and 25 bodies. Reading 80 will not sharpen a loop that already converged.

## Cross-checking against the recruiter and the live posting

The recruiter-provided schedule and the current job posting are authoritative for your specific loop, and reviews are not. Reviews lag reality: they span years and roles, and reviews written during a period of significant change may reflect a temporary disruption rather than the real process. So when your reconstruction and the recruiter disagree, the schedule wins for the stage list.

That does not make the reviews useless after the cross-check. It reassigns their job. Use the reviews to fill what the recruiter leaves unstated. Recruiters routinely name the stages ("a coding screen, then an onsite of four") without telling you the format, the difficulty, or the themes. That is precisely where the review bodies earn their keep. The division of labour is clean.

| Question | Authoritative source | Why |
|---|---|---|
| Which stages, in what order | Recruiter schedule + posting | Reviews lag reality and span years |
| How many rounds total | Recruiter schedule | Confirmed for your loop, not inferred |
| Format of each round | Role-matched reviews | Recruiters rarely describe format in detail |
| Difficulty per round | Recent role-level reviews | Not stated by recruiters |
| Recurring question themes | Tagged review bodies | Only candidates report these |

There is no published protocol for resolving these conflicts, so this default is a judgement call, not a law. But it is defensible: it trusts the confirmed, role-specific, present-tense source for structure and the crowd for texture. Weight recent, role-matched reviews highest when you fill the gaps, and discard bodies that predate a posting that contradicts them.

To find the recruiter or scheduler you are cross-checking against, and to see who recently survived the current loop, a targeted people search saves an afternoon of manual digging. [Refolk](/candidates) writes your resume from your own history and tailors it per posting, and the same index lets you surface the people around your loop directly.

Ask me this: `Technical recruiters at Scale AI` - [run the search](https://www.refolk.ai/start?q=Technical%20recruiters%20at%20Scale%20AI).

*Returns the schedulers you would cross-check your reconstructed stage list against in step 6, so you can confirm the loop with the person who owns it.*

## How this goes wrong: the failure modes

This is the section to slow down on, because a confident-looking reconstruction built on biased data is worse than no reconstruction. Seven failure modes recur, each with a check that catches it.

- **The company average masquerading as your role.** The headline blends all titles, and your role can sit 10 or more points off, as the 51%-versus-69.5% gap showed. Check: never quote a company number without the role-level N beside it.
- **A thin role sample read as signal.** Two companies looked worst on the platform largely because they had the fewest reviews, a selection-bias artifact. A "3.8 difficulty" resting on four reviews is noise. Check: flag anything under roughly 10 to 15 role reviews as directional only.
- **Rejection-inflated difficulty.** Rejected candidates rate harder, correlation minus 0.49 with offer rate, so a page skewed toward rejects reads scarier than the loop is. Check: read the outcome tag on each body and weight offers and rejections separately.
- **Extreme-only bodies.** People with extreme experiences, good or bad, post more often than those with moderate ones. Check: reconstruct from the modal, most-repeated stage list, not the vivid outlier.
- **A stale loop.** Reviews span years, and stages like debugging rounds, take-homes, and AI screens get added or dropped. Check: sort by date and discard bodies older than roughly 18 months when a recent posting contradicts them.
- **Source-count conflict.** Guides disagree, as the three-rounds-versus-five-onsite split showed. Check: require a stage to appear in two independent sources before you commit it.
- **Company-gamed positivity.** Employers sometimes encourage positive reviews. Check: compare percent-positive against difficulty and days-to-hire; suspiciously high positivity next to a long, hard loop is a red flag.

> **Watch out:** The two failures that cost the most
>
> The thin-sample read and the stale loop are the ones that quietly wreck prep. A scary difficulty score built on four reviews sends you into over-prep; a four-round reconstruction from an old review sends you into a seven-round loop unprepared. Both are invisible unless you check N and date on every figure.

The staleness failure has a measurable trend behind it. Technical roles saw 42% more interviews per hire in 2024 than in 2021, rising from 14 to 20 total interviews per opening. The behavioral round alone now accounts for 30 to 40% of total interview time at major tech companies, up from 10 to 15% five years earlier. So an older review that reconstructs "four rounds, mostly technical" is not just old, it is directionally wrong for a current loop. Recency-weighting is not optional.

## Trust levels: which numbers to build prep on

Attach the N and the recency to every number, and mark each cell high, medium, or low confidence before you plan a single hour of prep. The goal is not certainty, it is knowing which figures can carry weight and which are placeholders.

#### What each number can bear

Horizontal axis runs from Few role reviews to Many role reviews. Vertical axis runs from Old reviews to Recent reviews.

| Quadrant | What it means |
| --- | --- |
| Old, thin | Discard, or hold as a vague prior only |
| Recent, thin | Directional; confirm against the recruiter before you act |
| Old, deep | Trust the stage shape, distrust counts and timing |
| Recent, deep | High confidence; build prep directly on this |

*Trust rises with both sample size and recency; the two axes decide whether a figure drives prep or just informs it.*

Use a simple rule of thumb for the flag. Any figure resting on fewer than roughly 10 to 15 role-level reviews is directional only, whatever it says. A number with 30-plus recent role-matched bodies behind it can drive your prep schedule. Everything between is medium: useful, but confirm it against the recruiter or the posting before you commit prep time to it.

One more calibration point on scale. In Refolk's index there are 337,990 US software engineers against 106,531 recruiters, roughly 3.2 engineers per recruiter, so any loop draws its interviewers from a deep bench. That is a reason to spend your effort reconstructing the loop's shape rather than guessing which individual you will face. The stage list is stable and knowable; the specific interviewer is not, and knowing them buys you little.

**3.2x - US software engineers per recruiter in Refolk's index**

The interviewer bench is deep enough that reconstructing the loop shape beats guessing who you will meet.

Before you call the reconstruction done, run this check.

#### Before you plan prep against this loop

- [ ] The target is a single company-role-level string, not just a company.
- [ ] Every headline number is logged with its N and the date pulled.
- [ ] You have a role-level read, or an explicit too-few-for-this-role flag.
- [ ] The loop is built from the modal stage list, not from one vivid outlier.
- [ ] Each committed stage appears in at least two independent sources.
- [ ] Senior-only stages are marked separately from the core loop.
- [ ] Every stage-list conflict with the recruiter resolves in the recruiter's favour.
- [ ] Each cell carries a high, medium, or low confidence flag.
- [ ] Any figure on fewer than about 10 to 15 role reviews is marked directional.
- [ ] Reviews older than roughly 18 months are dropped where a recent posting contradicts them.

## Keeping the read current as your loop moves

A reconstruction is a snapshot, and your loop is live, so re-run the cheap steps as new information arrives. After each round, compare what actually happened to your predicted table and update the confidence flags. If the recruiter adds a stage the reviews never mentioned, that is a signal the loop changed recently and your older bodies are stale, not that you misread them.

Three triggers should send you back to the reviews. First, a recruiter mention of a new format, such as an AI screen or a take-home, that your reconstruction lacked. Second, a change in the live posting between your first read and your interview. Third, a gap of more than a few weeks between reconstruction and your first round, because interviews-per-hire keeps drifting upward and loops get longer, not shorter. Each trigger costs ten minutes to check against the modal list and is worth it.

Finally, feed what you learn back the other way. Once your loop is done, your own review is the recent, role-matched, offer-or-reject-tagged body that the next candidate needs most. The reason this method works at all is that people post; the reason it stays reliable is that recent people post. Write the review you wish you had found, with the stage sequence, the format, the duration, and the themes, and you close the loop for the person who reconstructs this same company next.

## Frequently asked questions

### How many interview rounds should I expect at a specific company?

There is no universal number; it depends on the company, the role, and the level. Reconstruct it from public reviews by tagging the round sequence in 15 to 30 free-text bodies and keeping only stages that recur across independent reports. As a market anchor, technical roles averaged 20 interviews per hire in 2024, up from 14 in 2021, so older reviews tend to undercount. Always confirm your final stage list against the recruiter's schedule, which is authoritative for your loop.

### How do I read interview difficulty ratings correctly?

Read them against the real distribution, not against the 1-to-5 scale. Only 10.5% of Glassdoor interviews were classified as difficult and 1.2% as very difficult, and the average rating is 2.8, so scores cluster near 3. A role reading 3.4 is genuinely near the top of the real distribution. Also account for rejection bias: rejected candidates rate harder, correlation minus 0.49 with offer rate, so a page skewed toward rejects reads scarier than the loop is.

### How many reviews do I need before a per-role number is trustworthy?

No platform publishes a per-role minimum, so treat it as a judgement call. Flag any figure resting on fewer than roughly 10 to 15 role-level reviews as directional only. Rating-scale means stabilise faster than proportions, but thin samples are dominated by self-selection: people with extreme experiences post more often. Netflix and Snap looked worst on Glassdoor largely because they had the fewest reviews, a textbook thin-sample trap.

### What do I do when the reviews and the recruiter's schedule disagree?

Let the recruiter-provided schedule and the current posting win for the stage list, because reviews lag reality and span years and roles. Use the reviews to fill the gaps the recruiter leaves unstated, mainly question themes and difficulty, not to override the confirmed rounds. This matters because sources genuinely conflict: one guide described a three-round loop for a company while others described a two-screen-plus-five-onsite loop for the same role.

### Can I predict interview length by company from public data?

Partly. Company pages report an average days-to-hire, but that figure blends all titles and can drift with hiring volume. Use it as a directional band, then adjust with role-matched, recent reviews that state duration in their bodies. For scale, Glassdoor's own overall average rose from 12 days in 2010 to 23 days in 2013, and U.S. time-to-hire runs roughly 44 days by one 2024-2025 benchmark, so treat any single number as a starting point, not a promise.

### Is a low positive-experience percentage a sign of a bad interview loop?

Not necessarily. Positive-experience percent is partly an offer-outcome proxy, not a pure loop-quality measure. There is a 0.75 correlation between onsite-to-offer ratio and positive experiences, and roughly 70% of candidates with a positive experience got offers versus 20 to 30% of neutral and negative ones. A low positive-percent often signals a selective funnel rather than a broken process, so read it alongside difficulty and days-to-hire.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/candidates/guides/reconstruct-interview-loop-reviews*
