# The Kaggle Profile Read: Applied Skill, Contest Overfit, or Noise

*You will be able to score any Kaggle profile by category, discount the farmed and overfit parts, and call it real skill, contest overfit, or noise.*

- Canonical URL: https://www.refolk.ai/guides/kaggle-profile-read-hiring
- Pillar: Engineering and open source
- Format: Framework
- Published: 2026-10-06
- Last reviewed: 2026-10-06
- Reading time: 15 min

You are looking at a data-science candidate's Kaggle profile and need to decide, in minutes, whether its rank and medals prove the applied machine-learning skill you are hiring for. This guide is a scoring rubric for technical screeners, engineering managers, technical founders, and technical sourcers. It gives you four category reads, two discounts to apply, and one defensible verdict: real applied skill, contest overfit, or noise.

Kaggle is a competition and community platform for data science. A profile carries a single badge, but that badge hides how it was earned. The whole job here is to stop reading the badge and start reading the record underneath it.

## What a Kaggle profile actually shows

A Kaggle profile shows one badge, and that badge is the single highest tier across four independent categories. It is possible to rank differently in each category, and the highest tier you achieve in any category is what appears under your avatar. So a "Grandmaster" label tells you someone hit the top tier somewhere, not where.

The four categories are Competitions, Notebooks, Datasets, and Discussions. The five tiers, from lowest to highest, are Novice, Contributor, Expert, Master, and Grandmaster. Tiers are awarded on the medals earned in each category, and the categories do not measure the same thing:

- **Competitions** measures predictive modeling against a held-out metric. This is the category that maps most directly to applied ML skill.
- **Notebooks** measures published analysis and code that other users upvote.
- **Datasets** measures datasets you publish and other users upvote.
- **Discussions** measures forum participation and other users upvote.

Three of the four categories rest on upvotes. Only one rests on leaderboard results. That is the single most important fact in this guide: the displayed label systematically over-reports competition skill, because Kaggle shows the max across categories and the cheapest category to climb is a forum.

> **Rule:** Read the category, never the badge
>
> A "Master" or "Grandmaster" badge is unreadable until you open the four-category breakdown. Treat the headline tier as a claim to verify, not a credential to accept.

## Tier scarcity and the cheapest path to each badge

The tiers are genuinely scarce at the top, which is why they carry weight, but each tier has a cheapest path that is far easier than the hardest path. As of April 2, 2025, out of 23.29 million Kaggle accounts, 2,973 have reached Master and 612 have reached Grandmaster. Against a base of millions of accounts, those are tiny fractions.

**612 - Kaggle Grandmasters out of 23.29M accounts (Apr 2, 2025)**

That is 0.0026% of accounts, which is why the tier is prized - but it can be earned in four different categories.

The table below shows each tier's scarcity and the cheapest way to display its badge. The "cheapest path" column is the one a screener must watch, because it is the route that proves the least.

| Tier | Count (Apr 2025) | Share of 23.29M accounts (derived) | Cheapest path to the badge |
| --- | --- | --- | --- |
| Grandmaster | 612 | 0.0026% | 15 gold notebooks, or 50 gold + 500 discussions, instead of 5 competition golds including a solo gold |
| Master | 2,973 | 0.0128% | 50 silver + 200 discussions, instead of 1 gold + 2 silver in Competitions |
| Expert | not published | n/a | 50 bronze discussions, comment-farmable |

Discussion Expert is the cheapest headline tier of all. It requires only 50 bronze medals, and one user reached it only after commenting on more than 90 notebooks and topics. That is engagement volume, not modeling. A candidate can legitimately carry an "Expert" badge earned entirely in a forum.

The scarcity cuts the other way for one specific tier. A Competitions Grandmaster requires five golds including at least one solo gold - a competition won with a single-person team. You cannot upvote your way to that. Of all four categories, the Competitions tier is the one that resists farming, which is exactly why it is worth the effort to confirm the badge sits there.

> The badge hides its own cheapest path, and the cheapest path is almost always a forum, not a leaderboard.

## Reading competition medals against team size

A competition medal means nothing until you know how many teams were in the field, because the medal bands scale with team count. A gold in a small event and a gold in a large event are different achievements that display identically. Medals are assigned on the final private leaderboard when the competition ends, not on the public board.

Here are the official bands:

| Team count | Bronze | Silver | Gold |
| --- | --- | --- | --- |
| 0-99 | top 40% | top 20% | top 10% |
| 100-249 | top 40% | top 20% | top 10 |
| 250-999 | top 100 | top 50 | top 10 + 0.2% |
| 1000+ | top 10% | top 5% | top 10 + 0.2% |

The practical spread is wide. A 500-team competition awards gold to the top 11 teams. A 5,000-team competition awards gold to the top 20 teams. In a sub-100-team event, gold is simply the top 10%. So a candidate with three golds from small Playground-style fields has done something real but far less selective than one gold in a 5,000-team featured contest.

#### Why team size changes what a gold proves

| Stage | Figure | Note |
| --- | --- | --- |
| 0-99 teams | top 10% | a wide gold band |
| 500 teams | top 11 | a tight gold band |
| 5000 teams | top 20 | near-elite, top fractions of a percent |

*The same gold medal demands a very different finish depending on how many teams entered.*

When you grade a Competitions record, do not count medals. Weight them. One 5,000-team gold outranks a handful of sub-100-team golds for predicting applied skill under real competitive pressure.

## Solo versus team, and the overfit check

Two things can inflate a competition record even after you weight for team size: a medal earned by a large team, and a placement that overfit the public leaderboard. Both have specific checks.

### Solo versus team

The whole competition team gets a medal regardless of who made the submissions. A five-person gold and a solo gold look the same on a profile. Only the leaderboard roster separates them. This matters most for the Grandmaster solo-gold requirement, which is the one place the platform forces individual proof. For every gold you care about, open the competition leaderboard and read the team. A candidate who is one of five names on every gold has different evidence than one with a solo win.

Note the asymmetry with Notebooks: only the notebook owner gets a notebook medal, while in a competition the whole team gets the medal no matter who submitted. So team structure inflates competition credit specifically.

### The shakeup check

Shakeup is the overfit lie-detector. The public leaderboard is computed from a subset of the test data, the final results use a disjoint subset, and the difference between them indicates whether a model overfit the public board. The shakeup metric is essentially the mean absolute percentage change in rank for all entrants between public and private standings.

The mechanism is quantified. The error between public leaderboard and true performance can be as large as O(sqrt(k/n)) for k submissions, so heavy submitters drift most. A candidate who probed the public board with hundreds of submissions and finished strong there may collapse privately. A placement that survives the private split is the real signal.

> **Watch out:** Never trust a mid-competition screenshot
>
> A strong public-leaderboard rank can collapse on the private set. Medals are awarded only on the final private leaderboard. If a candidate cites a rank that is not a final medal, it may be a public-board artifact that never held.

Ask me this: `Kaggle Competitions Grandmasters with a solo gold medal who now work as ML engineers in the United States.` - [run the search](https://www.refolk.ai/start?q=Kaggle%20Competitions%20Grandmasters%20with%20a%20solo%20gold%20medal%20who%20now%20work%20as%20ML%20engineers%20in%20the%20United%20States.).

*This returns people whose top-tier badge sits in the uncounterfeitable category and who have moved into production roles, so you start from verified modeling depth rather than a farmed badge.*

Verifying all of this by hand means opening leaderboards, reading rosters, checking team counts, and cross-referencing employers one profile at a time. [Refolk](/) lets you ask for the combination you actually want - the tier, the category, the solo gold, the current role - in plain English and get matched people back, so the scoring work starts from a pool that already clears the structural bars.

## The scoring procedure

Work the profile in this order. The first six steps are a screener's job and take under half an hour; the last two pull in a sourcer and a hiring manager. Note that practitioners disagree on step order: some read the headline badge first, while recruiters in the sources argue you should weight production-deployment evidence first and treat Kaggle as supporting. I put the badge decode first because it is fast and it tells you whether the rest of the read is even worth doing.

#### Scoring a Kaggle profile

1. **Identify the headline tier and its category** - Open the profile and read the badge under the avatar; it is the single highest tier across all four categories. Know whether the Master or Grandmaster label came from Competitions or from Discussions, Notebooks, or Datasets.
2. **Decompose by category** - Record tier and medal counts separately in Competitions, Notebooks, Datasets, and Discussions. You now have four independent reads, not one headline.
3. **Grade Competitions against team size** - For each competition medal, check the team count and which band it fell in. You can now tell a sub-100-team gold from a 5,000-team gold.
4. **Check solo versus team on golds** - Open each gold competition's leaderboard and inspect the candidate's team roster. You know whether a gold reflects individual skill or a large team, and whether a solo gold exists.
5. **Estimate overfit via shakeup** - Compare public versus private standing where available, or compute the shakeup. You know whether placements held on the private set or collapsed.
6. **Discount upvote-farmed tiers** - If the headline comes from Discussions or Notebooks, read it as engagement and communication, not modeling. The farmed portion is now separated from the competitive portion.
7. **Verify identity and employer** - Confirm the profile maps to a real person and the stated employer. Kaggle verifies ranked-competition accounts but not bio claims, so cross-check the employer against the bio.
8. **Land the verdict** - Combine the reads into real applied skill, contest overfit, or noise, and pair with evidence of production work that Kaggle does not test.

## What the verdict means

The verdict is a three-way call, and each outcome points to a different next action. The point of the rubric is to land one of these defensibly, not to produce a number.

#### The verdict grid

Horizontal axis runs from Thin or farmed signal to Deep weighted competition record. Vertical axis runs from Overfit or unverified to Private-set-proven.

| Quadrant | What it means |
| --- | --- |
| Noise | Treat Kaggle as neutral; judge on other evidence entirely |
| Contest overfit | Credit modeling instinct, probe for discipline and generalisation |
| Upvote-farmed tier | Read as communication and community, not modeling |
| Real applied skill | Strong modeling signal; now verify production separately |

*Competition depth on one axis, farmed or overfit signal on the other, gives four reads.*

- **Real applied skill.** A weighted Competitions record with large-field medals, confirmed solo or genuine team contribution, and placements that survived the private split. This is strong evidence of modeling, feature engineering, and iteration under a metric.
- **Contest overfit.** Strong public standings that collapsed privately, or a record built from heavy submission probing. Credit the instinct, but probe directly for generalisation and discipline.
- **Noise.** A headline tier earned in Discussions or Notebooks, small-field medals, or a record with no recent competitive activity. Not disqualifying, but Kaggle adds nothing here; judge the candidate on everything else.

One caveat sits over all three. Competitions do not test real-world concerns: model explainability, problem definition, acceptable false-positive rates, what data you are allowed to use, or the gap between real-time and historical data. Competition code also has no incentive to be production-shaped, since the only goal is to win. A Grandmaster who never deployed a model can be less useful than an engineer who shipped five models serving real users. So even "real applied skill" is a read on modeling depth only, to be paired with separate production evidence.

## How this read goes wrong

Most bad Kaggle reads come from eight specific traps. Each has a tell and a check. This is the section to keep open while you score.

1. **Headline-tier confusion.** A Grandmaster badge may be a Discussions or Notebooks Grandmaster, not Competitions. *False positive:* reading forum-upvote volume as modeling skill. *Check:* open the category breakdown; the highest tier in any category is what shows under the avatar.
2. **Small-field medals.** A gold in a sub-100-team event needs only the top 10%. *False positive:* counting all golds equally. *Check:* team count against the medal-band table.
3. **Team golds read as solo skill.** The whole team gets the medal regardless of who submitted. *False positive:* crediting one member for a five-person gold. *Check:* the leaderboard team roster; confirm solo golds.
4. **Public-leaderboard overfit.** A strong public rank can collapse privately. *False positive:* trusting a mid-competition rank or screenshot. *Check:* shakeup or final private standing.
5. **Upvote farming.** Discussion Expert is an easy way into the Expert tier. *False positive:* treating comment volume as applied ML. *Check:* whether the medals are competition medals or discussion medals.
6. **Stale points versus medals.** Points decay over time but medals are permanent. *False positive:* a high historical rank with no recent competitive activity. *Check:* medal dates.
7. **Production assumption.** A top competitor may have zero deployment experience. *False positive:* assuming competition skill means shipping skill. *Check:* separate production evidence from repos or work history.
8. **Identity mismatch.** Kaggle's own verification covers only ranked-competition accounts, not the bio's employer claim. *False positive:* accepting the stated employer because the account is verified. *Check:* cross-reference the employer independently.

> **Tip:** One tool does the overfit math for you
>
> The open davidthaler/shakeup project scrapes leaderboards and computes the shakeup metric. If you need to grade overfit at scale rather than eyeball one competition, this removes the manual public-versus-private comparison.

## Verifying identity and the market read

Confirm the profile is the person in front of you before you trust any of it. Since 2023, Kaggle introduced an account verification process for competitions that reward progression points, medals, and prizes, using Persona selfie identity verification solely to confirm a real, live human and to prevent creation of multiple accounts, plus phone and SMS verification. That is useful: it means a ranked Competitions record is tied to a verified human. But it does not verify the bio's employer claim, so cross-check the stated employer against the profile and against public records.

The market treats the tier as prized but not deciding, and the evidence backs that reading. Named demand exists - a referral posting sought a Kaggle Grandmaster data scientist at roughly $56 an hour. Yet the tier rarely stands alone as a headline.

In Refolk's index, US profiles with Data Scientist or ML Engineer titles mentioning Kaggle in their headline total 21, split evenly between 10 Data Scientists and 10 ML Engineers. The table below breaks down how the stronger "Competitions Master" wording appears across markets.

| Signal in headline | Market | Matching profiles | Notable employers |
| --- | --- | --- | --- |
| Mentions Kaggle (DS/MLE) | United States | 21 | Google DeepMind, Microsoft, Experian, Kensho |
| "Kaggle Competitions Master" | United States | 5 | eBay, KKR, Bidease |
| "Kaggle Competitions Master" | India | 0 | none in index |

**5 of 21 - US Kaggle-mentioning DS/MLE profiles that foreground "Competitions Master" (Refolk index)**

About 24%, and US-concentrated. Even the strongest tier is rarely the sole headline, consistent with it being a supporting signal.

The read: only about a quarter of candidates who mention Kaggle at all foreground the specific "Competitions Master" wording, and the phrasing concentrates in the US market. People who carry the strong, uncounterfeitable tier tend to pair it with a substantive employer and role rather than leading with the badge alone. That is the behaviour of a supporting signal, which is how you should weight it too. This is a small-sample, headline-text observation from the index, not a population census - treat it as directional.

## Before you call it done

Run this checklist against every Kaggle profile before you record a verdict. If any item fails, the verdict is not defensible yet.

#### Kaggle profile sign-off

- [ ] I opened the four-category breakdown and know which category earned the headline tier.
- [ ] Every competition medal is weighted against its team count, not counted flat.
- [ ] I checked the team roster on each gold and know which are solo and which are team.
- [ ] I compared public versus private standing, or shakeup, on the placements I am crediting.
- [ ] I separated any Discussions or Notebooks tier from the competition record.
- [ ] I checked medal dates and noted whether competitive activity is recent or stale.
- [ ] I verified the profile maps to a real person and cross-checked the stated employer.
- [ ] I paired the Kaggle read with separate evidence of production or deployment work.

## Keeping the read current

Two things in this guide are time-sensitive by mechanism, not by date. First, the population counts: 612 Grandmasters and 2,973 Masters are the April 2, 2025 figures, and both numbers rise over time, so re-pull the current counts from the public Kaggle population before you cite scarcity in a decision. The mechanism holds even as the numbers move - Grandmaster stays a tiny fraction of accounts.

Second, the medal bands and tier thresholds are set by Kaggle's progression system and have changed before. Confirm the current bands on the official progression page rather than trusting a memorised threshold. The shakeup logic, the solo-gold rule, and the four-category structure are stable, so the rubric's shape does not move even when the exact numbers do.

Finally, remember what the rubric is for. It tells you whether a Kaggle record proves modeling skill. It does not tell you whether the candidate can define a problem, ship a model, or run it in production, because Kaggle does not test those. Use this read to settle the modeling question fast and defensibly, then spend your remaining screening time on everything Kaggle cannot show you.

## Frequently asked questions

### What does Kaggle Grandmaster mean for recruiting?

It means a top-tier Kaggle rank, but only in one of four categories. A Competitions Grandmaster is genuinely scarce - 612 out of 23.29 million accounts as of April 2, 2025 - and requires five golds including a solo gold, which cannot be farmed. A Discussions or Notebooks Grandmaster is earned through upvotes, so always open the category breakdown before you read the badge as modeling skill.

### Does Kaggle rank translate to job skill?

Partly. A high Competitions rank proves strong modeling, feature engineering, and iteration under a fixed metric. It does not prove problem definition, model explainability, real-time data handling, or deployment, because competitions do not test those. Treat Kaggle as evidence of applied modeling and pair it with separate evidence of shipped, production work before you decide.

### Is a Kaggle Master better than production ML experience?

Not on its own. A Grandmaster who never deployed a model can be less useful than an engineer who shipped five models serving real users. Competition code has no incentive to be production-shaped. Use Kaggle to confirm modeling depth, then weight deployment evidence at least as heavily for any role that ships.

### How do I tell a Discussions tier from a Competitions tier?

Open the four-category breakdown on the profile. The badge under the avatar shows only the single highest tier across all categories, so it systematically over-reports competition skill. Discussion Expert needs only 50 bronze medals, which one user earned by commenting on 90-plus notebooks and topics. Confirm the medals are competition medals, not upvotes.

### How do I check whether a Kaggle gold was overfit?

Compare the public leaderboard standing to the final private standing, or compute the shakeup, the mean absolute percentage change in rank between the two. The public board uses a subset of test data and the final result uses a disjoint subset, so a large drop signals overfit. Heavy submitters drift most, since public-versus-true error scales with the number of submissions.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/kaggle-profile-read-hiring*
