# The Design Portfolio Score: Advance, Screen, or Pass a Designer

*You will score any designer's public portfolio on five dimensions and output a defensible advance, screen-with-caution, or pass decision plus a seniority level.*

- Canonical URL: https://www.refolk.ai/guides/design-portfolio-score
- Pillar: Recruiting and sourcing
- Format: Framework
- Published: 2026-08-29
- Last reviewed: 2026-08-29
- Reading time: 15 min

Deciding whether to spend a scarce interview slot on a designer, and at what level, is a judgement call you make from public evidence alone: a portfolio site, a Behance page, a Dribbble grid. This guide is for in-house recruiters, sourcers, talent leaders, and founders doing their own hiring. It gives you a reviewer-side rubric with named dimensions, scoring thresholds, and a seniority mapping, so you can turn any portfolio into a defensible advance, screen-with-caution, or pass decision plus a level you can put in front of a hiring manager.

Most guides on this topic coach the designer on building the portfolio. This one sits on the other side of the table. It treats the work as evidence and scores it the way a code reviewer scores a repo: what a signal proves, and what it looks like when it lies.

## Why a portfolio score needs a fixed rubric

A portfolio score is a repeatable read of public design work against fixed dimensions, so that two reviewers reach the same verdict. Without a rubric, "read the work" collapses into taste, and taste does not survive being handed to a hiring manager who asks why.

The problem is mechanical, not aesthetic. Reviewers place a candidate in seconds and finish the first pass in under a minute. That time budget caps what any first-pass score can measure. If you try to assess craft nuance in the same breath as role fit, the triage corrupts: the shiniest portfolio wins the slot, and polish is the easiest thing in a portfolio to fake and the least predictive of whether someone solves your actual problems.

So the rubric does two things. It separates the fast triage from the slow read, and it names concrete dimensions instead of adjectives. The dimensions below are what named hiring managers screen on, ordered by how early in the review they apply.

**18,906 - US profiles matching Product Designer or Senior Product Designer titles**

From Refolk's index of professional profiles; the raw pool is large, so the scoring bar is what protects your interview slots.

## The five scoring dimensions and what each proves

Score every portfolio on five dimensions: role fit, seniority signal, decision quality, impact delta, and craft. The first two are triage; the middle two are the seniority read; craft comes last and carries the least weight. Each dimension proves something specific, and each has a characteristic way of lying.

| Dimension | What it proves | What it looks like when it lies |
|---|---|---|
| Role fit | The candidate is the kind of designer this role needs | A generalist grid that fits any role and none |
| Seniority signal | Inferred scope, judgement, ownership | Big titles on small, well-defined work |
| Decision quality | Real judgement under competing goods | A tidy process diagram nobody believes happened in order |
| Impact delta | The work changed something | Final mockups with no before or after |
| Craft | Execution ability, one dimension only | A beautiful personal-brand site with no owned work |

Two of these deserve emphasis. Decision quality is what actually gets hired: strong portfolios expose the structure behind decisions, and where decisions are implicit, judgement cannot be evaluated. A missing impact delta is a critical red flag for mid and senior roles, because portfolios without impact collapse into surface-level review. Leslie Yang, a hiring manager who writes on seniority signals, names four indicators reviewers screen fast on: initiative, product thinking, scope, and innovation, adding that seniors must also index high on craft, empathy, storytelling, and leadership.

> **Rule:** Craft is scored last, never first
>
> Visual polish is the most fakeable dimension and the least diagnostic of whether someone solves your problems. Assess craft only after decision quality and scope, and never let it drive the triage.

## What the time budget lets you actually measure

The first screen can only assess role fit, a seniority signal, and relevance, because that is all the published time budget buys. Everything else is a second, slower pass reserved for candidates who survive triage.

The figures cluster tightly. Most hiring managers spend under one minute on a portfolio. One recruiter-focused source cites studies showing 80% of recruiters spend three minutes or less. And the placement threshold is brutal: if a reviewer cannot place you in under five seconds, they will not invest further effort. This is why the rubric front-loads role fit and seniority signal; they are the only dimensions that fit inside the window.

| Reviewer | Time budget | What it can assess |
|---|---|---|
| Hiring managers | under 1 min | role fit and seniority signal only |
| Recruiters (80%) | under 3 min | plus one relevant case skim |
| Placement threshold | under 5 sec | initial "can I place them" |

The practical consequence: run the fast filter first, and only pay the five-to-eight-minute case-study reading cost on the survivors. Treat the first screen as a gate, not a grade.

#### Where portfolios drop out

| Stage | Figure | Note |
| --- | --- | --- |
| Placed by role | 5 sec | can you say what kind of designer this is |
| Passed triage | 3 min | role fit, seniority signal, relevance |
| Case study read | 8 min | decision quality and impact delta |
| Scored to a level | 3 min | advance, screen, or pass |

*Each stage costs more reviewer time, so the fast filter protects the expensive read.*

## Reading seniority: scope, ambiguity, autonomy

Map evidence to a level using three axes that recur across leveling frameworks: scope of impact, level of ambiguity handled, and degree of autonomy. These are the axes that make two reviewers agree, because they describe the work rather than rate it.

Scope of impact runs from one task, to one team, to multiple teams, to the whole company. A clear signal of seniority is large scope: owning or co-owning a large feature set or product surface, an app, or a major vertical. Ambiguity runs from well-defined problems handed to the designer, to problems the designer defined themselves. Autonomy runs from close guidance and review to independent operation. Rippling's public design ladder uses concrete competencies in the same spirit: Impact, Influence, Deliver, Communicate, Operate, and Craft.

Seniors prove level through depth, not breadth. They show two or three deep, complex case studies, each demonstrating strategic thinking and cross-functional collaboration: untangling a legacy system, or designing a feature that touches five departments. They show decisions and trade-offs rather than process diagrams, because the places where you chose between two competing goods, or traded scope against time, are where seniority lives. Juniors own smaller, simpler pieces with more guidance; a focused portfolio of three to five strong case studies signals judgement, while 60 screens with no case studies signals its absence.

#### The seniority read

Horizontal axis runs from Narrow scope, well-defined problems to Broad scope, self-defined problems. Vertical axis runs from Shallow, many shots to Deep, few case studies.

| Quadrant | What it means |
| --- | --- |
| Junior signal | Advance to entry level, screen for growth |
| Range without depth | Screen with caution, probe for real ownership |
| Un-scorable on judgement | Route to a case walkthrough before deciding |
| Senior signal | Advance at level, verify ownership in interview |

*Depth in one problem space reads senior; breadth without depth reads junior.*

The named failure to avoid here is the adjective ladder: rating each level basic, strong, advanced, or exceptional. Two managers cannot agree where basic ends and strong begins, so calibration falls back to instinct. Force scope, ambiguity, and autonomy language into your notes and the disagreement mostly disappears.

## The scoring procedure

Run the eight steps below in order. The first is done before you open a portfolio; the rest move from cheap triage to the expensive read, ending in a written verdict a second reviewer could reproduce. Note that sources disagree on one point of order: some managers screen craft before decision quality, but the fast-filter sources put seniority and relevance first, and this procedure follows them.

#### Score a portfolio in eight passes

1. **Define target level and role fit first** - Before opening any portfolio, write the scope, autonomy, and decision-rights criteria for the open level. You should be able to state in two sentences what changes between the target level and the one below.
2. **Run first-screen triage** - Land on the homepage, read the positioning, and scan for one relevant case study, all under three minutes. Place the person by role, infer a seniority signal, and judge relevance; if you cannot place them in seconds, it ends here.
3. **Read the most relevant case study for decision quality** - Pick the single most relevant case and look for problem framing, trade-offs, and where judgement was exercised. You should be able to name two real decisions and the alternatives rejected.
4. **Score the impact delta** - Look for before and after metrics, or absent metrics, an observable behaviour change with a stated reason metrics were unavailable. Each scored case needs an impact line or a documented reason it lacks one.
5. **Assess scope and ownership** - Distinguish a "made the icons" contribution from owning a surface, and tag each case as owned, co-owned, or contributed. This tag feeds the seniority read.
6. **Check craft on a secondary surface** - Use Dribbble shots or high-fidelity screens for craft only, never as problem-solving evidence. Score craft as one dimension without letting it touch the seniority read.
7. **Handle NDA and research-heavy exceptions** - Look for a stated confidentiality note, a password gate, or an anonymised case, and do not penalise missing metrics when validated-learning outcomes are shown. Score confidential work on scope, role, and outcomes.
8. **Map to level and record the decision** - Place the candidate against scope, ambiguity, and autonomy, then output advance, screen-with-caution, or pass plus a level. Write a rationale a second reviewer could reproduce.

Once you have scored a handful of portfolios this way, the harder problem is finding enough qualified people to score. Describing the exact evidence you want in plain English is faster than boolean strings across three sites.

I ran this search: `Senior product designers who owned an end-to-end payments or checkout flow with a documented conversion lift.` - [see the full result list](https://www.refolk.ai/s/1d6qnaegnt).

*Returns designers whose public profiles show owned scope and an impact delta, which are the two dimensions hardest to fake in a portfolio.*

Because [Refolk](/) reads public LinkedIn, the open GitHub graph, and the open web, I can front-load the ownership and impact signals into the query itself, so the portfolios you open are already biased toward the ones that will score.

## What each surface proves and what it hides

Surface choice predicts what evidence exists before you read a word. Dribbble proves craft only; Behance can carry end-to-end problem-solving; a personal site can carry either or mislead through over-styling.

Dribbble shows isolated shots, which prove craft but not problem-solving. Its format structurally cannot carry the framing, trade-offs, and outcomes that decision quality needs, so a Dribbble-only candidate is un-scorable on judgement regardless of talent. Do not pass them for it; route them to a case walkthrough. Behance is built for full projects, with structured case-study pages that walk through problem, process, and outcomes, so it can carry the decision-quality read. A personal site is the wild card: a highly stylised, experimental personal-brand website often results in poor navigation and a lack of focus on the work itself, and if the portfolio's own user experience is bad, that is an immediate disqualifier from a designer.

> A Dribbble-only candidate is not a weak candidate; they are an un-scorable one until you see a case.

Two context facts help you calibrate. Behance was founded in 2006 and acquired by Adobe in December 2012; Dribbble was founded in 2009 and is social and shot-driven by design. Neither origin changes the rule: read shots for craft, read case pages for judgement, and never confuse the two.

## How this goes wrong: failure modes and false positives

The most valuable part of any scoring standard is the list of ways it misfires. Below are the failure modes that produce wrong verdicts, split into false positives that advance a weak candidate and false negatives that pass a strong one.

- **Polish read as competence.** A beautiful personal-brand site with no owned work reads as senior. Check: can you name two real decisions and their rejected alternatives? If not, it is craft, not judgement.
- **Dribbble shots scored as problem-solving.** Shots prove craft only. Treat them as one dimension, never as the seniority read.
- **Double Diamond as evidence.** A process diagram shows sequence, not judgement. A tidy linear process is itself a flag, because it hides where trade-offs were actually made.
- **Contribution inflation.** "I did research, UI, and testing" claimed in a silo can hide minimal real ownership. Check the team makeup and the candidate's specific role before tagging a case as owned.
- **Fabricated research.** Vague usability claims such as "we tested with five users and they liked it" with no specifics or honest gaps signal invented process. Look for concrete numbers, methods, and what did not work.
- **Over-leveling from range.** Breadth without depth in a problem space reads as junior in the current market. Do not let a wide grid inflate the level.

The false negatives are subtler and cost you senior talent. NDA and research-heavy work create systematic false-pass errors when you score on metrics alone.

> **Watch out:** A metrics-only rule rejects your senior enterprise pool
>
> Confidential work legitimately omits client names, assets, and sometimes metrics. You can craft a case study on scope, role, goals, and outcomes with generic or hypothetical visuals. A metrics-only score silently rejects senior fintech and enterprise designers, the exact pool most likely to be NDA-bound.

For juniors, impact does not always mean metrics; it can mean validated learning, improved usability, or clarified direction. The test for a confidential case is not "are there numbers" but "is the omission deliberate": look for a stated confidentiality note, a password gate, or an anonymised case. Morgane Peng, a design director who has written on showcasing NDA-protected work, is one of several practitioners who confirm confidential cases can be scored on described scope and outcomes rather than shown assets.

One more calibration point: student projects older than two years signal inexperience, not range, and a linear process presented as a full case study is a flag in its own right. The Fountain Institute's write-up on what gets senior product designers hired names both.

## Turning the score into a verdict and a level

Convert the five dimension scores into one of three verdicts plus a level, and write the rationale in scope, ambiguity, and autonomy language so a second reviewer can reproduce it. The verdict is a routing decision; the level is a hypothesis to test in the interview.

Advance means the case studies show owned scope, real decisions with rejected alternatives, and an impact delta or a documented reason it is absent. Screen-with-caution means the evidence is promising but one core dimension is unverifiable from public work, most often ownership or impact, so the interview must confirm it. Pass means a reliable red flag fired: only final mockups, a broken portfolio experience, or a case that is process diagrams with no visible judgement.

**Portfolio scoring record**

```
Candidate:
Target level and role:
Role fit (fits this role / generalist / mismatch):
Most relevant case study:
  Two real decisions + rejected alternatives:
  Impact delta (metric / behaviour change / documented reason absent):
  Ownership (owned / co-owned / contributed):
Scope of impact (task / team / multi-team / company):
Ambiguity handled (well-defined / self-defined):
Autonomy (guided / independent):
Craft (secondary surface only):
NDA or research note present? (yes / no):
Verdict: advance / screen-with-caution / pass
Proposed level + one-line rationale:
```

*Fill one per candidate. Keep the rationale in scope, ambiguity, and autonomy language so a second reviewer can reproduce the verdict.*

The seniority read is a claim, so hold it loosely. Depth in one problem space is the strongest public signal of senior work, and breadth without depth is the strongest signal to under-level. When the two conflict, tag it screen-with-caution and let the interview break the tie.

## Keeping the rubric calibrated and stocked

A scoring rubric drifts unless you re-calibrate it against outcomes and keep enough qualified candidates flowing to use it on. Two things keep this work current: checking your verdicts against how hires actually performed, and weighting your sourcing toward the scarce competencies your roles need.

Scarcity is measurable and uneven by title. In Refolk's index, 1,048 US Product Designers list Design Systems as a skill, against only 371 US UX Designers. If a role is systems-heavy, weight Product Designer sourcing accordingly, because the UX-titled pool is thin on that competency.

| Title (US) | Profiles with Design Systems skill |
|---|---|
| Product Designer | 1,048 |
| UX Designer | 371 |

The same index shows the raw geographic pool: the US Product Designer pool is 3.73 times the UK pool for the same two titles. That ratio is a sourcing input, not a quality signal, and it should shape how wide you cast before the rubric ever runs.

| Market | Matching profiles | Share of combined pool |
|---|---|---|
| United States | 18,906 | 78.9% |
| United Kingdom | 5,064 | 21.1% |

Before you call a scoring pass done, run this check.

#### Before you record the verdict

- [ ] You defined the target level's scope, ambiguity, and autonomy before opening the portfolio
- [ ] You can name two real decisions and the rejected alternatives from at least one case
- [ ] Each scored case is tagged owned, co-owned, or contributed
- [ ] Craft was scored from a shot surface only, never as the seniority read
- [ ] You checked for an NDA note before penalising missing metrics
- [ ] Your rationale uses scope, ambiguity, and autonomy, not basic, strong, or advanced
- [ ] A second reviewer could reproduce the verdict from your written record

Re-calibrate on a cadence: pull the scores you gave against how those hires performed, and adjust which red flags actually predicted a bad hire versus which just felt like one. The dimensions are stable, but the thresholds are yours to earn from your own outcomes.

## Frequently asked questions

### How long should I actually spend reviewing a designer's portfolio?

Budget under three minutes for the first screen and reserve deeper reading for candidates who survive it. Published figures show most hiring managers spend under one minute, and 80% of recruiters spend three minutes or less. The first pass only buys role fit, a seniority signal, and relevance. If the candidate advances, spend five to eight more minutes on the single most relevant case study reading for decision quality.

### What are the biggest UX portfolio red flags for a hiring manager?

Only final mockups with no path to the solution, a missing impact delta on mid or senior cases, and the portfolio's own broken navigation, which is an immediate disqualifier. Also watch for a linear Double Diamond presented as a case-study structure, contribution inflation where one person claims research, UI, and testing in a silo, and fabricated research such as vague 'we tested with five users and they liked it' claims with no specifics.

### How do I judge a designer's seniority from their portfolio?

Read for scope, ambiguity handled, and autonomy rather than adjectives. Seniors own or co-own a large surface, show two or three deep case studies with real trade-offs, and expose where they chose between competing goods. Juniors own smaller, well-defined pieces with more guidance. Depth in one problem space reads as senior; breadth without depth reads as junior in the current market.

### Should I penalise a designer whose work is under NDA?

No. Confidential work legitimately omits client names, specific assets, and sometimes metrics. Score it on scope, role, goals, and outcomes described with generic or hypothetical visuals. A metrics-only rule systematically rejects senior enterprise and fintech designers, who are the pool most likely to be NDA-bound. Look for a stated confidentiality note, a password gate, or an anonymised case as the signal that the omission is deliberate, not empty.

### Can I score a designer from Dribbble or Behance alone?

Dribbble proves craft only. Its shot format structurally cannot carry problem-solving, so a Dribbble-only candidate is un-scorable on judgment regardless of talent; route them to a case walkthrough. Behance supports full case-study pages that walk through problem, process, and outcomes, so it can carry decision-quality evidence. Score craft from either, but never read isolated shots as a seniority signal.

### Why do two reviewers disagree on the same portfolio?

Usually because the rubric uses an adjective ladder such as basic, strong, advanced, and exceptional. Two managers cannot agree where basic ends and strong begins, so calibration falls back to instinct. Switch to concrete axes: scope of impact, level of ambiguity handled, and degree of autonomy. Those are the axes that recur across leveling frameworks and let two reviewers grade the same portfolio the same way.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/design-portfolio-score*
