Refolk
FrameworkRecruiting and sourcing

The Design Portfolio Score: Advance, Screen, or Pass a Designer

You will score any designer's public portfolio on five dimensions and output a defensible advance, screen-with-caution, or pass decision plus a seniority level.

15 min readLast reviewed August 29, 2026Read as Markdown

Deciding whether to spend a scarce interview slot on a designer, and at what level, is a judgement call you make from public evidence alone: a portfolio site, a Behance page, a Dribbble grid. This guide is for in-house recruiters, sourcers, talent leaders, and founders doing their own hiring. It gives you a reviewer-side rubric with named dimensions, scoring thresholds, and a seniority mapping, so you can turn any portfolio into a defensible advance, screen-with-caution, or pass decision plus a level you can put in front of a hiring manager.

Most guides on this topic coach the designer on building the portfolio. This one sits on the other side of the table. It treats the work as evidence and scores it the way a code reviewer scores a repo: what a signal proves, and what it looks like when it lies.

Why a portfolio score needs a fixed rubric

A portfolio score is a repeatable read of public design work against fixed dimensions, so that two reviewers reach the same verdict. Without a rubric, "read the work" collapses into taste, and taste does not survive being handed to a hiring manager who asks why.

The problem is mechanical, not aesthetic. Reviewers place a candidate in seconds and finish the first pass in under a minute. That time budget caps what any first-pass score can measure. If you try to assess craft nuance in the same breath as role fit, the triage corrupts: the shiniest portfolio wins the slot, and polish is the easiest thing in a portfolio to fake and the least predictive of whether someone solves your actual problems.

So the rubric does two things. It separates the fast triage from the slow read, and it names concrete dimensions instead of adjectives. The dimensions below are what named hiring managers screen on, ordered by how early in the review they apply.

18,906
US profiles matching Product Designer or Senior Product Designer titles
From Refolk's index of professional profiles; the raw pool is large, so the scoring bar is what protects your interview slots.

The five scoring dimensions and what each proves

Score every portfolio on five dimensions: role fit, seniority signal, decision quality, impact delta, and craft. The first two are triage; the middle two are the seniority read; craft comes last and carries the least weight. Each dimension proves something specific, and each has a characteristic way of lying.

DimensionWhat it provesWhat it looks like when it lies
Role fitThe candidate is the kind of designer this role needsA generalist grid that fits any role and none
Seniority signalInferred scope, judgement, ownershipBig titles on small, well-defined work
Decision qualityReal judgement under competing goodsA tidy process diagram nobody believes happened in order
Impact deltaThe work changed somethingFinal mockups with no before or after
CraftExecution ability, one dimension onlyA beautiful personal-brand site with no owned work

Two of these deserve emphasis. Decision quality is what actually gets hired: strong portfolios expose the structure behind decisions, and where decisions are implicit, judgement cannot be evaluated. A missing impact delta is a critical red flag for mid and senior roles, because portfolios without impact collapse into surface-level review. Leslie Yang, a hiring manager who writes on seniority signals, names four indicators reviewers screen fast on: initiative, product thinking, scope, and innovation, adding that seniors must also index high on craft, empathy, storytelling, and leadership.

What the time budget lets you actually measure

The first screen can only assess role fit, a seniority signal, and relevance, because that is all the published time budget buys. Everything else is a second, slower pass reserved for candidates who survive triage.

The figures cluster tightly. Most hiring managers spend under one minute on a portfolio. One recruiter-focused source cites studies showing 80% of recruiters spend three minutes or less. And the placement threshold is brutal: if a reviewer cannot place you in under five seconds, they will not invest further effort. This is why the rubric front-loads role fit and seniority signal; they are the only dimensions that fit inside the window.

ReviewerTime budgetWhat it can assess
Hiring managersunder 1 minrole fit and seniority signal only
Recruiters (80%)under 3 minplus one relevant case skim
Placement thresholdunder 5 secinitial "can I place them"

The practical consequence: run the fast filter first, and only pay the five-to-eight-minute case-study reading cost on the survivors. Treat the first screen as a gate, not a grade.

Where portfolios drop out

  1. Placed by role
    5 sec

    can you say what kind of designer this is

  2. Passed triage
    3 min

    role fit, seniority signal, relevance

  3. Case study read
    8 min

    decision quality and impact delta

  4. Scored to a level
    3 min

    advance, screen, or pass

Each stage costs more reviewer time, so the fast filter protects the expensive read.

Reading seniority: scope, ambiguity, autonomy

Map evidence to a level using three axes that recur across leveling frameworks: scope of impact, level of ambiguity handled, and degree of autonomy. These are the axes that make two reviewers agree, because they describe the work rather than rate it.

Scope of impact runs from one task, to one team, to multiple teams, to the whole company. A clear signal of seniority is large scope: owning or co-owning a large feature set or product surface, an app, or a major vertical. Ambiguity runs from well-defined problems handed to the designer, to problems the designer defined themselves. Autonomy runs from close guidance and review to independent operation. Rippling's public design ladder uses concrete competencies in the same spirit: Impact, Influence, Deliver, Communicate, Operate, and Craft.

Seniors prove level through depth, not breadth. They show two or three deep, complex case studies, each demonstrating strategic thinking and cross-functional collaboration: untangling a legacy system, or designing a feature that touches five departments. They show decisions and trade-offs rather than process diagrams, because the places where you chose between two competing goods, or traded scope against time, are where seniority lives. Juniors own smaller, simpler pieces with more guidance; a focused portfolio of three to five strong case studies signals judgement, while 60 screens with no case studies signals its absence.

The seniority read

Deep, few case studiesShallow, many shots
Junior signal
Advance to entry level, screen for growth
Range without depth
Screen with caution, probe for real ownership
Un-scorable on judgement
Route to a case walkthrough before deciding
Senior signal
Advance at level, verify ownership in interview
Narrow scope, well-defined problemsBroad scope, self-defined problems
Depth in one problem space reads senior; breadth without depth reads junior.

The named failure to avoid here is the adjective ladder: rating each level basic, strong, advanced, or exceptional. Two managers cannot agree where basic ends and strong begins, so calibration falls back to instinct. Force scope, ambiguity, and autonomy language into your notes and the disagreement mostly disappears.

The scoring procedure

Run the eight steps below in order. The first is done before you open a portfolio; the rest move from cheap triage to the expensive read, ending in a written verdict a second reviewer could reproduce. Note that sources disagree on one point of order: some managers screen craft before decision quality, but the fast-filter sources put seniority and relevance first, and this procedure follows them.

Score a portfolio in eight passes

  1. Define target level and role fit first
    Before opening any portfolio, write the scope, autonomy, and decision-rights criteria for the open level. You should be able to state in two sentences what changes between the target level and the one below.
  2. Run first-screen triage
    Land on the homepage, read the positioning, and scan for one relevant case study, all under three minutes. Place the person by role, infer a seniority signal, and judge relevance; if you cannot place them in seconds, it ends here.
  3. Read the most relevant case study for decision quality
    Pick the single most relevant case and look for problem framing, trade-offs, and where judgement was exercised. You should be able to name two real decisions and the alternatives rejected.
  4. Score the impact delta
    Look for before and after metrics, or absent metrics, an observable behaviour change with a stated reason metrics were unavailable. Each scored case needs an impact line or a documented reason it lacks one.
  5. Assess scope and ownership
    Distinguish a "made the icons" contribution from owning a surface, and tag each case as owned, co-owned, or contributed. This tag feeds the seniority read.
  6. Check craft on a secondary surface
    Use Dribbble shots or high-fidelity screens for craft only, never as problem-solving evidence. Score craft as one dimension without letting it touch the seniority read.
  7. Handle NDA and research-heavy exceptions
    Look for a stated confidentiality note, a password gate, or an anonymised case, and do not penalise missing metrics when validated-learning outcomes are shown. Score confidential work on scope, role, and outcomes.
  8. Map to level and record the decision
    Place the candidate against scope, ambiguity, and autonomy, then output advance, screen-with-caution, or pass plus a level. Write a rationale a second reviewer could reproduce.

Once you have scored a handful of portfolios this way, the harder problem is finding enough qualified people to score. Describing the exact evidence you want in plain English is faster than boolean strings across three sites.

Because Refolk reads public LinkedIn, the open GitHub graph, and the open web, I can front-load the ownership and impact signals into the query itself, so the portfolios you open are already biased toward the ones that will score.

What each surface proves and what it hides

Surface choice predicts what evidence exists before you read a word. Dribbble proves craft only; Behance can carry end-to-end problem-solving; a personal site can carry either or mislead through over-styling.

Dribbble shows isolated shots, which prove craft but not problem-solving. Its format structurally cannot carry the framing, trade-offs, and outcomes that decision quality needs, so a Dribbble-only candidate is un-scorable on judgement regardless of talent. Do not pass them for it; route them to a case walkthrough. Behance is built for full projects, with structured case-study pages that walk through problem, process, and outcomes, so it can carry the decision-quality read. A personal site is the wild card: a highly stylised, experimental personal-brand website often results in poor navigation and a lack of focus on the work itself, and if the portfolio's own user experience is bad, that is an immediate disqualifier from a designer.

A Dribbble-only candidate is not a weak candidate; they are an un-scorable one until you see a case.

Two context facts help you calibrate. Behance was founded in 2006 and acquired by Adobe in December 2012; Dribbble was founded in 2009 and is social and shot-driven by design. Neither origin changes the rule: read shots for craft, read case pages for judgement, and never confuse the two.

How this goes wrong: failure modes and false positives

The most valuable part of any scoring standard is the list of ways it misfires. Below are the failure modes that produce wrong verdicts, split into false positives that advance a weak candidate and false negatives that pass a strong one.

  • Polish read as competence. A beautiful personal-brand site with no owned work reads as senior. Check: can you name two real decisions and their rejected alternatives? If not, it is craft, not judgement.
  • Dribbble shots scored as problem-solving. Shots prove craft only. Treat them as one dimension, never as the seniority read.
  • Double Diamond as evidence. A process diagram shows sequence, not judgement. A tidy linear process is itself a flag, because it hides where trade-offs were actually made.
  • Contribution inflation. "I did research, UI, and testing" claimed in a silo can hide minimal real ownership. Check the team makeup and the candidate's specific role before tagging a case as owned.
  • Fabricated research. Vague usability claims such as "we tested with five users and they liked it" with no specifics or honest gaps signal invented process. Look for concrete numbers, methods, and what did not work.
  • Over-leveling from range. Breadth without depth in a problem space reads as junior in the current market. Do not let a wide grid inflate the level.

The false negatives are subtler and cost you senior talent. NDA and research-heavy work create systematic false-pass errors when you score on metrics alone.

For juniors, impact does not always mean metrics; it can mean validated learning, improved usability, or clarified direction. The test for a confidential case is not "are there numbers" but "is the omission deliberate": look for a stated confidentiality note, a password gate, or an anonymised case. Morgane Peng, a design director who has written on showcasing NDA-protected work, is one of several practitioners who confirm confidential cases can be scored on described scope and outcomes rather than shown assets.

One more calibration point: student projects older than two years signal inexperience, not range, and a linear process presented as a full case study is a flag in its own right. The Fountain Institute's write-up on what gets senior product designers hired names both.

Turning the score into a verdict and a level

Convert the five dimension scores into one of three verdicts plus a level, and write the rationale in scope, ambiguity, and autonomy language so a second reviewer can reproduce it. The verdict is a routing decision; the level is a hypothesis to test in the interview.

Advance means the case studies show owned scope, real decisions with rejected alternatives, and an impact delta or a documented reason it is absent. Screen-with-caution means the evidence is promising but one core dimension is unverifiable from public work, most often ownership or impact, so the interview must confirm it. Pass means a reliable red flag fired: only final mockups, a broken portfolio experience, or a case that is process diagrams with no visible judgement.

Portfolio scoring record
Candidate:
Target level and role:
Role fit (fits this role / generalist / mismatch):
Most relevant case study:
  Two real decisions + rejected alternatives:
  Impact delta (metric / behaviour change / documented reason absent):
  Ownership (owned / co-owned / contributed):
Scope of impact (task / team / multi-team / company):
Ambiguity handled (well-defined / self-defined):
Autonomy (guided / independent):
Craft (secondary surface only):
NDA or research note present? (yes / no):
Verdict: advance / screen-with-caution / pass
Proposed level + one-line rationale:

Fill one per candidate. Keep the rationale in scope, ambiguity, and autonomy language so a second reviewer can reproduce the verdict.

The seniority read is a claim, so hold it loosely. Depth in one problem space is the strongest public signal of senior work, and breadth without depth is the strongest signal to under-level. When the two conflict, tag it screen-with-caution and let the interview break the tie.

Keeping the rubric calibrated and stocked

A scoring rubric drifts unless you re-calibrate it against outcomes and keep enough qualified candidates flowing to use it on. Two things keep this work current: checking your verdicts against how hires actually performed, and weighting your sourcing toward the scarce competencies your roles need.

Scarcity is measurable and uneven by title. In Refolk's index, 1,048 US Product Designers list Design Systems as a skill, against only 371 US UX Designers. If a role is systems-heavy, weight Product Designer sourcing accordingly, because the UX-titled pool is thin on that competency.

Title (US)Profiles with Design Systems skill
Product Designer1,048
UX Designer371

The same index shows the raw geographic pool: the US Product Designer pool is 3.73 times the UK pool for the same two titles. That ratio is a sourcing input, not a quality signal, and it should shape how wide you cast before the rubric ever runs.

MarketMatching profilesShare of combined pool
United States18,90678.9%
United Kingdom5,06421.1%

Before you call a scoring pass done, run this check.

Before you record the verdict

  • You defined the target level's scope, ambiguity, and autonomy before opening the portfolio
  • You can name two real decisions and the rejected alternatives from at least one case
  • Each scored case is tagged owned, co-owned, or contributed
  • Craft was scored from a shot surface only, never as the seniority read
  • You checked for an NDA note before penalising missing metrics
  • Your rationale uses scope, ambiguity, and autonomy, not basic, strong, or advanced
  • A second reviewer could reproduce the verdict from your written record

Re-calibrate on a cadence: pull the scores you gave against how those hires performed, and adjust which red flags actually predicted a bad hire versus which just felt like one. The dimensions are stable, but the thresholds are yours to earn from your own outcomes.

Questions practitioners ask

How long should I actually spend reviewing a designer's portfolio?

Budget under three minutes for the first screen and reserve deeper reading for candidates who survive it. Published figures show most hiring managers spend under one minute, and 80% of recruiters spend three minutes or less. The first pass only buys role fit, a seniority signal, and relevance. If the candidate advances, spend five to eight more minutes on the single most relevant case study reading for decision quality.

What are the biggest UX portfolio red flags for a hiring manager?

Only final mockups with no path to the solution, a missing impact delta on mid or senior cases, and the portfolio's own broken navigation, which is an immediate disqualifier. Also watch for a linear Double Diamond presented as a case-study structure, contribution inflation where one person claims research, UI, and testing in a silo, and fabricated research such as vague 'we tested with five users and they liked it' claims with no specifics.

How do I judge a designer's seniority from their portfolio?

Read for scope, ambiguity handled, and autonomy rather than adjectives. Seniors own or co-own a large surface, show two or three deep case studies with real trade-offs, and expose where they chose between competing goods. Juniors own smaller, well-defined pieces with more guidance. Depth in one problem space reads as senior; breadth without depth reads as junior in the current market.

Should I penalise a designer whose work is under NDA?

No. Confidential work legitimately omits client names, specific assets, and sometimes metrics. Score it on scope, role, goals, and outcomes described with generic or hypothetical visuals. A metrics-only rule systematically rejects senior enterprise and fintech designers, who are the pool most likely to be NDA-bound. Look for a stated confidentiality note, a password gate, or an anonymised case as the signal that the omission is deliberate, not empty.

Can I score a designer from Dribbble or Behance alone?

Dribbble proves craft only. Its shot format structurally cannot carry problem-solving, so a Dribbble-only candidate is un-scorable on judgment regardless of talent; route them to a case walkthrough. Behance supports full case-study pages that walk through problem, process, and outcomes, so it can carry decision-quality evidence. Score craft from either, but never read isolated shots as a seniority signal.

Why do two reviewers disagree on the same portfolio?

Usually because the rubric uses an adjective ladder such as basic, strong, advanced, and exceptional. Two managers cannot agree where basic ends and strong begins, so calibration falls back to instinct. Switch to concrete axes: scope of impact, level of ambiguity handled, and degree of autonomy. Those are the axes that recur across leveling frameworks and let two reviewers grade the same portfolio the same way.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next