Refolk
FrameworkEngineering and open source

The Design-Doc and RFC Read: Own It, Scope It, or Pass

After reading, you can find a candidate's public RFCs, ADRs, and design docs, score them across six dimensions, and reach a defensible own-it, scope-it, or pass verdict.

15 min readLast reviewed October 10, 2026Read as Markdown

Key takeaways

  • Architecture judgment lives in the alternatives, drawbacks, non-goals, and open-questions sections, not in heading completeness that templates fill automatically.
  • The alternatives section is the scarcest signal: it is where the Rust RFC process already penalizes authors who are disingenuous about drawbacks, so weak candidates are thinnest there.
  • Staff versus senior is a scope distinction, not a code-quality one: single-team blast radius reads senior, cross-team or multi-system blast radius reads staff.
  • In Refolk's index there are about 5.5 current US Senior Software Engineers for every current Staff Software Engineer (182,244 to 32,883), so title alone over-selects and an artifact read is what separates true staff scope.
  • The US staff pool is roughly 15.1x the UK staff pool (32,883 to 2,174), so UK staff hiring must lean on public RFC and ADR evidence rather than volume filtering.
  • A wrong verdict is expensive: the staff-to-senior gap widens from about 30 percent base pay to closer to 80 percent total comp once stock is counted.

You are looking at a senior or staff candidate and you want to know, before the panel, whether they can own architecture decisions at the level you are hiring for. This guide is for engineering managers, technical founders, dev-rel leads, and technical sourcers who need to grade that judgment from the written record. It gives you a scoring rubric and a discovery procedure to read one named person's public design docs, RFCs, and ADRs and reach a defensible verdict: own it, scope it, or pass.

Most writing on this topic explains what an RFC is and hands you a template. That is not the job here. The job is to grade the judgment of one specific person from the artifacts where systems-design reasoning actually lives, and to translate that read into a level call you can defend in a debrief.

Why read design docs instead of code or talks

Design docs, RFCs, and ADRs are the only public artifacts where architecture reasoning is written down in full. A code-review thread shows how someone responds to a change; a conference talk shows how they present a finished idea; a design doc shows the decision being made, with the alternatives, the trade-offs, and the open questions still exposed. That is the surface you want to grade.

Three terms, defined once. A design doc proposes a whole system and argues why one solution best satisfies the goals. An RFC (request for comments) is a design doc put through a public review process before a change is accepted. An ADR (architecture decision record) captures one architecturally significant decision in a short, numbered, durable record.

The supply math is the reason to do this at all. Title filtering over-selects.

5.5
US Senior Software Engineers per Staff Software Engineer in Refolk's index
182,244 current senior titles against 32,883 current staff titles, so the title alone cannot tell you who operates at staff scope.

With that ratio, a title filter leaves you with a crowded senior pool and a staff pool that is partly mislabeled in both directions: people doing staff-scope work without the title, and people carrying the title on single-team scope. A graded doc read is how you separate them. And the cost of getting it wrong is not a title quibble: the staff-to-senior gap widens from roughly 30 percent base pay to closer to 80 percent total comp once stock is counted. You are calibrating a large number.

The six dimensions that carry the signal

Score every doc on six dimensions: context, goals and non-goals, alternatives considered, decision and rationale, consequences and trade-offs, and open questions. Published templates converge on this spine, which means you can grade a Google design doc, a Nygard ADR, and a Rust RFC against the same rubric even though they name the sections differently.

The table below maps the three canonical formats onto the shared dimensions. Use it to find the right section fast regardless of which format the candidate wrote in.

DimensionGoogle design docNygard ADRRust RFC
Context/problemContext and scopeContextMotivation
Scope boundaryGoals and non-goalsimplicitimplicit in motivation
AlternativesAlternatives consideredMADR extensionRationale and alternatives
Decision rationaleThe actual designDecisionReference-level explanation
Trade-offs/costsTrade-offsConsequencesDrawbacks
Open questionssecondary sectionsnoneUnresolved questions

Notice what the table exposes. The base Nygard ADR format - four sections: Status, Context, Decision, Consequences - has no mandatory alternatives or open-questions section at all. MADR adds explicit decision drivers and considered-options-with-pros-and-cons precisely because the base format lets authors skip the reasoning. When you grade an ADR written in the raw Nygard format, you are grading what the author chose to include beyond the minimum.

Not all six dimensions carry equal weight. Context, decision rationale, and trade-offs are necessary but easy to produce once a system exists. The three that actually separate candidates are alternatives considered, non-goals, and open questions. Those are the sections templates cannot auto-fill with substance, and they are where judgment either shows up or does not.

Templates make every heading mandatory. Only the alternatives section makes the author think.

Why alternatives is the load-bearing section

The alternatives-considered section exists to force thinking about what options were available to achieve the goals. Its purpose is to say why the loser lost, specifically enough that a reader a year later can decide whether the reasoning still holds. That is a high bar, and it is where screening pressure already concentrates: the Rust RFC process explicitly notes that RFCs which are disingenuous about the drawbacks or alternatives tend to be poorly received. So a candidate who writes a real alternatives section has already survived a filter. A candidate who writes a hollow one has shown you the gap.

Why non-goals detect staff scope

Non-goals are not negated goals. They are things that could reasonably be goals but are explicitly chosen not to be goals. Anyone can list goals. Choosing what could belong in scope and declining it is the scope-setting act that separates the engineer who defines the problem from the one who executes a defined problem. When a doc names sharp non-goals, you are seeing the judgment that distinguishes staff-level framing from senior-level execution.

How to find one candidate's public docs

Search GitHub for the candidate as a pull-request author, not merely a committer, in RFC repositories and ADR folders. That distinction is the whole discovery problem, because authorship is what you are grading.

RFC repositories follow a predictable shape. The text/ directory contains accepted RFCs numbered sequentially, and accepted Rust RFCs get an id equal to their PR number. The pull requests themselves, both open and closed, represent active proposals and rejected ideas. ADRs are stored in the repository they relate to, numbered sequentially, and never deleted; superseded decisions are marked, not removed, commonly in a doc/adr/ folder.

Authorship is separable because of how these repos work. The RFC file is renamed to the PR number, and edits are made as new commits to the pull request with explanatory comments. So the PR author, the commit history, and the responses in the review thread together distinguish a real author from a co-author or a template-filler. Read the thread: did this person defend the design under hard questions, or did someone else do the arguing?

From public corpus to a graded shortlist of docs

  1. Docs where candidate appears
    all

    committer or author, any repo

  2. Candidate is the PR author
    fewer

    commit history confirms authorship

  3. Architecturally significant
    fewer

    whole system or one significant decision

  4. Scored on six dimensions
    fewest

    numeric score plus quoted evidence

Each stage removes docs that cannot carry an authorship or judgment signal.

This is the part that is slow by hand: separating PR authorship from committer noise across repos, then confirming the person defended the design in the thread. If you are starting from a name rather than a repo, Refolk lets you find people by the artifacts themselves rather than by title.

The scoring procedure

Run every candidate through the same seven steps. Steps 1 through 5 are reader or sourcer work on the artifacts; steps 6 and 7 are the hiring manager's level call. Budget roughly two to three hours for a candidate with two or three substantive docs.

Grade one candidate's architecture judgment

  1. Locate the artifacts
    Search GitHub for the candidate as PR author in RFC repos and ADR folders, checking numbered text/ directories and doc/adr/ paths. Done when you have a list of docs with the candidate named as author, not just committer.
  2. Separate own-authored from co-authored
    Check the PR author, the commit history, and whether the person defended the design in the review thread. Done when each doc is tagged solo, co-authored, or template-fill.
  3. Confirm architectural significance
    Verify each doc captures a real significant decision or proposes a whole system, not a bugfix. Done when trivial docs are discarded.
  4. Score the six judgment dimensions
    Grade context, goals and non-goals, alternatives, decision and rationale, consequences and trade-offs, and open questions. Done when you have a numeric score per dimension with a quoted line of evidence.
  5. Stress-test alternatives and drawbacks
    Check whether rejected options are named with reasons that still hold a year later, and whether real costs are admitted. Done when you have a pass or fail on honesty.
  6. Map scope against the level
    Classify blast radius: single-team reads senior, cross-team or multi-system reads staff. Done when the blast radius is classified against the role.
  7. Reach the verdict
    Combine scores, the honesty check, and scope into own architecture, needs scoped direction, or pass. Done when you have a defensible written call with the rubric attached.

Use a fixed rubric so two reviewers reach comparable scores. Here is one you can copy.

Six-dimension design-doc scoring rubric
Candidate: ______   Doc: ______   Authorship: solo / co-authored / template-fill   Date vs implementation: before / after

Context/problem        [0-3]  evidence: "..."
Goals and non-goals    [0-3]  evidence: "..."  (non-goals named? Y/N)
Alternatives considered[0-3]  evidence: "..."  (losers named with still-valid reasons? Y/N)
Decision and rationale [0-3]  evidence: "..."
Consequences/trade-offs[0-3]  evidence: "..."  (real costs admitted, not just benefits? Y/N)
Open questions         [0-3]  evidence: "..."

0 = absent  1 = heading present, no substance  2 = substantive  3 = reasoning a reader could falsify a year later
Honesty gate: fails if alternatives OR trade-offs scores 0-1. A failed gate caps the verdict at "needs scoped direction".
Blast radius: single-team (senior) / cross-team or multi-system (staff)

Score each dimension 0 to 3. Paste a verbatim quote as evidence for every score above 0.

The 0-to-3 anchors matter. A score of 1 means the heading is present but empty, which is the trap this whole guide exists to defuse. Reserve 3 for reasoning that is specific enough to be falsified later, because that is the standard the alternatives section is supposed to meet.

How this read goes wrong

The failure modes below are where good readers reach bad verdicts. Each one is a specific check, not a vague caution. Treat this as the core of the method.

Failure modeWhat it looks likeThe check
Heading-completeness as judgmentEvery section filled, nothing saidDoes alternatives name real options with reasons, or restate the decision?
Co-authored doc credited to oneStrong doc, weak candidateConfirm PR author plus who answered hard questions in the thread
Retrospective read as foresightClean doc, system already shippedCompare doc date to first commits of the system
Only selling the decisionAll benefits, no costsLook for admitted costs in consequences or drawbacks
Missing alternatives as styleNo road-not-taken sectionAre rejected options falsifiable a year later, or absent?
ADR as architecture overviewOne record sprawls across many decisionsOne decision per record, or a design dump?
Scope mismatchFlawless single-team doc for a staff roleClassify blast radius against the role you are filling
Rotted decision logADR says Postgres, code runs DynamoDBRead status fields and supersession links

Two of these deserve extra weight because they flip a verdict in opposite directions.

The first is heading-completeness. Google teams start from standard templates pre-filled with section descriptions, so every section being present is expected, not impressive. A reviewer who scores on structure will reward a template-filler as a strong architect. The defense is the honesty gate: if alternatives or trade-offs is empty, the completeness is cosmetic.

The second is the retrospective trap. The clearest Google design docs tend to be retrospective ones summarizing a system already running. A doc written after implementation reads as crisp judgment when it is really hindsight. Check the timestamp against the first commits of the system it describes, and weight pre-implementation reasoning more heavily, because that is the reasoning that was made under uncertainty.

The rotted-log failure is quieter but it costs you a different thing: credibility. A folder of numbered ADRs can describe PostgreSQL-over-MongoDB while the code now runs on DynamoDB. That does not mean the author has bad judgment; it means the record is stale. Read the status fields and supersession links before you treat an old decision as the candidate's current thinking.

Mapping the score to a level

Scope, not code quality, is what separates staff from senior. Senior engineers own deep execution within a single team; staff engineers own technical outcomes spanning multiple teams, systems, or business domains. Will Larson's framing is worth holding in mind: senior is the career level at most companies, and most companies have no expectation that you go from senior to staff. So a doc that proves excellent single-team execution is a success signal, not a staff signal.

Use the blast radius of the doc as the primary level evidence, and read problem framing as the secondary signal.

AttributeSenior signalStaff signal
Blast radiusSingle team or productMultiple teams or systems
Problem framingExecutes a defined problemDefines the problem
LevelingL5 / E5L6 at Google and Meta
Doc evidenceOne-system design docCross-cutting RFC or ADR

The two-variable call is cleaner as a matrix. Judgment quality is one axis; scope is the other. The combination is the verdict.

Verdict from judgment quality and scope

Strong judgment (falsifiable reasoning)Weak judgment (hollow alternatives)
Strong judgment, narrow scope
Own architecture at senior level; probe for staff scope in interview
Strong judgment, broad scope
Owns architecture at staff level; advance
Weak judgment, narrow scope
Pass, or needs scoped direction with close review
Weak judgment, broad scope
Scope claim without judgment to back it; probe hard or pass
Narrow scope (single team)Broad scope (multi-team)
A high score on a single-team doc is a strong senior hire, not a staff one.

The bottom-right quadrant is the dangerous one. A doc with broad blast radius but hollow alternatives is a candidate claiming staff scope without the reasoning to support it. Do not let the scope flatter the judgment. The honesty gate in the rubric exists to catch exactly this case.

Geography changes how hard you must lean on this read. The US staff pool is large; the UK pool is not.

15.1x
How much larger the US Staff Software Engineer pool is than the UK's in Refolk's index
32,883 US staff titles against 2,174 UK staff titles, so UK staff hiring cannot rely on volume filtering and must lean on artifact evidence.

Where the titled pool is thin, the doc read is not a nice-to-have; it is the primary way to find staff-scope operators who have not been promoted yet.

Keep the read current and defensible

Before you call the job done, verify the read against the checklist below. The goal is a written verdict that survives a hiring debrief: a score, a scope classification, and a quoted line of evidence for each.

Before you file the verdict

  • Each doc is confirmed PR-authored by the candidate, not just committed
  • The candidate defended the design in the review thread
  • Retrospective docs are dated against implementation and weighted down
  • Alternatives section names real losers with reasons that still hold
  • Consequences or drawbacks admits real costs, not only benefits
  • Each ADR captures one decision, not a sprawling overview
  • Status and supersession fields checked for rot
  • Blast radius classified against the specific role being filled
  • Verdict is one of own architecture, needs scoped direction, or pass, with the rubric attached

Two habits keep this method honest over time. First, re-read the artifact, not your notes, when a candidate resurfaces; decision logs rot, and a doc you scored well a year ago may now describe a system that was torn out. Second, calibrate across reviewers by having two people score the same doc independently before comparing, because the heading-completeness trap catches different people differently. If your scores diverge by more than one point on alternatives or non-goals, that is a sign one of you is reading structure where the other is reading substance.

The read is defensible because it is anchored in quoted evidence and a fixed rubric, not in impression. That is what makes it something you can put in front of a panel and something you can defend when the comp number attached to the verdict is large.

Questions practitioners ask

Where do I actually find an engineer's public design docs and RFCs?

Start on GitHub. Rust-style RFC repositories store accepted docs in a numbered text/ directory, and both open and closed pull requests represent active and rejected proposals. ADRs live in the repository they relate to, commonly in a doc/adr/ folder, numbered sequentially and never deleted. Filter by the candidate as pull-request author rather than committer, then read the review thread to confirm they defended the design themselves.

How do I tell a staff-level design doc from a senior-level one?

Look at blast radius and problem framing, not polish. A doc that designs one system for one team is a senior signal; a doc that coordinates multiple teams, systems, or business domains is a staff signal. Senior engineers execute on well-defined problems, while staff engineers often define what problems the team should solve. A flawless single-team doc proves senior, not staff.

Why not just screen on the Staff Software Engineer title?

Because the title under-covers the skill. In Refolk's index there are about 5.5 current US Senior Software Engineers for every current Staff Software Engineer, so many people operating at staff scope have not been promoted. A graded doc read surfaces the staff-scope judgment that a title filter misses, which matters most in smaller markets where the staff pool is thin.

What is the single highest-value section to read in a design doc?

The alternatives-considered section. Its whole point is to say why the losing options lost, specifically enough that a reader a year later can decide whether the reasoning still holds. Templates make every other heading mandatory and often pre-filled, so alternatives is where screening pressure already concentrates and where weak candidates are thinnest. If it is missing or just restates the decision, downgrade the verdict.

How do I avoid crediting a retrospective doc as foresight?

Check dates against implementation. Google's own experience is that the cleanest design docs tend to be retrospective ones summarizing a system that already runs, so a doc written after the fact can read as sharp judgment when it is really hindsight. Compare the doc timestamp to the first commits of the system it describes, and weight pre-implementation reasoning more heavily.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next