The Design-Doc and RFC Read: Own It, Scope It, or Pass
After reading, you can find a candidate's public RFCs, ADRs, and design docs, score them across six dimensions, and reach a defensible own-it, scope-it, or pass verdict.
Key takeaways
- Architecture judgment lives in the alternatives, drawbacks, non-goals, and open-questions sections, not in heading completeness that templates fill automatically.
- The alternatives section is the scarcest signal: it is where the Rust RFC process already penalizes authors who are disingenuous about drawbacks, so weak candidates are thinnest there.
- Staff versus senior is a scope distinction, not a code-quality one: single-team blast radius reads senior, cross-team or multi-system blast radius reads staff.
- In Refolk's index there are about 5.5 current US Senior Software Engineers for every current Staff Software Engineer (182,244 to 32,883), so title alone over-selects and an artifact read is what separates true staff scope.
- The US staff pool is roughly 15.1x the UK staff pool (32,883 to 2,174), so UK staff hiring must lean on public RFC and ADR evidence rather than volume filtering.
- A wrong verdict is expensive: the staff-to-senior gap widens from about 30 percent base pay to closer to 80 percent total comp once stock is counted.
You are looking at a senior or staff candidate and you want to know, before the panel, whether they can own architecture decisions at the level you are hiring for. This guide is for engineering managers, technical founders, dev-rel leads, and technical sourcers who need to grade that judgment from the written record. It gives you a scoring rubric and a discovery procedure to read one named person's public design docs, RFCs, and ADRs and reach a defensible verdict: own it, scope it, or pass.
Most writing on this topic explains what an RFC is and hands you a template. That is not the job here. The job is to grade the judgment of one specific person from the artifacts where systems-design reasoning actually lives, and to translate that read into a level call you can defend in a debrief.
Why read design docs instead of code or talks
Design docs, RFCs, and ADRs are the only public artifacts where architecture reasoning is written down in full. A code-review thread shows how someone responds to a change; a conference talk shows how they present a finished idea; a design doc shows the decision being made, with the alternatives, the trade-offs, and the open questions still exposed. That is the surface you want to grade.
Three terms, defined once. A design doc proposes a whole system and argues why one solution best satisfies the goals. An RFC (request for comments) is a design doc put through a public review process before a change is accepted. An ADR (architecture decision record) captures one architecturally significant decision in a short, numbered, durable record.
The supply math is the reason to do this at all. Title filtering over-selects.
With that ratio, a title filter leaves you with a crowded senior pool and a staff pool that is partly mislabeled in both directions: people doing staff-scope work without the title, and people carrying the title on single-team scope. A graded doc read is how you separate them. And the cost of getting it wrong is not a title quibble: the staff-to-senior gap widens from roughly 30 percent base pay to closer to 80 percent total comp once stock is counted. You are calibrating a large number.
The six dimensions that carry the signal
Score every doc on six dimensions: context, goals and non-goals, alternatives considered, decision and rationale, consequences and trade-offs, and open questions. Published templates converge on this spine, which means you can grade a Google design doc, a Nygard ADR, and a Rust RFC against the same rubric even though they name the sections differently.
The table below maps the three canonical formats onto the shared dimensions. Use it to find the right section fast regardless of which format the candidate wrote in.
| Dimension | Google design doc | Nygard ADR | Rust RFC |
|---|---|---|---|
| Context/problem | Context and scope | Context | Motivation |
| Scope boundary | Goals and non-goals | implicit | implicit in motivation |
| Alternatives | Alternatives considered | MADR extension | Rationale and alternatives |
| Decision rationale | The actual design | Decision | Reference-level explanation |
| Trade-offs/costs | Trade-offs | Consequences | Drawbacks |
| Open questions | secondary sections | none | Unresolved questions |
Notice what the table exposes. The base Nygard ADR format - four sections: Status, Context, Decision, Consequences - has no mandatory alternatives or open-questions section at all. MADR adds explicit decision drivers and considered-options-with-pros-and-cons precisely because the base format lets authors skip the reasoning. When you grade an ADR written in the raw Nygard format, you are grading what the author chose to include beyond the minimum.
Not all six dimensions carry equal weight. Context, decision rationale, and trade-offs are necessary but easy to produce once a system exists. The three that actually separate candidates are alternatives considered, non-goals, and open questions. Those are the sections templates cannot auto-fill with substance, and they are where judgment either shows up or does not.
Templates make every heading mandatory. Only the alternatives section makes the author think.
Why alternatives is the load-bearing section
The alternatives-considered section exists to force thinking about what options were available to achieve the goals. Its purpose is to say why the loser lost, specifically enough that a reader a year later can decide whether the reasoning still holds. That is a high bar, and it is where screening pressure already concentrates: the Rust RFC process explicitly notes that RFCs which are disingenuous about the drawbacks or alternatives tend to be poorly received. So a candidate who writes a real alternatives section has already survived a filter. A candidate who writes a hollow one has shown you the gap.
Why non-goals detect staff scope
Non-goals are not negated goals. They are things that could reasonably be goals but are explicitly chosen not to be goals. Anyone can list goals. Choosing what could belong in scope and declining it is the scope-setting act that separates the engineer who defines the problem from the one who executes a defined problem. When a doc names sharp non-goals, you are seeing the judgment that distinguishes staff-level framing from senior-level execution.
How to find one candidate's public docs
Search GitHub for the candidate as a pull-request author, not merely a committer, in RFC repositories and ADR folders. That distinction is the whole discovery problem, because authorship is what you are grading.
RFC repositories follow a predictable shape. The text/ directory contains accepted RFCs numbered sequentially, and accepted Rust RFCs get an id equal to their PR number. The pull requests themselves, both open and closed, represent active proposals and rejected ideas. ADRs are stored in the repository they relate to, numbered sequentially, and never deleted; superseded decisions are marked, not removed, commonly in a doc/adr/ folder.
Authorship is separable because of how these repos work. The RFC file is renamed to the PR number, and edits are made as new commits to the pull request with explanatory comments. So the PR author, the commit history, and the responses in the review thread together distinguish a real author from a co-author or a template-filler. Read the thread: did this person defend the design under hard questions, or did someone else do the arguing?
From public corpus to a graded shortlist of docs
- allDocs where candidate appears
committer or author, any repo
- fewerCandidate is the PR author
commit history confirms authorship
- fewerArchitecturally significant
whole system or one significant decision
- fewestScored on six dimensions
numeric score plus quoted evidence
This is the part that is slow by hand: separating PR authorship from committer noise across repos, then confirming the person defended the design in the thread. If you are starting from a name rather than a repo, Refolk lets you find people by the artifacts themselves rather than by title.
The scoring procedure
Run every candidate through the same seven steps. Steps 1 through 5 are reader or sourcer work on the artifacts; steps 6 and 7 are the hiring manager's level call. Budget roughly two to three hours for a candidate with two or three substantive docs.
Grade one candidate's architecture judgment
- Locate the artifactsSearch GitHub for the candidate as PR author in RFC repos and ADR folders, checking numbered text/ directories and doc/adr/ paths. Done when you have a list of docs with the candidate named as author, not just committer.
- Separate own-authored from co-authoredCheck the PR author, the commit history, and whether the person defended the design in the review thread. Done when each doc is tagged solo, co-authored, or template-fill.
- Confirm architectural significanceVerify each doc captures a real significant decision or proposes a whole system, not a bugfix. Done when trivial docs are discarded.
- Score the six judgment dimensionsGrade context, goals and non-goals, alternatives, decision and rationale, consequences and trade-offs, and open questions. Done when you have a numeric score per dimension with a quoted line of evidence.
- Stress-test alternatives and drawbacksCheck whether rejected options are named with reasons that still hold a year later, and whether real costs are admitted. Done when you have a pass or fail on honesty.
- Map scope against the levelClassify blast radius: single-team reads senior, cross-team or multi-system reads staff. Done when the blast radius is classified against the role.
- Reach the verdictCombine scores, the honesty check, and scope into own architecture, needs scoped direction, or pass. Done when you have a defensible written call with the rubric attached.
Use a fixed rubric so two reviewers reach comparable scores. Here is one you can copy.
Candidate: ______ Doc: ______ Authorship: solo / co-authored / template-fill Date vs implementation: before / after Context/problem [0-3] evidence: "..." Goals and non-goals [0-3] evidence: "..." (non-goals named? Y/N) Alternatives considered[0-3] evidence: "..." (losers named with still-valid reasons? Y/N) Decision and rationale [0-3] evidence: "..." Consequences/trade-offs[0-3] evidence: "..." (real costs admitted, not just benefits? Y/N) Open questions [0-3] evidence: "..." 0 = absent 1 = heading present, no substance 2 = substantive 3 = reasoning a reader could falsify a year later Honesty gate: fails if alternatives OR trade-offs scores 0-1. A failed gate caps the verdict at "needs scoped direction". Blast radius: single-team (senior) / cross-team or multi-system (staff)
Score each dimension 0 to 3. Paste a verbatim quote as evidence for every score above 0.
The 0-to-3 anchors matter. A score of 1 means the heading is present but empty, which is the trap this whole guide exists to defuse. Reserve 3 for reasoning that is specific enough to be falsified later, because that is the standard the alternatives section is supposed to meet.
How this read goes wrong
The failure modes below are where good readers reach bad verdicts. Each one is a specific check, not a vague caution. Treat this as the core of the method.
| Failure mode | What it looks like | The check |
|---|---|---|
| Heading-completeness as judgment | Every section filled, nothing said | Does alternatives name real options with reasons, or restate the decision? |
| Co-authored doc credited to one | Strong doc, weak candidate | Confirm PR author plus who answered hard questions in the thread |
| Retrospective read as foresight | Clean doc, system already shipped | Compare doc date to first commits of the system |
| Only selling the decision | All benefits, no costs | Look for admitted costs in consequences or drawbacks |
| Missing alternatives as style | No road-not-taken section | Are rejected options falsifiable a year later, or absent? |
| ADR as architecture overview | One record sprawls across many decisions | One decision per record, or a design dump? |
| Scope mismatch | Flawless single-team doc for a staff role | Classify blast radius against the role you are filling |
| Rotted decision log | ADR says Postgres, code runs DynamoDB | Read status fields and supersession links |
Two of these deserve extra weight because they flip a verdict in opposite directions.
The first is heading-completeness. Google teams start from standard templates pre-filled with section descriptions, so every section being present is expected, not impressive. A reviewer who scores on structure will reward a template-filler as a strong architect. The defense is the honesty gate: if alternatives or trade-offs is empty, the completeness is cosmetic.
The second is the retrospective trap. The clearest Google design docs tend to be retrospective ones summarizing a system already running. A doc written after implementation reads as crisp judgment when it is really hindsight. Check the timestamp against the first commits of the system it describes, and weight pre-implementation reasoning more heavily, because that is the reasoning that was made under uncertainty.
The rotted-log failure is quieter but it costs you a different thing: credibility. A folder of numbered ADRs can describe PostgreSQL-over-MongoDB while the code now runs on DynamoDB. That does not mean the author has bad judgment; it means the record is stale. Read the status fields and supersession links before you treat an old decision as the candidate's current thinking.
Mapping the score to a level
Scope, not code quality, is what separates staff from senior. Senior engineers own deep execution within a single team; staff engineers own technical outcomes spanning multiple teams, systems, or business domains. Will Larson's framing is worth holding in mind: senior is the career level at most companies, and most companies have no expectation that you go from senior to staff. So a doc that proves excellent single-team execution is a success signal, not a staff signal.
Use the blast radius of the doc as the primary level evidence, and read problem framing as the secondary signal.
| Attribute | Senior signal | Staff signal |
|---|---|---|
| Blast radius | Single team or product | Multiple teams or systems |
| Problem framing | Executes a defined problem | Defines the problem |
| Leveling | L5 / E5 | L6 at Google and Meta |
| Doc evidence | One-system design doc | Cross-cutting RFC or ADR |
The two-variable call is cleaner as a matrix. Judgment quality is one axis; scope is the other. The combination is the verdict.
Verdict from judgment quality and scope
The bottom-right quadrant is the dangerous one. A doc with broad blast radius but hollow alternatives is a candidate claiming staff scope without the reasoning to support it. Do not let the scope flatter the judgment. The honesty gate in the rubric exists to catch exactly this case.
Geography changes how hard you must lean on this read. The US staff pool is large; the UK pool is not.
Where the titled pool is thin, the doc read is not a nice-to-have; it is the primary way to find staff-scope operators who have not been promoted yet.
Keep the read current and defensible
Before you call the job done, verify the read against the checklist below. The goal is a written verdict that survives a hiring debrief: a score, a scope classification, and a quoted line of evidence for each.
Before you file the verdict
- Each doc is confirmed PR-authored by the candidate, not just committed
- The candidate defended the design in the review thread
- Retrospective docs are dated against implementation and weighted down
- Alternatives section names real losers with reasons that still hold
- Consequences or drawbacks admits real costs, not only benefits
- Each ADR captures one decision, not a sprawling overview
- Status and supersession fields checked for rot
- Blast radius classified against the specific role being filled
- Verdict is one of own architecture, needs scoped direction, or pass, with the rubric attached
Two habits keep this method honest over time. First, re-read the artifact, not your notes, when a candidate resurfaces; decision logs rot, and a doc you scored well a year ago may now describe a system that was torn out. Second, calibrate across reviewers by having two people score the same doc independently before comparing, because the heading-completeness trap catches different people differently. If your scores diverge by more than one point on alternatives or non-goals, that is a sign one of you is reading structure where the other is reading substance.
The read is defensible because it is anchored in quoted evidence and a fixed rubric, not in impression. That is what makes it something you can put in front of a panel and something you can defend when the comp number attached to the verdict is large.
Questions practitioners ask
Where do I actually find an engineer's public design docs and RFCs?
Start on GitHub. Rust-style RFC repositories store accepted docs in a numbered text/ directory, and both open and closed pull requests represent active and rejected proposals. ADRs live in the repository they relate to, commonly in a doc/adr/ folder, numbered sequentially and never deleted. Filter by the candidate as pull-request author rather than committer, then read the review thread to confirm they defended the design themselves.
How do I tell a staff-level design doc from a senior-level one?
Look at blast radius and problem framing, not polish. A doc that designs one system for one team is a senior signal; a doc that coordinates multiple teams, systems, or business domains is a staff signal. Senior engineers execute on well-defined problems, while staff engineers often define what problems the team should solve. A flawless single-team doc proves senior, not staff.
Why not just screen on the Staff Software Engineer title?
Because the title under-covers the skill. In Refolk's index there are about 5.5 current US Senior Software Engineers for every current Staff Software Engineer, so many people operating at staff scope have not been promoted. A graded doc read surfaces the staff-scope judgment that a title filter misses, which matters most in smaller markets where the staff pool is thin.
What is the single highest-value section to read in a design doc?
The alternatives-considered section. Its whole point is to say why the losing options lost, specifically enough that a reader a year later can decide whether the reasoning still holds. Templates make every other heading mandatory and often pre-filled, so alternatives is where screening pressure already concentrates and where weak candidates are thinnest. If it is missing or just restates the decision, downgrade the verdict.
How do I avoid crediting a retrospective doc as foresight?
Check dates against implementation. Google's own experience is that the cleanest design docs tend to be retrospective ones summarizing a system that already runs, so a doc written after the fact can read as sharp judgment when it is really hindsight. Compare the doc timestamp to the first commits of the system it describes, and weight pre-implementation reasoning more heavily.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.