# The Design-Doc and RFC Read: Own It, Scope It, or Pass

*After reading, you can find a candidate's public RFCs, ADRs, and design docs, score them across six dimensions, and reach a defensible own-it, scope-it, or pass verdict.*

- Canonical URL: https://www.refolk.ai/guides/design-doc-rfc-architecture-read
- Pillar: Engineering and open source
- Format: Framework
- Published: 2026-10-10
- Last reviewed: 2026-10-10
- Reading time: 15 min
- Keywords: evaluate architecture skills from design docs, assess system design judgment from RFCs, ADR evaluation senior engineer, staff vs senior architecture judgment, find engineer public RFCs

## Key takeaways

- Architecture judgment lives in the alternatives, drawbacks, non-goals, and open-questions sections, not in heading completeness that templates fill automatically.
- The alternatives section is the scarcest signal: it is where the Rust RFC process already penalizes authors who are disingenuous about drawbacks, so weak candidates are thinnest there.
- Staff versus senior is a scope distinction, not a code-quality one: single-team blast radius reads senior, cross-team or multi-system blast radius reads staff.
- In Refolk's index there are about 5.5 current US Senior Software Engineers for every current Staff Software Engineer (182,244 to 32,883), so title alone over-selects and an artifact read is what separates true staff scope.
- The US staff pool is roughly 15.1x the UK staff pool (32,883 to 2,174), so UK staff hiring must lean on public RFC and ADR evidence rather than volume filtering.
- A wrong verdict is expensive: the staff-to-senior gap widens from about 30 percent base pay to closer to 80 percent total comp once stock is counted.

You are looking at a senior or staff candidate and you want to know, before the panel, whether they can own architecture decisions at the level you are hiring for. This guide is for engineering managers, technical founders, dev-rel leads, and technical sourcers who need to grade that judgment from the written record. It gives you a scoring rubric and a discovery procedure to read one named person's public design docs, RFCs, and ADRs and reach a defensible verdict: own it, scope it, or pass.

Most writing on this topic explains what an RFC is and hands you a template. That is not the job here. The job is to grade the judgment of one specific person from the artifacts where systems-design reasoning actually lives, and to translate that read into a level call you can defend in a debrief.

## Why read design docs instead of code or talks

Design docs, RFCs, and ADRs are the only public artifacts where architecture reasoning is written down in full. A code-review thread shows how someone responds to a change; a conference talk shows how they present a finished idea; a design doc shows the decision being made, with the alternatives, the trade-offs, and the open questions still exposed. That is the surface you want to grade.

Three terms, defined once. A design doc proposes a whole system and argues why one solution best satisfies the goals. An RFC (request for comments) is a design doc put through a public review process before a change is accepted. An ADR (architecture decision record) captures one architecturally significant decision in a short, numbered, durable record.

The supply math is the reason to do this at all. Title filtering over-selects.

**5.5 - US Senior Software Engineers per Staff Software Engineer in Refolk's index**

182,244 current senior titles against 32,883 current staff titles, so the title alone cannot tell you who operates at staff scope.

With that ratio, a title filter leaves you with a crowded senior pool and a staff pool that is partly mislabeled in both directions: people doing staff-scope work without the title, and people carrying the title on single-team scope. A graded doc read is how you separate them. And the cost of getting it wrong is not a title quibble: the staff-to-senior gap widens from roughly 30 percent base pay to closer to 80 percent total comp once stock is counted. You are calibrating a large number.

## The six dimensions that carry the signal

Score every doc on six dimensions: context, goals and non-goals, alternatives considered, decision and rationale, consequences and trade-offs, and open questions. Published templates converge on this spine, which means you can grade a Google design doc, a Nygard ADR, and a Rust RFC against the same rubric even though they name the sections differently.

The table below maps the three canonical formats onto the shared dimensions. Use it to find the right section fast regardless of which format the candidate wrote in.

| Dimension | Google design doc | Nygard ADR | Rust RFC |
| --- | --- | --- | --- |
| Context/problem | Context and scope | Context | Motivation |
| Scope boundary | Goals and non-goals | implicit | implicit in motivation |
| Alternatives | Alternatives considered | MADR extension | Rationale and alternatives |
| Decision rationale | The actual design | Decision | Reference-level explanation |
| Trade-offs/costs | Trade-offs | Consequences | Drawbacks |
| Open questions | secondary sections | none | Unresolved questions |

Notice what the table exposes. The base Nygard ADR format - four sections: Status, Context, Decision, Consequences - has no mandatory alternatives or open-questions section at all. MADR adds explicit decision drivers and considered-options-with-pros-and-cons precisely because the base format lets authors skip the reasoning. When you grade an ADR written in the raw Nygard format, you are grading what the author chose to include beyond the minimum.

Not all six dimensions carry equal weight. Context, decision rationale, and trade-offs are necessary but easy to produce once a system exists. The three that actually separate candidates are alternatives considered, non-goals, and open questions. Those are the sections templates cannot auto-fill with substance, and they are where judgment either shows up or does not.

> Templates make every heading mandatory. Only the alternatives section makes the author think.

### Why alternatives is the load-bearing section

The alternatives-considered section exists to force thinking about what options were available to achieve the goals. Its purpose is to say why the loser lost, specifically enough that a reader a year later can decide whether the reasoning still holds. That is a high bar, and it is where screening pressure already concentrates: the Rust RFC process explicitly notes that RFCs which are disingenuous about the drawbacks or alternatives tend to be poorly received. So a candidate who writes a real alternatives section has already survived a filter. A candidate who writes a hollow one has shown you the gap.

### Why non-goals detect staff scope

Non-goals are not negated goals. They are things that could reasonably be goals but are explicitly chosen not to be goals. Anyone can list goals. Choosing what could belong in scope and declining it is the scope-setting act that separates the engineer who defines the problem from the one who executes a defined problem. When a doc names sharp non-goals, you are seeing the judgment that distinguishes staff-level framing from senior-level execution.

## How to find one candidate's public docs

Search GitHub for the candidate as a pull-request author, not merely a committer, in RFC repositories and ADR folders. That distinction is the whole discovery problem, because authorship is what you are grading.

RFC repositories follow a predictable shape. The text/ directory contains accepted RFCs numbered sequentially, and accepted Rust RFCs get an id equal to their PR number. The pull requests themselves, both open and closed, represent active proposals and rejected ideas. ADRs are stored in the repository they relate to, numbered sequentially, and never deleted; superseded decisions are marked, not removed, commonly in a doc/adr/ folder.

Authorship is separable because of how these repos work. The RFC file is renamed to the PR number, and edits are made as new commits to the pull request with explanatory comments. So the PR author, the commit history, and the responses in the review thread together distinguish a real author from a co-author or a template-filler. Read the thread: did this person defend the design under hard questions, or did someone else do the arguing?

#### From public corpus to a graded shortlist of docs

| Stage | Figure | Note |
| --- | --- | --- |
| Docs where candidate appears | all | committer or author, any repo |
| Candidate is the PR author | fewer | commit history confirms authorship |
| Architecturally significant | fewer | whole system or one significant decision |
| Scored on six dimensions | fewest | numeric score plus quoted evidence |

*Each stage removes docs that cannot carry an authorship or judgment signal.*

This is the part that is slow by hand: separating PR authorship from committer noise across repos, then confirming the person defended the design in the thread. If you are starting from a name rather than a repo, [Refolk](/) lets you find people by the artifacts themselves rather than by title.

Ask me this: `Engineers who authored accepted RFCs in the rust-lang/rfcs repository and now work at US infrastructure startups` - [run the search](https://www.refolk.ai/start?q=Engineers%20who%20authored%20accepted%20RFCs%20in%20the%20rust-lang%2Frfcs%20repository%20and%20now%20work%20at%20US%20infrastructure%20startups).

*Returns named people with authored, merged RFCs in a public corpus, so you start the read from confirmed authorship rather than a self-reported title.*

## The scoring procedure

Run every candidate through the same seven steps. Steps 1 through 5 are reader or sourcer work on the artifacts; steps 6 and 7 are the hiring manager's level call. Budget roughly two to three hours for a candidate with two or three substantive docs.

#### Grade one candidate's architecture judgment

1. **Locate the artifacts** - Search GitHub for the candidate as PR author in RFC repos and ADR folders, checking numbered text/ directories and doc/adr/ paths. Done when you have a list of docs with the candidate named as author, not just committer.
2. **Separate own-authored from co-authored** - Check the PR author, the commit history, and whether the person defended the design in the review thread. Done when each doc is tagged solo, co-authored, or template-fill.
3. **Confirm architectural significance** - Verify each doc captures a real significant decision or proposes a whole system, not a bugfix. Done when trivial docs are discarded.
4. **Score the six judgment dimensions** - Grade context, goals and non-goals, alternatives, decision and rationale, consequences and trade-offs, and open questions. Done when you have a numeric score per dimension with a quoted line of evidence.
5. **Stress-test alternatives and drawbacks** - Check whether rejected options are named with reasons that still hold a year later, and whether real costs are admitted. Done when you have a pass or fail on honesty.
6. **Map scope against the level** - Classify blast radius: single-team reads senior, cross-team or multi-system reads staff. Done when the blast radius is classified against the role.
7. **Reach the verdict** - Combine scores, the honesty check, and scope into own architecture, needs scoped direction, or pass. Done when you have a defensible written call with the rubric attached.

Use a fixed rubric so two reviewers reach comparable scores. Here is one you can copy.

**Six-dimension design-doc scoring rubric**

```
Candidate: ______   Doc: ______   Authorship: solo / co-authored / template-fill   Date vs implementation: before / after

Context/problem        [0-3]  evidence: "..."
Goals and non-goals    [0-3]  evidence: "..."  (non-goals named? Y/N)
Alternatives considered[0-3]  evidence: "..."  (losers named with still-valid reasons? Y/N)
Decision and rationale [0-3]  evidence: "..."
Consequences/trade-offs[0-3]  evidence: "..."  (real costs admitted, not just benefits? Y/N)
Open questions         [0-3]  evidence: "..."

0 = absent  1 = heading present, no substance  2 = substantive  3 = reasoning a reader could falsify a year later
Honesty gate: fails if alternatives OR trade-offs scores 0-1. A failed gate caps the verdict at "needs scoped direction".
Blast radius: single-team (senior) / cross-team or multi-system (staff)
```

*Score each dimension 0 to 3. Paste a verbatim quote as evidence for every score above 0.*

The 0-to-3 anchors matter. A score of 1 means the heading is present but empty, which is the trap this whole guide exists to defuse. Reserve 3 for reasoning that is specific enough to be falsified later, because that is the standard the alternatives section is supposed to meet.

> **Rule:** Score authorship before you score content
>
> Never score a doc's dimensions until you have confirmed the candidate is the PR author and defended it in the review thread. A strong doc with weak authorship is a false positive, and it is the most common one.

## How this read goes wrong

The failure modes below are where good readers reach bad verdicts. Each one is a specific check, not a vague caution. Treat this as the core of the method.

| Failure mode | What it looks like | The check |
| --- | --- | --- |
| Heading-completeness as judgment | Every section filled, nothing said | Does alternatives name real options with reasons, or restate the decision? |
| Co-authored doc credited to one | Strong doc, weak candidate | Confirm PR author plus who answered hard questions in the thread |
| Retrospective read as foresight | Clean doc, system already shipped | Compare doc date to first commits of the system |
| Only selling the decision | All benefits, no costs | Look for admitted costs in consequences or drawbacks |
| Missing alternatives as style | No road-not-taken section | Are rejected options falsifiable a year later, or absent? |
| ADR as architecture overview | One record sprawls across many decisions | One decision per record, or a design dump? |
| Scope mismatch | Flawless single-team doc for a staff role | Classify blast radius against the role you are filling |
| Rotted decision log | ADR says Postgres, code runs DynamoDB | Read status fields and supersession links |

Two of these deserve extra weight because they flip a verdict in opposite directions.

The first is heading-completeness. Google teams start from standard templates pre-filled with section descriptions, so every section being present is expected, not impressive. A reviewer who scores on structure will reward a template-filler as a strong architect. The defense is the honesty gate: if alternatives or trade-offs is empty, the completeness is cosmetic.

> **Watch out:** A full template is the default, not an achievement
>
> Because teams pre-fill templates with section descriptions, heading presence proves nothing about judgment. Grade only what the author added to the alternatives, non-goals, and open-questions sections.

The second is the retrospective trap. The clearest Google design docs tend to be retrospective ones summarizing a system already running. A doc written after implementation reads as crisp judgment when it is really hindsight. Check the timestamp against the first commits of the system it describes, and weight pre-implementation reasoning more heavily, because that is the reasoning that was made under uncertainty.

The rotted-log failure is quieter but it costs you a different thing: credibility. A folder of numbered ADRs can describe PostgreSQL-over-MongoDB while the code now runs on DynamoDB. That does not mean the author has bad judgment; it means the record is stale. Read the status fields and supersession links before you treat an old decision as the candidate's current thinking.

## Mapping the score to a level

Scope, not code quality, is what separates staff from senior. Senior engineers own deep execution within a single team; staff engineers own technical outcomes spanning multiple teams, systems, or business domains. Will Larson's framing is worth holding in mind: senior is the career level at most companies, and most companies have no expectation that you go from senior to staff. So a doc that proves excellent single-team execution is a success signal, not a staff signal.

Use the blast radius of the doc as the primary level evidence, and read problem framing as the secondary signal.

| Attribute | Senior signal | Staff signal |
| --- | --- | --- |
| Blast radius | Single team or product | Multiple teams or systems |
| Problem framing | Executes a defined problem | Defines the problem |
| Leveling | L5 / E5 | L6 at Google and Meta |
| Doc evidence | One-system design doc | Cross-cutting RFC or ADR |

The two-variable call is cleaner as a matrix. Judgment quality is one axis; scope is the other. The combination is the verdict.

#### Verdict from judgment quality and scope

Horizontal axis runs from Narrow scope (single team) to Broad scope (multi-team). Vertical axis runs from Weak judgment (hollow alternatives) to Strong judgment (falsifiable reasoning).

| Quadrant | What it means |
| --- | --- |
| Strong judgment, narrow scope | Own architecture at senior level; probe for staff scope in interview |
| Strong judgment, broad scope | Owns architecture at staff level; advance |
| Weak judgment, narrow scope | Pass, or needs scoped direction with close review |
| Weak judgment, broad scope | Scope claim without judgment to back it; probe hard or pass |

*A high score on a single-team doc is a strong senior hire, not a staff one.*

The bottom-right quadrant is the dangerous one. A doc with broad blast radius but hollow alternatives is a candidate claiming staff scope without the reasoning to support it. Do not let the scope flatter the judgment. The honesty gate in the rubric exists to catch exactly this case.

Geography changes how hard you must lean on this read. The US staff pool is large; the UK pool is not.

**15.1x - How much larger the US Staff Software Engineer pool is than the UK's in Refolk's index**

32,883 US staff titles against 2,174 UK staff titles, so UK staff hiring cannot rely on volume filtering and must lean on artifact evidence.

Where the titled pool is thin, the doc read is not a nice-to-have; it is the primary way to find staff-scope operators who have not been promoted yet.

## Keep the read current and defensible

Before you call the job done, verify the read against the checklist below. The goal is a written verdict that survives a hiring debrief: a score, a scope classification, and a quoted line of evidence for each.

#### Before you file the verdict

- [ ] Each doc is confirmed PR-authored by the candidate, not just committed
- [ ] The candidate defended the design in the review thread
- [ ] Retrospective docs are dated against implementation and weighted down
- [ ] Alternatives section names real losers with reasons that still hold
- [ ] Consequences or drawbacks admits real costs, not only benefits
- [ ] Each ADR captures one decision, not a sprawling overview
- [ ] Status and supersession fields checked for rot
- [ ] Blast radius classified against the specific role being filled
- [ ] Verdict is one of own architecture, needs scoped direction, or pass, with the rubric attached

Two habits keep this method honest over time. First, re-read the artifact, not your notes, when a candidate resurfaces; decision logs rot, and a doc you scored well a year ago may now describe a system that was torn out. Second, calibrate across reviewers by having two people score the same doc independently before comparing, because the heading-completeness trap catches different people differently. If your scores diverge by more than one point on alternatives or non-goals, that is a sign one of you is reading structure where the other is reading substance.

The read is defensible because it is anchored in quoted evidence and a fixed rubric, not in impression. That is what makes it something you can put in front of a panel and something you can defend when the comp number attached to the verdict is large.

## Frequently asked questions

### Where do I actually find an engineer's public design docs and RFCs?

Start on GitHub. Rust-style RFC repositories store accepted docs in a numbered text/ directory, and both open and closed pull requests represent active and rejected proposals. ADRs live in the repository they relate to, commonly in a doc/adr/ folder, numbered sequentially and never deleted. Filter by the candidate as pull-request author rather than committer, then read the review thread to confirm they defended the design themselves.

### How do I tell a staff-level design doc from a senior-level one?

Look at blast radius and problem framing, not polish. A doc that designs one system for one team is a senior signal; a doc that coordinates multiple teams, systems, or business domains is a staff signal. Senior engineers execute on well-defined problems, while staff engineers often define what problems the team should solve. A flawless single-team doc proves senior, not staff.

### Why not just screen on the Staff Software Engineer title?

Because the title under-covers the skill. In Refolk's index there are about 5.5 current US Senior Software Engineers for every current Staff Software Engineer, so many people operating at staff scope have not been promoted. A graded doc read surfaces the staff-scope judgment that a title filter misses, which matters most in smaller markets where the staff pool is thin.

### What is the single highest-value section to read in a design doc?

The alternatives-considered section. Its whole point is to say why the losing options lost, specifically enough that a reader a year later can decide whether the reasoning still holds. Templates make every other heading mandatory and often pre-filled, so alternatives is where screening pressure already concentrates and where weak candidates are thinnest. If it is missing or just restates the decision, downgrade the verdict.

### How do I avoid crediting a retrospective doc as foresight?

Check dates against implementation. Google's own experience is that the cleanest design docs tend to be retrospective ones summarizing a system that already runs, so a doc written after the fact can read as sharp judgment when it is really hindsight. Compare the doc timestamp to the first commits of the system it describes, and weight pre-implementation reasoning more heavily.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/design-doc-rfc-architecture-read*
