Refolk
TeardownEngineering and open source

Reproducing a Repo's Headline Benchmark Before You Trust It

You can take one repo's published speedup claim, attempt to reproduce it end to end on your own hardware, and return a defensible verdict with evidence.

18 min readLast reviewed October 2, 2026Read as Markdown

Key takeaways

  • In the documented repo cases, the headline claim died at provenance, not at measurement - the harness was absent, the import was broken, or the numbers came from ad-hoc runs with no re-measure path.
  • Google Benchmark defaults to 1 repetition with no dispersion, while Criterion.rs defaults to 100 samples with bootstrap confidence intervals, so the same claim reads as proven or noisy depending only on which harness the author chose.
  • Cross-machine speedups are not portable: one speculative-decoding config that delivered 3.5x on NVIDIA H100s delivered only 1.1x on AMD MI300X with no code change, which is why the same-hardware control is the deciding test.
  • Re-measuring one tool on the real product path turned a reported 5.6x slowdown into 13 ms versus ripgrep's 16.7 ms - the gap was the harness, not the algorithm.
  • An undated multiple is structurally unreproducible: a kernel scoring 2x over last quarter's reference can score 1.4x against a faster successor with no change to the kernel.
  • In Refolk's index, Performance Engineer resolves to 1,119 current US profiles versus 84 in Germany, but the title is polluted by motorsport roles, so sourcing on title alone over-counts the people who can actually reproduce a software benchmark.

You found a library whose README claims a large speedup, and before you build on it you need to reproduce that number yourself. This guide is for engineers who have to return a defensible answer, not a feeling. It carries one performance claim all the way through a reproduction attempt - the clone, the missing harness, the version pin, the mismatched hardware, the raw runs, and the wrong turns - and ends with a three-way verdict you can attach evidence to.

Most guides on this topic stop at a checklist and tell you benchmarks "can be misleading" in the abstract. That is true and useless. What you need is to know the exact forks where a claim falls apart, and what the surviving number actually proves once it does. The short version, borne out by every documented case below: the claim usually dies at provenance, long before you ever measure anything.

What a reproducible benchmark actually requires

A benchmark is reproducible when a second engineer can rebuild your result from recorded facts, not when it is merely fast. Multiple current acceptance-criteria specs converge on the same minimum list, so this is not a matter of taste.

To call a performance number reproducible, you need, at minimum, to record exact software versions, the git SHA, and the hardware environment; define your timing boundaries (setup versus solve versus transfer); specify warmup, repetition count, and seed policy; and report dispersion, not just a single figure. A fuller protocol adds dataset split and sample IDs, the model or checkpoint and runtime identity, batch and concurrency, sequence lengths, the memory-measurement method, and preprocessing overhead. Gernot Heiser's canonical list of "systems benchmarking crimes" names two offenses that matter here directly: missing platform specification, and reporting relative numbers only.

Translate that into two questions you ask of any README claim:

  • Can I regenerate it? Is there a committed script that produces this number, with the versions and rig it was produced on?
  • Can I compare it? Is the number absolute and bounded, or is it a bare multiple with no baseline, no hardware, and no date?

If the answer to either is no, you already have most of your verdict.

The claim I am going to reproduce

The worked example here is a real pattern, not a toy. The repo I will carry through ships a README that cites a GPU speedup, specifically numbers measured on an RTX 4090. That is the claim. The job is to land a verdict of holds, hardware-inflated, or unreproducible, and to leave behind artifacts a colleague could rerun.

Those three verdict categories are this guide's framework, not a named public standard, so treat the labels as load-bearing editorial. What is documented is what each verdict needs attached to it: raw result artifacts with hashes, normalized JSON or CSV, a reproduction command, median and p95, toolchain, OS, machine, and commit. The verdict is my judgment; the evidence is not negotiable.

Where reproduction attempts fall out

  1. Claim located in README
    100%

    the headline number

  2. Harness and results exist in-tree
    fewer

    many die here

  3. Builds from pinned commit
    fewer

    version and import breaks

  4. Reproduces near author rig
    fewer

    the sanity pass

  5. Survives same-hardware control
    fewest

    the real algorithmic gain

Most claims never reach measurement, because the harness is absent or the provenance is ad-hoc.

Step one kills most claims: provenance first

Before you install anything, find the script that produced the number. In the documented repo failures, the breakage was absence, not a wrong algorithm, which means your first twenty minutes decide most verdicts.

Three real cases make the point. In one repo (kvcompress #13), the README cited specific GPU numbers, but the only committed result files were labelled CPU-only and measured on macOS, and the reproduction script imported a module that did not exist. The claim was not wrong so much as undefended - there was no artefact in the tree that produced the headline figure, and the one script that pretended to was broken. In another (LogosLang #73), the headline speedups "came from ad-hoc runs": no benches folder, no cargo bench target, no script that re-measures them. In a third (ThinkRL #134), the benchmark entry point pointed at a module that did not exist.

Here is the concrete provenance pass for my RTX 4090 claim:

  1. grep -r "4090" and grep the exact numeric figure across the tree. If the figure appears only in README.md and nowhere in a results file or a script, that is the first red flag.
  2. Find the results directory. Open the files. Read the labels. In the kvcompress pattern, the committed results said CPU-only while the README said GPU - a README-versus-results mismatch you catch only by reading both.
  3. Find the harness. Look for a benches/ folder, a cargo bench target, a scripts/bench.sh, or a Makefile target. If none exists, the claim has no re-measure path and you are already at a verdict.
Repo (issue)Failure modeHow it surfaced
kvcompress #13Claimed GPU numbers, only CPU results in-tree, broken scriptREADME vs results/ mismatch
LogosLang #73Ad-hoc numbers, no benches targetNo re-measure path
aside-codemode #42Harness artifact inflated the gapRe-measure flipped the verdict

For my example, assume provenance survives: there is a benches/ target, a committed results file whose labels match the README, and a documented command. That is the minority case, and it is the only case where the rest of this guide applies.

Pin the build to the exact commit

Pin to a full commit SHA, not a version tag. A full SHA guarantees the exact code is used even if the tag or the code is moved later, which is the whole point of reproduction.

Resolve the README's claimed version to a full SHA, clone, and check out that worktree. Then confirm git rev-parse HEAD matches the pinned SHA and write down the versions of everything you installed in a small pin manifest. CI, if you have it, should fail when the manifest disagrees with what is actually resolved.

The subtle trap is that there can be two commits, not one. The benchmark suite and the runner that drives it may live at different revisions. In one documented case (graalvm #9972), between a pinned suite commit and a runner commit one week later, the analysis code "moved substantially." If you record only one SHA, you cannot tell which code produced the number. So record both benchmarkSuiteCommit and runnerCommit, and make every execution build its source worktree from that exact commit, never from the launcher's current HEAD.

Pin manifest (fill before any run)
claim_source: README.md line, quoted verbatim
benchmarkSuiteCommit: <full 40-char SHA>
runnerCommit: <full 40-char SHA, or "same as suite">
library_version: <tag> -> <resolved SHA>
baseline_version: <name> <version> -> <resolved SHA or package digest>
dataset: <name> :: <split/sample IDs> :: <digest>
seeds: corpus=<n> native=<n> provider=<n>
toolchain: <compiler/runtime + version>
os: <distro + kernel>
cpu: <model> | gpu: <model>

Keep this next to your raw results. One row per input that can change.

Note the separate seed fields. A reproducibility spec (sts2-harness #125) requires typed fields for corpus, native, and provider seeds and tells you to never substitute one for another. Collapsing them into one "seed=42" is how two honest engineers get different numbers and blame each other.

Reconstruct the harness, then do a sanity pass

Rebuild the environment from the manifest until the benchmark runs one clean trial, then run it on hardware close to the README's rig before you trust your own. Reproducing the author's number first tells you whether your setup is sane; skipping it means every later deviation is ambiguous.

For the RTX 4090 claim, this means finding a 4090 - a cloud instance will do - and running the committed harness there first. One of two things happens. Either you land inside the claimed band, and you have a working baseline to carry to your own hardware, or you do not, and you now have a concrete deviation to explain rather than a vague "I couldn't get it to work." Both are progress. The deviation is data.

This is also where you confirm the measured unit is the product path and not startup. In the aside-codemode #42 case, a tool reported "5.6x slower," but re-measuring on the real path turned it into 13 ms versus ripgrep's 16.7 ms. The reported gap was the harness: process startup and transport counted inside the unit of work. Before you report any ratio, check that the thing being timed is the thing you would actually run.

Faster frequently means the harness, not the code. Align the measured unit to the path you will ship.

Run the measured protocol, and know your harness's defaults

Run with warmup discarded, multiple repetitions, fixtures and IO excluded from timing, and record every raw run so you end with a distribution and dispersion rather than a single number. The default settings of your benchmarking tool decide, on their own, whether a claim looks proven or noisy.

This is the quiet trap. The same claim, measured honestly, can read either way depending only on which harness the author reached for. Criterion.rs ships at 100 samples with 3 seconds of warm-up, 5 seconds of measurement, 100,000 bootstrap resamples, and 0.95 confidence, and it reports confidence intervals and standard deviation by default. Google Benchmark ships with repetitions defaulting to 1, which means no dispersion at all unless you explicitly ask for more.

HarnessDefault repetitions/samplesDefault warmupDispersion by default
Criterion.rs100 samples3 sYes (CI, std dev, median)
Google Benchmark1 repetitionnone (iteration estimation only)No (needs flag)
1
Default repetitions in Google Benchmark
One run reports no dispersion, so a single fast result inside normal variance reads as a win.

Practitioner specs put the floor at five repetitions and insist you report p50, p95, and p99, not a mean. So if the README's number came from Google Benchmark at defaults, the author never had dispersion to begin with, and "reproducing" a single run tells you nothing. Set repetitions to at least five, discard warmup, and record the raw values. For my 4090 claim I run the harness with explicit repetitions on both the author's rig and my own, and I keep every raw number in a file I will later hash.

Once you have a distribution on your hardware and the author's, you have done the honest part. Finding the people who do this work well, rather than the people whose title merely says they do, is its own problem - and one worth solving before you staff a reproduction effort.

The same-hardware control is the deciding test

To separate algorithmic speedup from hardware, put the library and the baseline it claims to beat on one machine and measure the ratio there. A cross-machine table in a README proves nothing about the algorithm, because performance is strongly dependent on factors unrelated to the code.

The documented control is explicit: include at least one same-hardware comparison to separate algorithmic from hardware effects. The deep-learning efficiency literature says the same thing in stronger terms - ensure the methods being compared use identical hardware and software to the greatest extent possible, because timing is strongly dependent on hardware and software factors that have nothing to do with the algorithm. A 2026 replay study fixed 64 vCPUs and 256GB of memory across four cloud profiles and varied only the processor, precisely so the processor was the only free variable. The GSO benchmark sidesteps absolute speedups entirely by comparing generated optimizations against expert implementations in the same environment.

Why this is the deciding test, not a nicety: the same config does not travel. One speculative-decoding setup delivered 3.5x on NVIDIA H100s and only 1.1x on AMD MI300X with no code change. A README that shows a big number on one accelerator and a baseline on another is measuring the gap between two machines and calling it an algorithm.

Reading a speedup claim

High gain on your hardwareLow gain on your hardware
Hardware-inflated, low value
grade hardware-inflated; the gain was the machine
Hardware-inflated but still fast
grade hardware-inflated; re-benchmark on your target before adopting
Real but marginal
grade holds; decide if a small algorithmic win is worth the dependency
Real and large
grade holds; adopt, and archive the control as proof
Fails same-hardware controlSurvives same-hardware control
Place your result by whether the number survives a same-hardware control and whether it holds on your rig.

For the RTX 4090 claim, the control is concrete: run the library and its stated baseline on the same single GPU, measure the ratio, and compare that to the README's cross-machine figure. If the same-hardware ratio collapses toward 1.0 while the README showed a large number, the gain lived in the hardware comparison, and the verdict is hardware-inflated regardless of how fast the absolute number looks.

Watch the quieter hardware confounds too. On virtualized hosts, being allocated 2 vCPUs says nothing about how fast those cores tick, so pin the CPU governor, record the scaling driver, and confirm your vCPU entitlement is not the variable moving your numbers.

The reproduction procedure end to end

Here is the whole attempt as an ordered procedure. Each step has a done condition so you know when to move on, and the early steps are the ones that end most attempts.

Reproduce a headline benchmark, clone to verdict

  1. Locate the claim and its provenance
    Read the exact README line and confirm whether a harness, raw results, and the rig exist in-tree. Done when you can point to the script that produced the number or confirm it is absent.
  2. Pin the build to a commit
    Resolve the version to a full commit SHA, clone, and check out that worktree; record benchmarkSuiteCommit and runnerCommit separately if they differ. Done when git rev-parse HEAD matches the pinned SHA and a pin manifest exists.
  3. Reconstruct the harness and environment
    Install exact dependency versions, seeds, and dataset, and record OS, CPU, toolchain, and git SHA. Done when the benchmark builds and runs one trial without error.
  4. Reproduce the author's rig numbers first
    Run on hardware closest to the README's and confirm you land near the published figure. Done when you are within the claimed band or have a concrete deviation to explain.
  5. Run the measured protocol on your hardware
    Discard warmup, run at least five repetitions, exclude fixtures and IO from timing, and record every raw run. Done when you have a distribution with dispersion, not a single number.
  6. Run the same-hardware control
    Put the library and its baseline on one machine to isolate algorithmic gain. Done when you have a speedup ratio measured on a single rig.
  7. Grade and attach evidence
    Compute p50 and p99 with dispersion, assign holds, hardware-inflated, or unreproducible, and archive raw artifacts with hashes plus the reproduction command. Done when a second engineer could rerun from your artifacts.

How this goes wrong: the failure modes

Most false verdicts come from trusting an artifact that looks authoritative but is not. Each failure mode below has a specific check and a documented case, because a signal is only useful when you also know what it looks like when it lies.

  • Harness absent. Lies as: a results folder that looks authoritative but holds synthetic or CPU-only numbers under a GPU README. Check: grep for the exact figure in-tree and confirm a committed script regenerates it (kvcompress #13).
  • Ad-hoc provenance. Lies as: a plausible number with no re-run path, so a future regression is invisible. Check: is there a bench target or a scripts/bench.sh; if not, the claim is unverifiable (LogosLang #73).
  • Wrong commit measured. Lies as: you reproduce a number from current HEAD and assume it matches the release. Check: build the worktree from the pinned SHA and record runner versus suite commit separately (graalvm #9972).
  • Single-run trust. Lies as: one fast run inside normal variance reading as a win. Check: require repetitions and dispersion; Google Benchmark's default of 1 gives you none.
  • Cross-hardware inflation. Lies as: a README's cross-machine table reading as algorithmic gain. Check: run the same-hardware control; 3.5x became 1.1x across accelerators with no code change.
  • Frequency and turbo confound. Lies as: faster numbers that are really a higher clock or a turbo window. Check: pin the governor, record the scaling driver, and confirm vCPU entitlement is not the moving variable.
  • Harness-not-algorithm gap. Lies as: startup or transport overhead counted as the unit of work. Check: align the measured unit to the product path; a 5.6x gap was the harness (aside-codemode #42).
  • Drifting baseline. Lies as: a stale "2x faster" that only held against an old reference. Check: date the comparison and re-run the current baseline; a 2x can fall to 1.4x against a faster successor with no code change.

That last one is worth dwelling on, because it means even a correct number ages out. An undated multiple is structurally unreproducible: the reference it was measured against has moved on, and you have no way to reconstruct what it beat. When a README gives a bare multiple with no date and no baseline version, the honest verdict is often unreproducible even if the code is excellent.

Grade the claim and attach the evidence

Assign one of three verdicts and attach the artifacts that let someone else check your work. The verdict is your judgment; the evidence is what makes it defensible.

  • Holds. The same-hardware control shows a real algorithmic gain within the claimed band, with dispersion that does not swallow it. Attach the control ratio and the distributions.
  • Hardware-inflated. The absolute number may be real on specific hardware, but the same-hardware control collapses the ratio, or the README compared across machines. The gain is the machine, not the code. Attach both the cross-machine and same-hardware numbers.
  • Unreproducible. The harness is absent or broken, provenance is ad-hoc, or the claim is an undated bare multiple with no baseline. Attach the provenance findings - the missing script, the mismatched results folder, the grep that came up empty.

Before you call the verdict final

  • The exact figure from the README was found in a committed script, or its absence is documented
  • The build was made from a pinned full commit SHA, with suite and runner commits recorded separately
  • A pin manifest lists every input version, digest, and seed, with corpus/native/provider seeds kept distinct
  • At least five repetitions were run with warmup discarded and fixtures excluded from timing
  • You report p50, p95, and p99 with dispersion, not a single mean
  • A same-hardware control ratio exists, measured with the library and baseline on one machine
  • Raw result artifacts are archived with hashes alongside a single reproduction command
  • The verdict (holds / hardware-inflated / unreproducible) is stated with its supporting artifact named

Staffing this work and keeping it current

Reproduction is a specialist skill, and the title that names it is noisier than it looks. In Refolk's index, "Performance Engineer" resolves to 1,119 current profiles in the US, 243 in the UK, and 84 in Germany - but the label conflates software performance engineers with motorsport ones, where the top UK and German employers include Red Bull Racing, Aston Martin F1, and TOYOTA RACING. Treat those counts as upper bounds for the software reader.

CountryCurrent countvs Germany
United States1,11913.3x
United Kingdom2432.9x
Germany841.0x

The practical consequence: sourcing on title alone over-counts the people who can actually reproduce a software benchmark. You want people who publish dispersion numbers and have handled cross-accelerator work, not people who share a job title with a pit-lane engineer. Describing exactly that in plain English - and excluding motorsport - is where Refolk removes the friction that a title filter cannot. When you need a reviewer for a contested claim, asking for "performance engineers in Germany who do software tuning, exclude motorsport" returns the right people the first time.

To keep a verdict current, re-run two things on a schedule: the same-hardware control against the current baseline version, since a 2x can drift to 1.4x against a faster successor with no code change, and the pin manifest, since the upstream tag may have moved off the SHA you measured. A verdict without a date is a claim that will quietly expire. Attach the date, the commit, and the baseline version, and your reproduction stays defensible long after the README's number stops being true.

Questions practitioners ask

How do I reproduce a benchmark from a GitHub repo if there's no benchmark script?

You cannot reproduce it, and that is a verdict, not a dead end. If there is no bench target, no scripts/bench.sh, and no results file that a committed script regenerates, the claim has no re-measure path and grades unreproducible. Grep the repo for the exact figure first to confirm it is not hiding in a notebook or a results folder, then document the absence as your evidence.

How do I separate algorithmic gains from hardware in a speedup claim?

Run a same-hardware control: put the library and the baseline it claims to beat on one machine and measure the ratio there. A README's cross-machine table proves nothing about the algorithm, because the same config has moved from 3.5x on one accelerator to 1.1x on another with no code change. Only the single-rig ratio isolates the algorithmic component.

Why can't I reproduce the number in a repo's README?

The most common reasons, in order, are that the harness is absent or broken, the numbers came from ad-hoc runs with no script, or you measured the wrong commit. Fewer failures are an actually wrong algorithm. Check provenance before you spend hours on the environment, because in the documented cases the claim died at step one.

How many repetitions do I need to trust a benchmark result?

Five is a common floor, and you should report dispersion such as p50, p95, and p99 rather than a single mean. Be aware of defaults: Google Benchmark runs 1 repetition and reports no dispersion unless you pass a flag, while Criterion.rs runs 100 samples with bootstrap confidence intervals. One fast run inside normal variance can read as a win when it is noise.

How do I pin a repo to the exact commit its benchmark was run at?

Resolve the README's version tag to a full commit SHA and build your worktree from that SHA, never from current HEAD. A full SHA guarantees the exact code even if the tag is moved later. If the benchmark suite and the runner live at different commits, record both, because analysis code can move substantially between them.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next