Refolk
ReferenceEngineering and open source

The AI-Engineer Footprint Reference: What Each LLM Artifact Proves

You will be able to look at any artifact in a candidate's public AI footprint and state what it proves, how it lies, and when it goes stale.

15 min readLast reviewed October 7, 2026Read as Markdown

Key takeaways

  • In Refolk's index, US AI engineers self-report LangChain 41 times more often than RAG (872 versus 21), so the framework keyword is cheap and a committed eval harness is the scarce, discriminating signal.
  • An artifact's imports date it precisely: pre-September-2024 LangChain or pre-August-2024 LlamaIndex patterns prove dated familiarity no matter how polished the README.
  • Around 70% of enterprise AI work involves observability, yet most portfolios omit cost and latency entirely, which makes their presence a strong production marker.
  • With 15,926 public mcp-server repos and 43% of servers carrying command-injection flaws, a forked MCP server proves little, while an authored one that handles auth proves judgment.
  • Explicit RAG tagging runs 3.5x higher in the US than Germany (21 versus 6) in Refolk's index, so one verified eval-gated RAG repo outranks dozens of keyword matches.
  • The consensus heuristic is 3 to 5 deeply evaluated projects with a live URL and an eval table, because one production RAG project beats five tutorial clones.

You are screening an engineer's public AI and LLM work, and the footprint is crowded. This reference is for engineering managers, technical founders, developer-relations leads, and technical sourcers who need to triage a candidate's AI artifacts in minutes, not infer skill from a blog post that lists signals without saying when each one lies. It gives you one row per artifact type: what it proves about production capability, the failure mode that makes it look stronger than it is, and how long the signal stays fresh.

The scale is the first problem. GitHub counted 4.3 million AI-related repositories, nearly doubled in under two years, with 1.1 million public repositories importing large language model SDKs and 693,867 of those created in the past twelve months. The footprint is not thin. The skill inside it is.

Why the AI footprint needs its own reference

The AI footprint needs its own reference because the artifacts flooding profiles share names with production systems but rarely share their guts. A RAG app, an agent, an eval harness, a fine-tune, an MCP server, and a vector pipeline each prove a different thing, and each lies in a different way. Reading them as one undifferentiated pile of "AI work" is how a tutorial clone gets a phone screen.

The ecosystem grew fast enough that the ratio of keyword to capability inverted. Lightcast data shows generative AI skill postings rose from 16,000 in 2023 to 66,000 in 2024, large language modeling from 5,000 to 20,000, and prompt engineering from 1,400 to nearly 6,300. Demand pulled every engineer toward listing the stack. Supply of engineers who can actually ship a reliable AI feature with guardrails, monitoring, and measurable quality did not keep pace.

41:1
How much more often US AI engineers self-report LangChain than RAG
In Refolk's index, 872 US AI/ML/LLM engineers list LangChain as a skill; 21 list Retrieval Augmented Generation. The framework keyword is cheap; the capability is scarce.

That ratio is the spine of this document. The keyword is cheap because engineers list the tool they imported, not the capability they proved. The eval table is scarce because it is expensive to produce and trivial to verify. Every row in the reference below is really a way of asking the same question: did this person import a framework, or did they ship and measure a system?

The artifact reference: what each type proves

Each artifact type sits on a spectrum from "ran a model locally" to "operated a system in production." The table below is the fast lookup. Jump to the row for the artifact in front of you, read what it proves, and note the failure mode that inflates it.

ArtifactWhat it provesHow it misleads
RAG appRetrieval design, grounding, and if evaluated, production judgmentFluent answers with no faithfulness or recall metric read as production when they are a demo
AgentOrchestration, tool use, and control flow under uncertainty"Works" with no cost caps or p50/p95 signals it was only run locally
Eval harnessThe scarcest skill: measuring quality and catching regressionsA test file that never gates CI, or a golden set of three queries, is cosmetic
Fine-tune / LoRATraining-loop fluency and dataset handlingA fine-tune with no eval table proves you ran a script, not that it improved anything
MCP serverProtocol integration and, if authored with auth, production judgmentA forked server among 15,926 similar repos proves familiarity, not skill
Vector pipelineChunking, embedding, and indexing at ingestA pipeline with no retrieval quality metric is plumbing without a gauge

The ordering matters. An eval harness is the highest-value artifact precisely because it is the one most candidates skip. The failure a harness catches is specific: the system answers beautifully in the launch review, then quietly degrades. An embedding update shifts retrieval, a knowledge-base refresh changes chunk boundaries, a prompt tweak makes the generator embellish, and every one of those ships without an error. An engineer who built a regression gate has met that failure and designed against it. An engineer with a RAG demo and no evals has not yet learned it exists.

Reading the stack as layers

A single AI system stacks concerns, and strong artifacts show evidence at more than one layer. The weak ones stop at the model.

The production AI stack, outermost first

  1. Operations
    Observability, cost tracking, p50/p95 latency, a live deployment URL
  2. Evaluation
    Golden set, faithfulness and recall metrics, a CI gate that fails on regression
  3. Orchestration
    Retrieval design, tool use, retry and backoff, cost caps
  4. Model
    The LLM call itself and the imported SDK
Strong artifacts show evidence above the model layer; demos stop at the call.

Most portfolios live entirely at the model layer. Around 70% of enterprise AI work involves observability, yet most portfolios omit it entirely. When you see the operations layer addressed at all, you are looking at a candidate who has operated something, not just invoked it.

The production markers that separate build from demo

The markers that separate production work from a demo cluster tightly, and practitioners converge on the same list. Look for all of them, and note which are present rather than scoring on any single one.

  • Production RAG: hybrid retrieval plus reranking plus citations, not a single-vector lookup.
  • An eval suite: faithfulness, context precision or recall, and a stated hallucination rate.
  • A live deployment URL: a running endpoint, not a screenshot.
  • README as product spec: problem statement, architecture diagram, eval numbers, and cost.
  • Observability and cost tracking: instrumentation that records what each request cost and how it behaved.
  • Error handling: retry, backoff, and cost caps.
  • Quantitative results: numbers shown, not claimed.
  • Open-source PR and issue history: evidence of collaboration, not a solo tutorial clone.

The single most reliable tell of a demo dressed as production is missing evaluation. A model with no stated accuracy, no baseline comparison, and no discussion of failure cases reads as unfinished even if the code runs. Evaluation is the part of the job that separates an engineer from a script. The second most reliable tell is cost and latency silence: production AI systems live and die on cost per request and response time, and if a repo never mentions either, the candidate is signaling they have only run a model locally.

The keyword is cheap and the eval table is scarce, so build your screen around what is expensive to fake.

This is where a sourcing tool earns its place. If you are starting from scratch rather than reviewing an inbound candidate, you want to find engineers who already show these markers, not read a thousand profiles to find the few who do. Refolk lets you ask for the capability in plain English and get back people whose public footprint actually carries it, so you spend your screening time on the markers above instead of on triage.

How to screen an AI footprint in under an hour

To screen an AI footprint, work from classification to verdict in a fixed order so you never over-read a single artifact. The procedure below takes roughly an hour per candidate and produces a one-line reason per artifact.

Screen an AI footprint end to end

  1. Inventory and classify artifacts
    List each pinned or active repo and tag its type: RAG, agent, eval harness, fine-tune, MCP server, vector pipeline, or prompt library. Done when every repo has a type label.
  2. Verify authorship
    Check commit history, authorship concentration, and whether the repo is a fork or template clone. Done when you can say the candidate wrote the core, not just forked it.
  3. Check for an eval harness
    Look for a golden or test set, metrics such as faithfulness, recall, and hallucination rate, and a CI regression gate. Done when you find a gate that fails the build on regression, or confirm its absence.
  4. Check production markers
    Look for observability and cost tracking, latency handling, retry and backoff, cost caps, and a live deployment URL. Done when you can state which production concerns are addressed.
  5. Read the README as a spec
    Check for a problem statement, architecture, eval numbers, and cost or latency. Done when you know whether numbers are claimed or shown.
  6. Date-check freshness
    Compare imports and commit dates against the framework deprecation calendar. Done when each artifact is labeled current or dated.
  7. Cross-reference discussion
    Read issue and PR threads for real collaboration versus a solo tutorial. Done when collaboration style is characterized.
  8. Triage
    Assign interview, hold, or reject with a one-line reason per artifact. Done when the candidate has a single verdict backed by the strongest verified artifact.

Sources disagree on weighting, and the disagreement is worth knowing. Some reviewers lead with the README as the product spec; others lead with commit and issue behavior as the authorship proof. Treat the README as the claim and the commit history as the evidence. When they conflict, the commit history wins.

What survives each screening pass

  1. Repos importing LLM SDKs
    1,100,000

    The raw public footprint

  2. New LLM-SDK repos in 12 months
    693,867

    Recent enough to date-check

  3. mcp-server topic repos
    15,926

    One artifact type, still abundant and often forked

  4. Authored, eval-gated artifacts
    scarce

    What actually survives steps 2 through 4

Volume narrows sharply once authorship and evaluation become the filter.

How each artifact misleads: the false positives

Each artifact type has a characteristic false positive, and knowing it is more valuable than knowing what the artifact proves. The table below is the failure-mode lookup. When an artifact looks strong, check its matching row before you advance the candidate.

Failure modeWhat it looks likeThe check that exposes it
Stars read as skill400 stars, many forks, zero authored production codeCommit authorship concentration, not star count
Cosmetic eval harnessA test file that never gates CI, or a golden set of three queriesA gate that fails the build, and dataset size
Demo dressed as RAGFluent answers, no faithfulness or recall metric, no monitoringSilent-degradation handling and segmented live metrics
Dated framework fluencyPolished LangChain repo on pre-v0.3 imports (LLMChain, ServiceContext)Imports against the 2024 and 2025 deprecation cutoffs
Template clone with a good READMEStrong README, no issue or PR discussion, single bulk commitPR and issue history, and commit cadence
Keyword inflationRAG, LangChain, and agents listed with no artifactWhether any repo backs the keyword at all

Two of these deserve extra weight because they catch the most polished fakes.

Stars and forks are not authorship. Not all GitHub activity in ML is worth your time. Someone can star 400 repositories and watch videos about PyTorch without writing a single line of production code. The fix is mechanical: open the contributors view, read the commit cadence, and ask whether this person wrote the core logic or forked a working project and renamed it.

Keyword inflation is the default state, not the exception. In Refolk's index, LangChain is self-reported roughly 41 times more often than RAG. The keyword is cheap and the eval table is not. When a profile lists RAG, LangChain, and agents but no repo carries a golden set or a live URL, you are reading a keyword, not a capability.

Signal freshness: when each artifact goes stale

AI artifacts go stale fast, and freshness is dictated by framework re-architectures rather than a fixed clock. AI orchestration frameworks change faster than the content teaching them can keep up. LlamaIndex and LangChain each re-architected their entire package layout in 2024 and removed their original headline abstractions within a year, so a course recorded twelve months ago often fails on the first import. The same logic applies to a candidate's repo: imports that predate a cutoff prove dated familiarity, not current skill.

CutoffWhat changedWhat a pre-cutoff import proves
LlamaIndex v0.11 (Aug 2024)Removed ServiceContextDated LlamaIndex familiarity
LangChain v0.3 (Sep 2024)Moved Pydantic v1 to v2, broke user codeRepo predates the current type system
LangChain v1.0 (Oct/Nov 2025)Moved LLMChain, RetrievalQA, AgentExecutor to langchain-classicLegacy chain patterns, not the blessed create_agent path

No published study quantifies an exact signal half-life per artifact, so treat this as a reasoned inference from the deprecation calendar rather than a measured number. The mechanism is sound even where the precise value is not established: if you see ServiceContext or LLMChain in the imports, you are looking at skill that was current before the cutoff and has not been refreshed since. Re-run this check every time you screen, because the calendar keeps moving.

Eval metrics age more slowly than orchestration code. A golden set and a faithfulness measurement remain meaningful even when the surrounding framework has moved, because the discipline of measuring does not expire the way an import path does. When you find a dated repo whose evaluation approach is still sound, weight the discipline over the stale import.

The sourcing picture: how thin the verified pool really is

The verified pool is thin, which is the whole reason this screen pays off. In Refolk's index, 21 US AI/ML/LLM engineers list Retrieval Augmented Generation as an explicit skill, against 872 who list LangChain. Explicit RAG tagging runs 3.5 times higher in the US than Germany, where only 6 engineers carry the tag. Top employers for the US RAG pool include Atlassian, Grammarly, Apple, Databricks, and Truveta; in Germany the names include Recare, Fraunhofer HHI, and Gallatin AI.

3.5x
How much more often US AI engineers tag RAG than German ones
21 in the US versus 6 in Germany in Refolk's index. A small absolute pool in both markets, which is why a single verified artifact outranks dozens of keyword matches.

The implication for screening is direct. Because the keyword is cheap and verified capability is scarce, a single eval-gated RAG repo with a live URL outranks a long list of profiles that merely mention the stack. You are not looking for the candidate with the most AI keywords. You are looking for the one whose smallest verified artifact survives every row of the failure-mode table. The market agrees with this emphasis: generative AI engineer postings rose 7 times from 2022 to 2024, and AI skills in other IT roles rose 35 times, which means demand concentrates on shipping capability, exactly the thing the production markers measure.

Before you call it done

Before you move a candidate to interview or reject, run this checklist against the strongest artifact in their footprint. It is the fastest way to confirm you screened for capability and not for keywords.

AI footprint screen, final pass

  • Every repo in the footprint has a type label (RAG, agent, eval, fine-tune, MCP server, vector pipeline, prompt library)
  • The core of the strongest artifact is authored by the candidate, not forked or bulk-committed
  • At least one artifact has an eval harness with a golden set larger than a handful of queries and a CI gate that fails on regression
  • Production markers are inventoried: observability, cost tracking, latency, retry and backoff, cost caps, live URL
  • README numbers are shown, not merely claimed
  • Imports and commit dates are checked against the framework deprecation cutoffs and labeled current or dated
  • Issue and PR history has been read for collaboration style, not just the file tree
  • Any MCP server is confirmed authored and handling auth, not a fork among thousands
  • The verdict has a one-line reason tied to a specific verified artifact

Keep this reference current by re-checking the deprecation calendar whenever you screen, because the cutoffs that date an artifact are the fastest-moving part of this document. The artifact types and their failure modes are stable; the import paths that mark a repo as fresh are not. When a new framework re-architecture lands, add its cutoff to the freshness table and re-read your pending candidates against it. The engineers worth interviewing will have repos that pass the newest cutoff, not just the ones that were current when they shipped.

The one durable rule underneath all of this: screen for what is expensive to fake. A live deployment plus an eval table is costly to produce and cheap for you to verify. A keyword costs nothing and proves nothing. Build every screen on the gap between those two, and the crowded footprint sorts itself out.

Questions practitioners ask

How do I evaluate an AI engineer's GitHub when every profile looks the same?

Start by classifying each repo by artifact type, then verify authorship before reading anything else. In Refolk's index, US AI engineers self-report LangChain 41 times more often than RAG, so the keyword tells you almost nothing. The discriminating signals are an eval harness that gates CI, a live deployment URL, and cost or latency numbers. These are expensive to fake and cheap to verify, which is exactly what makes them worth your time.

What is the difference between a RAG repo and a demo project?

A demo answers fluently in a launch review and ships silent regressions after. A production RAG repo has hybrid retrieval with reranking and citations, an eval suite measuring faithfulness and context recall, a CI gate that fails the build on regression, observability, cost tracking, and a live URL. The tell for a demo dressed as production is a model with no stated accuracy, no baseline, and no discussion of failure cases. Evaluation is the part that separates an engineer from a script.

How long does an LLM project repo stay a valid signal of current skill?

Short, and tied to framework re-architectures rather than a fixed clock. LlamaIndex removed ServiceContext in v0.11 in August 2024, LangChain moved to Pydantic v2 in v0.3 in September 2024, and LangChain v1.0 moved legacy chains to langchain-classic. An artifact whose imports predate these cutoffs proves dated familiarity, not current skill. Re-check by reading the import lines against the deprecation calendar whenever you screen.

Do GitHub stars prove an AI engineer's skill?

No. Someone can star 400 repositories and watch PyTorch videos without writing a line of production code. Stars and watched repos measure interest, not authorship. Check commit authorship concentration instead: who wrote the core logic, in what cadence, and whether the repo is a fork or template clone with a single bulk commit. The star count belongs in the noise column.

Is an MCP server in a candidate's profile a strong signal?

Only if they authored it and it handles auth. There are 15,926 public repos with the mcp-server topic, and 43% of MCP servers carry command-injection flaws. A forked server among thousands of similar repos proves little. An authored server that handles OAuth or other auth and sets guardrails proves production judgment, because it shows the engineer treated an exposed surface as a real one.

How many projects should a strong AI engineer portfolio have?

Depth beats breadth. The consensus is 3 to 5 deeply evaluated projects, each with a live URL and an eval table, pinned to no more than 6. One production RAG project with proper evals beats five tutorial clones. Reviewers click into a repo and read its commit and issue history rather than counting the tree, so a small set of verifiable artifacts outranks a long list of keyword matches.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next