- Location
- United States
- Workplace
- Remote
- Employment
- Full time
- Level
- Manager
- Posted
- 2 days ago
About this role
About Fundamental
Fundamental is an AI research lab pioneering the future of enterprise decision-making. Our flagship model, NEXUS is the world's most powerful Large Tabular Model (LTM) - purpose-built for the structured records that contain trillions of dollars in business value. With $275m in funding from leading investors and trusted by Fortune 100 companies, Fundamental is giving businesses the Power to Predict.
At Fundamental, you'll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world's largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.
About the role
NEXUS is already in production, and we are working to substantially improve its predictive quality, latency, and cost efficiency. You will lead the team that owns how NEXUS is measured, building the shared evaluation platform that engineering, research, and Applied AI all rely on: consistent benchmarks, curated datasets, and standards for metrics, data splits, leakage prevention, and benchmark contamination that hold up under scrutiny.
Research needs to trust results before committing compute to it, engineering needs regressions caught before they reach production, and when a customer's own data science team benchmarks NEXUS against their own models, your platform is what Applied AI relies upon. You will benchmark NEXUS against competing approaches, using fair tuning budgets, data access, and latency measurement protocols, and turn what you find into research priorities and release recommendations.
This is a player-coach role: you will hire and manage a small team while staying hands-on with the code, the experiment design, and the methodology yourself. There is no evaluation function to inherit here - what you build becomes the standard the rest of the company measures NEXUS against.
Key responsibilities
Build a shared evaluation platform for engineering, research, and Applied AI. Continuously add and maintain models and curated datasets so teams can run benchmarks and investigate results independently.
Define evaluation standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination.
Build reproducible pipelines with versioned inputs and artifacts, and integrate regression checks into research and release workflows.
Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency measurement protocols.
Support Applied AI’s customer POC evaluations with tooling, methodological guidance, and analysis.
Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics.
Turn findings into research priorities, release recommendations, and evidence-backed customer improvement plans.
Hire and develop the team, set priorities, and stay hands-on with code and experimental design.
Must have
Experience owning evaluation for tabular ML systems used in production or consequential customer decisions.
Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation.
Practical experience finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related entities across splits.
Strong Python and SQL skills, familiarity with scikit-learn and gradient-boosted trees, and experience building reliable ML tooling or platforms used by other teams.
Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement.
Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved.
Clear written and spoken communication with researchers, engineers, and customer data scientists, including the willingness to challenge claims the evidence does not support.
Nice to have
Experience evaluating tabular foundation models or AutoML systems.
Experience measuring how model optimisations affect predictive quality and inference performance.
Experience with relational or multi-table data, and Snowflake or Databricks environments.
Experience evaluating automated or agent-driven ML workflows, including failures that aggregate metrics can hide.
Benefits
Competitive compensation with salary and equity
Comprehensive health coverage for you and your dependents
Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys
Relocation support for employees moving to join the team in one of our office locations
A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action
As published by Fundamental. Applications are handled on their site.
Skills this posting mentions
About Fundamental
For decades companies have relied on archaic tools to inform decisions and make bets on the future. Until now. Fundamental empowers businesses to turn gambles into guarantees and determine their future with far greater accuracy than ever before. Built by DeepMind alumni and trusted by Fortune 100 enterprises, NEXUS is our most powerful Large Tabular Model (LTM). By revealing the hidden language of tables, NEXUS unlocks trillions of dollars of value by giving businesses the Power to Predict™.
All 21 openings at FundamentalOne click, then it is written
Apply to Fundamental with a resume written for this role.
Queue Evaluations Team Lead and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Fundamental’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Fundamental
See all- 6 days ago
- 3 weeks ago
- 7 weeks ago
Forward Deployed Data Scientist - Oil & Gas, Houston
Houston, TexasHybrid
$135k - $225k/yrMid levelData and ML - 2 months ago
- 3 months ago
- 3 months ago
Similar roles elsewhere
See more- Yesterday
Principal Software Engineer I - CI/CD - Platform Engineering Productivity
ElasticUnited States
$160k - $304k/yrPrincipalEngineering - Yesterday
Staff Software Engineer, Streaming Infrastructure
AirbnbUnited States
$204k - $255k/yrStaffEngineering
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Fundamental publishes. Refolk is not the employer and does not handle their hiring. Applications go to Fundamental directly.