- Location
- United States
- Workplace
- Remote
- Employment
- Full time
- Level
- Mid level
- Posted
- 6 weeks ago
About this role
Company Overview:
We are building Protege to solve the biggest unmet need in AI - getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI - and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
About the Role
We're hiring a Forward Deployed Machine Learning Engineer in our Benchmarks and Evaluations vertical. You'll be the first MLE dedicated to this vertical and will work directly with the GM and our researchers to scale Protege’s position as a renowned leader in the space.
At Protege, we believe that real world data is one of the largest bottlenecks to AI progress. Our data and data expertise position us to be neutral arbiters for the market, helping model builders understand the current performance of their models, identify what data will improve performance, and show that improvement over time. Benchmarks and evaluations power that cycle. As an early engineer in the Benchmarks and Evaluations vertical, this role is an opportunity to help build the technical foundation for a critical area that greatly benefits current and future customers.
What You'll Do
Work on the eval foundation
• Partner with the GM and early customers to define what constitutes strong evals in different domains
• Work with Protege researchers to design and build benchmarks
• Build the standards on how different modalities should be processed
Own infrastructure
• Build the backend the vertical runs on which includes data pipelines, execution environments, storage, and orchestration
• Stand up sandboxed environments for agentic evals, where models need tools, code execution, or multi-step tasks
Go from fast iteration to product
• Find repeatable eval patterns, infrastructure gaps, and product opportunities from live engagements
• Partner with DataLab (our research team) on domain-specific data and research questions
What Success Looks Like
In the first 90 days, we expect the following:
• Build an understanding of the evals landscape, the GM's strategy, and customer demand
• Build an understanding of what our platform and data partners can support today, and where the gap is for eval building
• Identify the largest technical bets and ship multiple iterations of the eval infrastructure
• Own the engineering portion of customer engagements end to end
What You Bring
Must Haves
• 4+ years of engineering experience
• Hands-on ML work evaluating models
• Have previously owned backend and infrastructure
• High ambiguity tolerance and bias to action
• Comfort working with urgency to meet the pace and volume of the market demands
• Strong written communication
Nice to Haves
• Prior experience building benchmarks, evals, or human data pipelines for LLMs
• Time at a frontier lab, an eval-focused team, or a research org
• Founding or early engineer experience at a fast-moving startup
• Familiarity with agentic systems, RL environments, code-execution sandboxes, TEE/TREs
Protege's Values
Pass the Loved Ones' Test
We act with integrity and do the right thing - especially when it's hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.
As published by Protege. Applications are handled on their site.
Skills this posting mentions
About Protege
The biggest unmet need in AI today is getting access to the right training data. Data holders often don’t know where to start and are rightly concerned about governance, intellectual property, and security implications. AI companies can spend years finding and negotiating access to the data they need. Protege is solving these problems by providing an easy-to-use platform to connect data holders with vetted data users.
All 16 openings at ProtegeOne click, then it is written
Apply to Protege with a resume written for this role.
Queue Forward Deployed Machine Learning Engineer and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Protege’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Protege
See all- 6 weeks ago
- 6 weeks ago
- 7 weeks ago
- 7 weeks ago
- 2 months ago
- 2 months ago
Technical Product Manager, Data Ingestion & Quality
United StatesRemote
$120k - $200k/yrMid levelProduct
Similar roles elsewhere
See more- Today
Sr. Staff Data Scientist - Ads Measurement, Signals, Privacy
RedditRemote - United StatesRemote
$233k - $326k/yrStaffData and ML - Today
Staff Data Scientist - Ads Measurement, Signals, Privacy
RedditRemote - United StatesRemote
$217k - $304k/yrStaffData and ML - Today
Senior Staff Data Scientist - Consumer Experimentation
RedditRemote - United StatesRemote
$233k - $326k/yrStaffData and ML - Today
Senior Machine Learning Engineer, Ads Optimization
RedditRemote - United StatesRemote
$217k - $303k/yrSeniorData and ML - Today
Staff Machine Learning Engineer, Retrieval
RedditRemote - United StatesRemote
$230k - $322k/yrStaffData and ML - Today
Senior Staff Machine Learning Engineer, Ads Ranking
RedditRemote - United StatesRemote
$293k - $410k/yrStaffData and ML
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Protege publishes. Refolk is not the employer and does not handle their hiring. Applications go to Protege directly.