Member of Technical Staff - Research, Inference
Modal · New York
About this role
ABOUT US: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable https://modal.com/blog/lovable-case-study, Ramp https://modal.com/blog/how-ramp-built-a-full-context-background-coding-agent-on-modal, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C https://modal.com/blog/modal-series-c at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g.,Seaborn https://github.com/mwaskom/seaborn,Luigi https://github.com/spotify/luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. THE ROLE: Most of the value of owning a model shows up at serving time. We're building a platform that covers the whole life of an LLM -- train it, deploy it, observe it -- and inference is where teams feel the difference every day. We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell. You will do hands-on inference research at Modal, working with the research lead to pick high-impact bets and owning them end to end. The bets that matter most are the ones that move cost per token and tail latency on t
Excerpt from the posting Modal published. Read the full description on their site before applying.
Apply to this role, tailored
Queue Member of Technical Staff - Research, Inference at Modal and I will read the posting, rewrite your resume against it, draft the cover letter, and score the fit before you send anything.
Applications finish as drafts. Nothing is sent until you read it and press send. New accounts start with 500 free credits.
More roles at Modal
See all- 2 days ago
- 2 days ago
- 3 days ago
- 3 weeks ago
- 6 weeks ago
- 2 months ago
Similar roles elsewhere
See more- Yesterday
Senior Applied ML/AI Scientist - Search
FaireNew York City, NY +1
$211k - $291k/yr2 locationsSeniorScience and research - 2 days ago
Senior Manager, Forward Deployed Research
Snorkel AINew York City, NY (Hybrid)Remote
$185k - $322k/yrManagerScience and research - 4 days ago
- 4 days ago
Listed from the job board Modal publishes. Refolk is not the employer and does not handle their hiring. Applications go to Modal directly.