- Location
- Europe
- Workplace
- Remote
- Employment
- Full time
- Level
- Mid level
- Posted
- 7 weeks ago
About this role
About Cantina:
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
About the Role:
We are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.
What You’ll Do:
Design and maintain inference infrastructure for generative audio model architectures.
Implement and manage high-performance inference engines.
Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.
Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.
Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.
Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.
Optimize inference performance for both streaming and batch applications.
What You’ll Bring:
Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.
Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.
Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
Experience with GPU-accelerated inference and performance profiling techniques.
Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.
Compensation:
The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
Benefits for U.S.-based roles:
Competitive salary and generous company equity
Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina
42 days of paid time off, including:
15 PTO days
10 sick days
15 company holidays
2 floating holidays
Generous parental leave & fertility support
401(k) retirement savings plan
Lifestyle spending account - $500/month to use however you’d like
Complimentary lunch and snacks for in-office employees
One Medical membership, and more!
As published by Cantina. Applications are handled on their site.
Skills this posting mentions
About Cantina
The most advanced AI character creator. Unleash AI bots that talk, feel, and capture their adventures with selfies.
All 19 openings at CantinaOne click, then it is written
Apply to Cantina with a resume written for this role.
Queue Machine Learning Engineer, Ops and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Cantina’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Cantina
See all- 6 weeks ago
- 6 weeks ago
- 6 weeks ago
- 2 months ago
- 2 months ago
Senior Creative Strategist, Performance Marketing
San Francisco, CaliforniaRemote
$150k - $200k/yrSeniorMarketing - 2 months ago
Similar roles elsewhere
See more- 6 weeks ago
Machine Learning Engineer, Speech - Joint Audio-Video Modeling
CantinaEuropeRemote
Mid levelData and ML
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Cantina publishes. Refolk is not the employer and does not handle their hiring. Applications go to Cantina directly.