- Location
- San Francisco
- Level
- Senior
- Posted
- 6 days ago
About this role
About the Role
Together AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-performance AI compute environments.
This is a hands-on engineering role for someone with deep networking expertise who can also troubleshoot across Linux, Kubernetes, automation, and application boundaries. You will work on large-scale, multi-vendor data center networks and help ensure they remain highly available, reliable, scalable, and performant.
The ideal candidate has strong networking fundamentals, experience operating complex networks at scale, and a structured, evidence-based approach to troubleshooting. You should be comfortable owning problems from initial investigation through root cause and resolution, including situations where the issue may extend beyond the network itself.
Requirements
- 8+ years of professional experience designing, building, and supporting large-scale production data center, cloud, service-provider, or high-performance computing networks (excluding enterprise networks).
- Deep understanding of TCP/IP and strong experience with technologies such as BGP, OSPF, VXLAN, EVPN, ECMP, and QoS.
- Experience designing and supporting multi-tenant network environments using technologies such as VRFs, VLANs, overlays, and policy-based segmentation.
- Hands-on experience deploying and troubleshooting network platforms from vendors such as Arista, Cisco, Juniper, and NVIDIA.
- Strong troubleshooting skills using tools such as Wireshark, tcpdump, MTR, curl, nmap, and standard Linux networking utilities.
- Ability to diagnose connectivity, latency, packet-loss, routing, and performance issues across the network, host, and application layers.
- Experience developing or maintaining network automation using Python, Ansible, or similar tools.
- Experience working through a Git-based software development lifecycle, including branching, code review, validation, linting, testing, CI/CD, deployment, and rollback.
- Working knowledge of Kubernetes networking, including pods, services, CNIs, and basic connectivity troubleshooting.
- Foundational knowledge of RDMA networking and technologies such as RoCE or InfiniBand.
- Experience with cloud networking in AWS, GCP, or Azure.
- Strong Linux administration and troubleshooting skills.
Responsibilities
- Design, deploy, operate, and maintain global, multi-vendor, multi-protocol networks supporting high-performance AI compute infrastructure.
- Troubleshoot complex network and application-connectivity issues, identify root causes, and drive problems through resolution.
- Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints.
- Participate in architecture and design reviews to ensure solutions meet requirements for performance, availability, scalability, security, and operational supportability.
- Develop and maintain automation, validation, and operational tooling that improves network reliability and reduces manual effort.
- Evaluate network hardware, software, optics, and emerging technologies for use in production environments.
- Establish standards and operational best practices for network design, deployment, monitoring, change management, and incident response.
- Lead projects addressing complex technical challenges and contribute directly to the network engineering roadmap.
- Partner with infrastructure, systems, security, and application teams to troubleshoot issues that cross traditional ownership boundaries.
Preferred
- Hands-on experience deploying or operating RoCE and/or InfiniBand fabrics.
- Experience supporting GPU clusters, HPC environments, distributed storage, or other high-bandwidth and latency-sensitive workloads.
- Understanding of AI training and inference traffic patterns and the demands they place on network infrastructure.
- Experience operating networks spanning thousands of devices, multiple data centers, and multiple geographic regions.
- Familiarity with AI-assisted engineering tools and the ability to validate, test, and safely deploy AI-generated automation or code.
About Together AI
Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.
Compensation
We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $190,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy
As published by Together AI. Applications are handled on their site.
Skills this posting mentions
About Together AI
Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform - from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories. Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves. Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.
All 65 openings at Together AIOne click, then it is written
Apply to Together AI with a resume written for this role.
Queue Senior Network Engineer and I read the posting, rewrite your resume against it, draft the cover letter, and score the fit. Then you press send, or press one button and I fill in Together AI’s form for you.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- 25 sent a week, free
- No card
- Nothing sent until you say so
More roles at Together AI
See all- Today
- Today
- 3 days ago
- 4 days ago
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Bangalore IndiaRemote
2 locationsStaffEngineering - 4 days ago
- 4 days ago
Similar roles elsewhere
See morePut this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.
Listed from the job board Together AI publishes. Refolk is not the employer and does not handle their hiring. Applications go to Together AI directly.