Senior Software Engineer - AI Compute, Together Cloud
Together AI San FranciscoUSD 220,000–270,000 / yearSenior
Key requirements
- Golang
- Aws
- Azure
- Gcp
- Kubernetes
- Terraform
- Ci/Cd
- Linux
- Llm
- Salesforce
About the Role
Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products.
As a Senior Software Engineer focusing on AI Compute in the Together Cloud org, you will build and own major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs.
The hard problems in rapidly scaling heterogeneous GPU fleets are software problems: fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes. We solve them by building global and in-DC services, Kubernetes operators, and high-performance SDN libraries, forking hypervisors and Linux kernels, and building infra tailored to inference and fine-tuning. You'll own massive greenfield projects across their full lifecycle – scoping the problem and writing PRDs with our PMs, designing the system, and building it through to GA launch.
Responsibilities
Build the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that makes GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware.
Build and maintain our in-DC IaaS layer: the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers, including VMs, parallel filesystems, VPCs, and InfiniBand partitions. Implement and harden the bring-up path for a new Vera Rubin data center with thousands of GPUs.
Scale the distributed GPU scheduling and global management plane: the control-plane services behind on-demand and reserved clusters, including the automation that onboards new capacity and raises per-cluster limits.
Harden the monitoring and automated remediation layer for fault tolerance: automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference fault-tolerant.
Own your components end-to-end: write the design docs, break the work into milestones that ship incrementally, and improve the reliability of what's already in production.
Raise the bar around you: code review, design feedback, and mentoring junior engineers.
Build the tooling other teams rely on: testing frameworks, developer tools, and documentation that make our systems robust and usable across teams, plus contributions to the core, open-source Together AI platform.
To be successful, you'll need to be deeply technical, ready to own projects truly end-to-end, and an excellent communicator — strong software development fundamentals, strong systems knowledge and troubleshooting instincts, and the collaboration and diplomacy skills to work across teams.
Requirements
5+ years of professional software development experience, with strong proficiency in at least one backend programming language (Golang desired).
Demonstrated ownership of large-scale projects driven end-to-end to completion, writing high-performance, well-tested, production-quality code.
Demonstrated experience building and operating high-performanc
See your match score for this role.
Xecodai maps the interview stages and shows what is preventing a 95% match.
