Senior Backend Engineer, Inference Platform

togetherai · San Francisco

Posted
43 days ago
Last confirmed live
2 days ago

What this role involves

Senior Backend Engineer responsible for building and optimizing the inference platform for generative AI models. Works with large-scale GPU infrastructure, contributes to open-source inference projects, and requires deep experience in distributed systems and low-level performance optimization.

Skills this posting asks for

  • rust
  • go
  • python
  • typescript
  • kubernetes
  • container orchestration
  • cuda
  • triton
  • nccl
  • infiniband
  • nvlink
  • mpi
  • sglang
  • vllm
  • nvidia dynamo
  • load balancing
  • auto-scaling
  • multi-tenant
  • prefix caching
  • distributed systems
  • os concepts
  • llms
  • generative models

Requirements

  • 5 years of experience
  • Level: senior

From the employer’s posting

About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the…

Read the full description on togetherai’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at togetherai

All 19 roles at togetherai