Machine Learning Engineer - Inference

togetherai · San Francisco · $160k–$230k

Posted
43 days ago
Last confirmed live
2 days ago
Published range
$160k–$230k

What this role involves

Together AI is seeking a Machine Learning Engineer to join their Inference Engine team, focusing on optimizing and enhancing the performance of AI inference systems. The role involves designing and building production systems for the AI inference engine, developing runtime inference services, and collaborating with researchers and engineers. Required skills include Python, PyTorch, and experience with high-performance code and low-level OS concepts. Preferred knowledge includes AI inference systems like TGI, vLLM, and techniques like speculative decoding and CUDA/Triton programming.

Skills this posting asks for

  • python
  • pytorch
  • multi-threading
  • memory management
  • networking
  • storage
  • performance
  • scale
  • tgi
  • vllm
  • tensorrt-llm
  • optimum
  • speculative decoding
  • cuda
  • triton
  • rust
  • cython
  • compilers

Requirements

  • 3 years of experience
  • Level: mid

From the employer’s posting

About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models…

Read the full description on togetherai’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at togetherai

All 19 roles at togetherai