Machine Learning Engineer - Inference
togetherai · San Francisco · $160k–$230k
- Posted
- 43 days ago
- Last confirmed live
- 2 days ago
- Published range
- $160k–$230k
What this role involves
Together AI is seeking a Machine Learning Engineer to join their Inference Engine team, focusing on optimizing and enhancing the performance of AI inference systems. The role involves designing and building production systems for the AI inference engine, developing runtime inference services, and collaborating with researchers and engineers. Required skills include Python, PyTorch, and experience with high-performance code and low-level OS concepts. Preferred knowledge includes AI inference systems like TGI, vLLM, and techniques like speculative decoding and CUDA/Triton programming.
Skills this posting asks for
- python
- pytorch
- multi-threading
- memory management
- networking
- storage
- performance
- scale
- tgi
- vllm
- tensorrt-llm
- optimum
- speculative decoding
- cuda
- triton
- rust
- cython
- compilers
Requirements
- 3 years of experience
- Level: mid
From the employer’s posting
About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models…
Read the full description on togetherai’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at togetherai
- Senior Product Manager, Model APIs & Developer ExperienceSan Francisco
- Senior Software Engineer - Together Cloud InfrastructureSan Francisco
- Software Engineer, Customer InsightsSan Francisco
- Senior Product Engineer, FullstackSan Francisco
- Staff Software Engineer, Inference / Compute Infrastructure EngineeringSan Francisco
- Senior Network EngineerSan Francisco