Senior Machine Learning Engineer, LLM Inference Optimization
nebius · Palo Alto, California, United States
- Posted
- 3 days ago
- Last confirmed live
- 2 days ago
What this role involves
Nebius is hiring a Senior Machine Learning Engineer to optimize LLM and VLM inference for their AI cloud platform. The role involves owning optimization projects for model families and customer endpoints, deploying inference engines, and implementing model compression and serving optimizations. The position is based in Amsterdam with global R&D hubs.
Skills this posting asks for
- llm
- vlm
- vllm
- sglang
- tensorrt-llm
- triton inference server
- nvidia dynamo
- quantization
- quantization-aware training
- distillation
- low-bit serving
- speculative decoding
- kv-cache optimization
- prefix caching
- chunked prefill
- continuous batching
- disaggregated prefill/decode
- gpu
- distributed systems
- model training frameworks
- rl
- benchmarking
Requirements
- Level: senior
From the employer’s posting
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…
Read the full description on nebius’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at nebius
- IT Infrastructure Engineer – RMA & Hardware DiagnosticsMinnesota, United States
- AI/ML Specialist Solutions ArchitectRemote
- Frontend Engineer - User InterfaceRemote
- Infrastructure Site Reliability EngineerUnited States
- Partner Solutions ArchitectRemote
- Principal ML Solutions Architect - Token FactoryUnited States