Senior Machine Learning Engineer, LLM Inference Optimization

nebius · Palo Alto, California, United States

Posted
3 days ago
Last confirmed live
2 days ago

What this role involves

Nebius is hiring a Senior Machine Learning Engineer to optimize LLM and VLM inference for their AI cloud platform. The role involves owning optimization projects for model families and customer endpoints, deploying inference engines, and implementing model compression and serving optimizations. The position is based in Amsterdam with global R&D hubs.

Skills this posting asks for

  • llm
  • vlm
  • vllm
  • sglang
  • tensorrt-llm
  • triton inference server
  • nvidia dynamo
  • quantization
  • quantization-aware training
  • distillation
  • low-bit serving
  • speculative decoding
  • kv-cache optimization
  • prefix caching
  • chunked prefill
  • continuous batching
  • disaggregated prefill/decode
  • gpu
  • distributed systems
  • model training frameworks
  • rl
  • benchmarking

Requirements

  • Level: senior

From the employer’s posting

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…

Read the full description on nebius’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at nebius

All 30 roles at nebius