ML Infrastructure Engineer
nebius · Remote
- Posted
- 3 days ago
- Last confirmed live
- 2 days ago
What this role involves
Seeking an ML Infrastructure Engineer to lead GPU platform benchmarking for ML/AI workloads. Responsibilities include profiling GPU performance, debugging and optimizing workloads, and developing tools for performance metrics. Requires deep understanding of deep learning frameworks and GPU stack.
Skills this posting asks for
- pytorch
- jax
- megatron-lm
- tensorrt-llm
- cuda
- nccl
- docker
- kubernetes
- python
- vllm
- sglang
- tensorrt
From the employer’s posting
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…
Read the full description on nebius’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at nebius
- IT Infrastructure Engineer – RMA & Hardware DiagnosticsMinnesota, United States
- AI/ML Specialist Solutions ArchitectRemote
- Frontend Engineer - User InterfaceRemote
- Infrastructure Site Reliability EngineerUnited States
- Partner Solutions ArchitectRemote
- Principal ML Solutions Architect - Token FactoryUnited States