Senior Principal AI Engineer
cerence · Remote
- Posted
- 18 days ago
- Last confirmed live
- 4 days ago
What this role involves
This role focuses on designing and operating distributed training systems for large neural networks across GPU clusters. Responsibilities include optimizing multi-node, multi-GPU execution, diagnosing bottlenecks, and improving training stability. The ideal candidate has deep experience with distributed systems, GPU orchestration tools, and training frameworks like PyTorch Distributed, Megatron-LM, and DeepSpeed.
Skills this posting asks for
- slurm
- kubernetes
- ray
- runai
- nccl
- rdma
- infiniband
- nvlink
- pytorch distributed
- megatron-lm
- deepspeed
- activation checkpointing
- zero offload
- distributed systems
- ml systems
- gpu clusters
- data parallelism
- tensor parallelism
- pipeline parallelism
- hpc
Requirements
- Level: senior
- Remote policy: remote
Apply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at cerence
- Principal Software Engineer – Robot Applications & Voice AIRemote
- Principal Solutions Architect – Robotics & Voice AIRemote
- Sr. Principal Software EngineerRemote