Principal Site Reliability Engineer, Machine Learning
Cambridge Mobile Telematics · Cambridge, MA
- Posted
- 19 days ago
- Last confirmed live
- 1 day ago
What this role involves
Cambridge Mobile Telematics is seeking a Principal Site Reliability Engineer for Machine Learning to ensure the operational health of Ray clusters on AWS EKS and Databricks workloads. The role involves owning SLOs, maintaining observability, and automating infrastructure using Terraform and CI/CD. Candidates need 7+ years of SRE experience and expertise in AWS, Kubernetes, and Python.
Skills this posting asks for
- aws
- ec2
- ecs
- eks
- sqs
- lambda
- dynamodb
- rds
- aurora
- s3
- iam
- kubernetes
- docker
- terraform
- python
- cloudwatch
- datadog
- ray
- databricks
- linux
- ci/cd
Requirements
- 7 years of experience
- Level: principal
From the employer’s posting
CMT is looking for a Principal Site Reliability Engineer I, Machine Learning to help us change the world. CMT has helped protect over 65 million drivers and prevent over 126,000 crashes worldwide. We build AI to solve some of the most difficult challenges in mobility — und…
Read the full description on Cambridge Mobile Telematics’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Cambridge Mobile Telematics
- IT Security EngineerCambridge, MA
- Principal Security Engineer, Product & ApplicationCambridge, MA
- Principal Security Engineer, Cloud & InfrastructureCambridge, MA
- Principal Machine Learning Engineer, Foundation ModelsCambridge, MA
- Principal Technical Product Manager, Machine LearningCambridge, MA
- Principal Software Engineer, Full StackCambridge, MA