Senior Site Reliability Engineer - Core Cloud Platform
Lambda · Remote
- Posted
- 25 days ago
- Last confirmed live
- 1 day ago
What this role involves
Lambda is hiring a Senior Site Reliability Engineer to improve the reliability and scalability of their Core Cloud Platform, which powers compute provisioning and infrastructure orchestration across physical data centers. The role involves operating Kubernetes, building observability, defining SLOs, and leading incident response. The position requires presence in San Francisco, San Jose, or Bellevue office 4 days per week.
Skills this posting asks for
- kubernetes
- terraform
- argo cd
- flux
- helm
- kustomize
- opentelemetry
- prometheus
- grafana
- datadog
- go
- python
- gitops
- ci/cd
- linux
- etcd
- rbac
- oidc
- chaos engineering
Requirements
- 7 years of experience
- Level: senior
- Remote policy: onsite
From the employer’s posting
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintellige…
Read the full description on Lambda’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Lambda
- Staff Technical Program Manager - Business TechnologyRemote
- Senior Platform Engineer - Core InfrastructureRemote
- Senior Site Reliability Engineer - FleetRemote
- Senior Software Engineer - Managed KubernetesRemote
- Senior Site Reliability Engineer - Managed KubernetesRemote
- Staff Software Engineer - ComputeRemote