Staff Observability Platform Engineer

nscaleoperationsukltd · US

Posted
4 days ago
Last confirmed live
2 days ago

What this role involves

Nscale is seeking a Staff Observability Platform Engineer to build and evolve observability platforms for GPU clusters and AI workloads. The role involves designing scalable observability solutions, partnering with SRE and infrastructure teams, and leading technical initiatives. Requires 6+ years of experience with tools like Prometheus, Grafana, and OpenTelemetry, and proficiency in Go or Python.

Skills this posting asks for

  • prometheus
  • thanos
  • victoriametrics
  • grafana
  • loki
  • tempo
  • opentelemetry
  • clickhouse
  • elastic
  • go
  • python

Requirements

  • 6 years of experience
  • Level: staff

From the employer’s posting

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost manag…

Read the full description on nscaleoperationsukltd’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at nscaleoperationsukltd

All 27 roles at nscaleoperationsukltd