Software Engineer, GPU Infrastructure

Fluidstack · San Francisco, CA

Posted
34 days ago
Last confirmed live
1 day ago

What this role involves

This role involves owning compute fleet health, building automation for fault detection and repair, designing GPU qualification platforms, and managing Redfish/BMC tooling at scale. The engineer will work with Kubernetes and bare metal to ensure reliability and performance of large GPU clusters.

Skills this posting asks for

  • kubernetes
  • bare metal
  • redfish
  • bmc
  • gpu
  • automation
  • metrics
  • alerting
  • firmware
  • telemetry
  • log collection
  • performance baselining
  • npi
  • observability
  • orchestration

From the employer’s posting

ABOUT FLUIDSTACK We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if mod…

Read the full description on Fluidstack’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Fluidstack

All 35 roles at Fluidstack