Software Engineer, Compute (GPU)
Fluidstack · San Francisco, CA
- Posted
- 34 days ago
- Last confirmed live
- 1 day ago
What this role involves
This role involves owning the health of the compute fleet, building automation for GPU repair and qualification, and developing low-level firmware and BMC tooling. The engineer will design metrics pipelines, alerting, and a unified health view for Kubernetes-orchestrated and bare metal GPU workloads. The goal is to make hyperscale AI compute operable and reliable at massive scale.
Skills this posting asks for
- gpu
- kubernetes
- bare metal
- redfish
- bmc
- metrics pipelines
- alerting
- automation
- gpu qualification
- burn-in
- performance baselining
- npi
- firmware-level telemetry
- log collection
- reliability engineering
- incident discipline
- compute fleet health
- tooling
From the employer’s posting
ABOUT FLUIDSTACK We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if mod…
Read the full description on Fluidstack’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Fluidstack
- Software Engineer, Applied AISan Francisco, CA
- Product Manager, Business OperationsNew York, NY
- Data EngineerAustin, TX
- Technical Program Manager, NPIAustin, TX
- Network Engineer, WirelessRemote
- Product Manager, Compute OperationsNew York, NY