Staff AI Infrastructure Engineer

biohub · Redwood City, CA (Hybrid)

Posted
82 days ago
Last confirmed live
Today

What this role involves

This role owns the reliability, observability, and scaling of large-scale multi-GPU AI clusters for biological research. It involves debugging deep infrastructure issues across storage, networking, and compute layers, building automation and tooling, and collaborating with AI researchers to support frontier training runs. The mission is to cure disease using AI.

Skills this posting asks for

  • slurm
  • kubernetes
  • infiniband
  • gpu
  • distributed training
  • hpc
  • linux
  • python
  • storage
  • networking
  • automation
  • configuration-as-code
  • capacity planning
  • observability
  • incident response
  • vendor management
  • root cause analysis

Requirements

  • Level: staff

From the employer’s posting

Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological founda…

Read the full description on biohub’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at biohub

All 5 roles at biohub