Inference Optimization ML Engineer
Rhoda AI · Mountain View
- Posted
- 103 days ago
- Last confirmed live
- Today
What this role involves
Rhoda AI is hiring an Inference Optimization MLE to optimize large multimodal models for production deployment. The role focuses on end-to-end inference performance, including latency, throughput, and efficiency improvements. The candidate will work closely with research and robotics teams to bridge the gap between training and real-world deployment.
Skills this posting asks for
- inference optimization
- ml systems
- pytorch
- jax
- quantization
- distillation
- pruning
- model compilation
- tensorrt
- torch.compile
- xla
- attention mechanisms
- kv caching
- cuda
- triton
- inference serving
- triton inference server
- vllm
- torchserve
- multimodal models
- video model inference
- edge deployment
- cloud deployment
- speculative decoding
Requirements
- 3 years of experience
- Level: mid
From the employer’s posting
At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists cap…
Read the full description on Rhoda AI’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Rhoda AI
- Senior Fleet Software EngineerMountain View
- Senior DevOps EngineerMountain View
- Product Manager, Rhoda PlatformMountain View
- Robot Software EngineerMountain View
- Robotics Software Test Engineer Mountain View
- Inference Infrastructure EngineerMountain View