Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
plaud · Remote
- Posted
- 107 days ago
- Last confirmed live
- 1 day ago
What this role involves
Plaud is seeking a Machine Learning Engineer focused on inference and serving for speech LLMs. The role involves building and deploying high-throughput, low-latency inference engines, optimizing GPU performance, and implementing real-time audio streaming. The position sits between the ML training and backend infrastructure teams.
Skills this posting asks for
- inference engines
- large language models
- speech models
- continuous batching
- kv cache management
- pagedattention
- gpu architectures
- nvidia ampere
- nvidia hopper
- memory hierarchy
- vllm
- tensorrt-llm
- sglang
- nvidia triton inference server
- websockets
- webrtc
- neural audio codecs
- speculative decoding
- lookahead decoding
- chunked prefill
- post-training quantization
- fp8
- int8
- awq
From the employer’s posting
ABOUT PLAUD INC. Plaud is building the world's most trusted AI work companion for professionals to elevate productivity and performance through note-taking solutions, loved by over 1,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud is building the next-genera…
Read the full description on plaud’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at plaud
- Senior Data Analyst, B2B Growth - San FranciscoRemote
- Product Designer (Contract)- San FranciscoRemote
- SRE - SeattleRemote
- Full-Stack Engineer - Palo AltoRemote
- Senior SRE Engineer - San FranciscoRemote
- Senior Data Scientist, B2B Product - San FranciscoRemote