EngineeringRemote (US/EU)Full-time
Senior ML Engineer, Inference
Our inference platform serves 48 million requests a day with an 87ms average first token. You'll own the systems that keep that number falling while models — and traffic — keep growing. This is deep infrastructure work with a direct line to every user's experience.
What you'll do
- Design and operate GPU inference clusters across 32 edge regions
- Drive latency work: speculative decoding, KV-cache strategies, smart batching
- Build the model-routing layer that picks the best model per request
- Own capacity planning and cost efficiency for a multi-tenant fleet
- Partner with research to productionize new models within days, not months
What we're looking for
- 5+ years building distributed systems, 2+ with ML inference in production
- Fluency in Python and one systems language (Go, Rust or C++)
- Hands-on experience with CUDA, TensorRT, vLLM or similar serving stacks
- Strong grasp of queueing, autoscaling and observability at scale
- Clear written communication for RFC-driven, async decision making
Nice to have
- Experience with multi-region traffic engineering
- Contributions to open-source inference tooling
- Familiarity with speech or diffusion model serving
Apply
Apply for Senior ML Engineer, Inference
Five fields, five minutes. We review every application by hand and reply within a week — usually faster.