Skip to content

Command Palette

Search for a command to run...

EngineeringRemote (US/EU)Full-time

Senior ML Engineer, Inference

Our inference platform serves 48 million requests a day with an 87ms average first token. You'll own the systems that keep that number falling while models — and traffic — keep growing. This is deep infrastructure work with a direct line to every user's experience.

What you'll do

  • Design and operate GPU inference clusters across 32 edge regions
  • Drive latency work: speculative decoding, KV-cache strategies, smart batching
  • Build the model-routing layer that picks the best model per request
  • Own capacity planning and cost efficiency for a multi-tenant fleet
  • Partner with research to productionize new models within days, not months

What we're looking for

  • 5+ years building distributed systems, 2+ with ML inference in production
  • Fluency in Python and one systems language (Go, Rust or C++)
  • Hands-on experience with CUDA, TensorRT, vLLM or similar serving stacks
  • Strong grasp of queueing, autoscaling and observability at scale
  • Clear written communication for RFC-driven, async decision making

Nice to have

  • Experience with multi-region traffic engineering
  • Contributions to open-source inference tooling
  • Familiarity with speech or diffusion model serving
Apply

Apply for Senior ML Engineer, Inference

Five fields, five minutes. We review every application by hand and reply within a week — usually faster.

By applying you agree to our candidate privacy policy. We'll only use your details for this application.