Open positions

[ AI INFRASTRUCTURE ]

AI Inference Routing & Scheduling Lead

Team
AI Infrastructure
Location
Mountain View, CA
Apply for this role

About the job

Renice AI develops system and infrastructure solutions designed for the unique demands of advanced AI inference workloads. We work closely with external research, software, and hardware partners to shape the next generation of AI systems, from silicon to model weights through full-scale deployments.

This role focuses on designing and operating the routing and scheduling systems that turn large accelerator clusters into reliable, efficient token services at data center scale.

About the role

We are seeking an AI Inference Routing & Scheduling Lead to build the control plane for large-scale, multi-tenant LLM serving.

You will develop algorithms and production systems for admission control, request routing, workload scheduling, capacity allocation, fairness, and failure recovery across accelerator clusters. Your work will directly influence how we operate vLLM, SGLang, Kubernetes, and future serving runtimes as one reliable service.

This role sits at the intersection of distributed systems, online optimization, queueing theory, ML inference, and production reliability. It requires strong technical judgment, hands-on ownership, and the ability to translate mathematical models into simple, measurable production policies.

Key responsibilities

  • Build and own the production routing and scheduling control plane for multi-tenant LLM inference at data center scale.
  • Design online algorithms that balance time to first token, inter-token latency, throughput, KV-cache reuse, accelerator utilization, cost, and SLA risk.
  • Develop quantitative policies for admission control, overload protection, load shedding, cache-aware routing, priority, fairness, batching, and tenant resource allocation.
  • Define an engine-neutral telemetry and service-profile contract for vLLM, SGLang, and future serving runtimes.
  • Own routing correctness and resilience, including stale endpoint detection, draining, failover, recovery, and exactly-once streaming semantics.
  • Build trace replay, simulation, benchmarking, and shadow-decision systems to validate policies before production rollout.
  • Partner with machine learning, infrastructure, Kubernetes, networking, reliability, and security teams to translate service requirements into dependable systems.
  • Lead and grow a small team of two to three engineers, set technical direction, and maintain high standards for systems rigor and operational quality.

Qualifications

  • Experience owning or building production routing, scheduling, load-balancing, or resource-management systems at large scale.
  • Deep knowledge of distributed systems and at least one of online algorithms, queueing theory, optimization, control, or operations research.
  • Understanding of LLM inference behavior, including prefill and decode asymmetry, continuous batching, KV-cache growth and reuse, and tail-latency tradeoffs.
  • Comfort working across abstraction layers, from request-level SLAs and control-plane APIs to accelerator topology and engine behavior.
  • Strong systems programming skills in Go, Python, C++, Rust, or a comparable language, with experience debugging production performance and reliability.
  • Ability to turn ambiguous, open-ended objectives into measurable models, experiments, and production policies.
  • Clear communication and the ability to influence internal teams and external infrastructure partners.

Preferred skills

  • Experience with hyperscale ML serving, recommendation, search, ads, or cloud AI control planes.
  • Experience with vLLM, SGLang, Triton, Ray Serve, KServe, llm-d, Kubernetes, or similar inference and orchestration systems.
  • Familiarity with accelerators, HBM and KV-cache management, NVLink or NVSwitch, InfiniBand, and high-performance Ethernet.
  • A record of applied research, publications, or patents in distributed systems, scheduling, networking, ML systems, or operations research.
  • Prior experience leading or mentoring senior engineers and working across research and production teams.

About Renice AI

Renice AI builds the integrated stack for AI inference: the model, the software and the hardware designed as one system, rather than one layer at a time.

Nobody designed the AI stack as a whole. Models, software, chips, memory and networking are each owned by a different part of the industry, and each locked in choices that were rational at the time but were never made for serving. Added up, those choices cap how many tokens the world can make. We redesign across them so the machines the world already has produce more tokens in the datacenter and on the desktop.

[ APPLY ]

AI Inference Routing & Scheduling Lead

Send us your background and a link to your resume. The form opens a pre-filled email; nothing is stored on this website.

Your email client will open with these details. Review the message before sending.