
About the Role
You will design and build the core infrastructure powering AI inference across Cloudflare's global network. This includes real-time voice, frontier open LLMs, and customer-deployed models running on a heterogeneous fleet of GPUs and accelerators in hundreds of cities worldwide. You will solve complex problems in distributed systems and high-performance computing, such as sub-second model cold starts, multi-accelerator workload scheduling, and efficient KV cache management.
Responsibilities
- Platform Architecture: Develop and maintain core components of the serverless inference platform to ensure high availability and scalability.
- Optimization & Performance: Optimize model scheduling systems to increase efficiency and resource utilization; improve request routing logic to reduce latency.
- System Reliability: Drive measurable improvements in platform resilience by identifying and mitigating systemic risks.
- Observability: Refine the observability stack (metrics, logging, tracing) and fine-tune alerts to proactively resolve production issues.
- Technical Leadership: Lead complex, cross-functional technical projects from concept through deployment.
- Mentorship: Mentor junior engineers and contribute to a collaborative engineering culture.
Desirable Skills & Experience
- Proven experience in systems engineering, focusing on distributed, high-performance systems.
- Expert proficiency in Rust programming, particularly in asynchronous environments.
- Deep understanding of networking and application protocols (TCP, HTTP, WebSocket).
- Experience with scaling and performance optimization techniques, including load balancing and caching.
Bonus Points
- Demonstrable experience with container orchestration platforms, specifically Kubernetes and/or Nomad.
- Familiarity with architectural challenges in large-scale inference serving (e.g., LLMs and diffusion models).
Benefits
Open to
Austin · United States · London · United Kingdom
Sign in to track applications and earn points.