
Runpod is the AI Developer Cloud, empowering more than one million developers to experiment, train, fine-tune, deploy, and scale AI on a single platform. We're looking for a Director of Infrastructure Engineering to lead and scale Runpod’s core cloud and bare-metal environments.
This role owns the critical foundational layers of our platform—Site Reliability Engineering (SRE), global networking, High-Performance Computing (HPC) networks, and distributed storage engines. You’ll build the operating rhythm, culture, and technical direction that ensures Runpod remains highly available, performant, and capable of scaling to meet massive GPU computing demands.
Responsibilities
- Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.
- Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads.
- Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks.
- Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment.
- Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes.
- Continuously Improve Systems & Flow: Drive measurable improvements in infrastructure reliability and delivery metrics, such as deployment frequency, MTTR, infrastructure as code (IaC) coverage, and system uptime.
- Architectural Stewardship: Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters, ensuring seamless scalability.
- Cross-Functional Partnership: Coordinate cleanly with product delivery and platform teams to ensure infrastructure primitives are robust, well-documented, and highly available.
Requirements
- Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads.
- Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
- HPC & Advanced Networking: Strong architectural understanding of ultra-low latency networking, InfiniBand, RoCE, spine-leaf architectures, and BGP.
- Storage Systems Knowledge: Experience building or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) for AI/ML I/O loads.
- SRE / DevOps Culture: Strong foundation in reliability engineering, Terraform, Ansible, Kubernetes, and modern observability stacks.
- Remote-First Operating Excellence: Experience building culture, accountability, and momentum across distributed technical teams.
- Communication & Collaboration: Clear communication, strong stakeholder management, and decisive leadership during high-stakes incidents.
- Eligibility: Eligible to work in the United States (visa sponsorship is not available).
What We Offer
- Competitive base salary range of $225,000 - $325,000.
- Meaningful equity / stock options.
- Generous medical, dental & vision plans.
- Flexible PTO policy.
- $1,200 home office & equipment stipend.
Timezone overlap
UTC-8–-4
Benefits
Equity, Health, Dental, Vision, PTO, Home office, Equipment, Unlimited PTO, Visa
Open to
US
Sign in to track applications and earn points.