Runpod logo
Runpod·

Datacenter Infrastructure Specialist - Runpod

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. We closed a $100M Series A in June 2026, and we're building the platform the next generation of developers will depend on.

We are looking for a Datacenter Infrastructure Specialist to be the operational linchpin of our global fleet. Reporting to the Manager of Infrastructure Capacity & Management, you will serve as the technical authority bridging our hardware partners and internal engineering teams.

As our Datacenter Infrastructure Specialist, you will own the technical lifecycle and operational health of Runpod’s rapidly expanding, high-density GPU fleet. You will work directly with cutting-edge GPU cloud technologies, advanced RDMA fabrics, and modern observability stacks to solve complex hardware challenges at scale.

Responsibilities

  • Hardware Validation & Benchmarking: Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads.
  • Uptime & SLA Enforcement: Monitor fleet health to identify performance degradation. Audit downtime and provide technical data to protect customer SLAs.
  • AI-Driven Operations: Work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet.
  • Incident Support: Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions.
  • Partner Technical Support: Support the growth of our infrastructure partners.

Requirements

  • Professional Background: 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
  • Datacenter Networking: Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred.
  • GPU & AI Stack: Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning.
  • Systems & Diagnostics: Solid Linux system administration skills and experience with containerization (Docker). Comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
  • Effective Communication: Clear written and verbal communication skills to explain hardware or networking issues to technical partners and internal leadership.
  • Operational Flexibility: Willingness to participate in an on-call rotation as our global fleet scales.
  • Strategic Problem-Solver: Detail-oriented and proactive in identifying potential failures before they impact customers.

Preferred Qualifications

  • Experience working in a fast-paced startup environment building operational workflows.
  • Experience managing or optimizing bare-metal High-Performance Computing (HPC) environments at scale.
  • Experience with observability tools like Grafana, Prometheus, or Datadog.
  • Proficiency in Python, Go (Golang), or Bash for infrastructure automation.

What We Offer

  • Competitive base pay ranging from €105,324.00 to €140,432.00.
  • Meaningful equity in a fast-growing AI infrastructure company.
  • Comprehensive medical, dental, and vision plans (100% covered for employees).
  • Flexible PTO policy.
  • $1,200 Home Office & Equipment Stipend.
  • Remote-first collaborative team culture.

Timezone overlap

UTC+0–+3

Culture

Async-friendly

Open to

Europe

Sign in to track applications and earn points.

More roles at Runpod

Similar remote roles