GitLab logo
GitLab·Verified

Principal Site Reliability Engineer, Platform Engineering - GitLab

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation.

An overview of this role

We’re looking for a Principal Engineer with deep expertise in Site Reliability, Backend, or Platform Engineering to help shape the next phase of GitLab Dedicated, our fully managed single-tenant SaaS offering. This highly influential technical leadership role will set direction for how we scale a growing fleet of isolated, customer-specific environments while maintaining the reliability, security, and compliance our customers depend on.

You’ll lead platform and operating-model transformation across resilience and failover, tenant orchestration, change management, self-service tooling, and platform integrations. You’ll also help align Dedicated with GitLab’s evolution toward more modular and cell-based architectures, strengthen service ownership, and establish scalable patterns that reduce operational complexity. As a Principal Engineer, you’ll influence across teams, guide complex technical decisions, mentor senior engineers, and help raise the technical maturity of the platform as we scale.

What you will do

  • Set technical direction for GitLab Dedicated, shaping architecture and platform strategy as we scale a growing fleet of isolated, single-tenant environments.
  • Lead platform transformations across resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations.
  • Drive scalable, modular architecture that aligns Dedicated with GitLab’s broader Cells strategy while preserving its security, isolation, and compliance requirements.
  • Strengthen service ownership and operational maturity, helping engineering teams build, operate, and improve the production systems they own.
  • Identify and address systemic reliability and scalability risks using production signals, incident patterns, and architectural insight.
  • Establish reusable platform patterns and automation that reduce operational toil and allow Dedicated to scale efficiently.
  • Lead complex technical decisions across teams, balancing reliability, security, cost, maintainability, and customer needs.
  • Advance engineering excellence across the organization through architectural leadership, mentorship, and influence with senior engineers and engineering leaders.

What you will bring

  • Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering, with experience designing and operating large-scale production systems.
  • Hands-on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices.
  • Strong software engineering fundamentals, with experience building production systems or infrastructure tooling in languages such as Go, Ruby, Python, or similar.
  • Strong distributed systems and systems-design expertise, with sound judgment around reliability, failure isolation, scalability, and operational complexity.
  • A track record of technical leadership across multiple teams, setting direction and driving complex initiatives through influence.
  • Experience leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems through major growth.
  • Experience leading changes that improve how engineering teams own and operate production systems, strengthening reliability, operational readiness, and accountability at scale.
  • Exceptional technical communication and influence, with the ability to build alignment, mentor senior engineers, and guide complex architectural decisions.

Timezone overlap

UTC-8–+1

Culture

Async-friendly

Open to

NA · UK

Sign in to track applications and earn points.

More roles at GitLab

Similar remote roles