
About the Role
GitLab is seeking Site Reliability Engineers (SREs) to join our Infrastructure Platforms team. This is a single application for roles ranging from Intermediate to Senior Staff. We evaluate candidates holistically to match you with the team and level that best aligns with your experience and our current hiring needs.
As an SRE at GitLab, you will combine software engineering with operational excellence to keep our user-facing services and production systems reliable, scalable, and efficient.
What You'll Do
- Reliability & Scale: Maintain the reliability, scalability, and efficiency of GitLab’s production systems.
- Automation: Build tooling and infrastructure-as-code to reduce toil and replace manual processes.
- Kubernetes Operations: Operate and troubleshoot production systems on Kubernetes, including deployments and scaling.
- Observability: Contribute to our observability stack using metrics, logs, and SLOs to detect issues proactively.
- Incident Response: Participate in on-call rotations, triage alerts, and conduct post-incident reviews to drive systemic improvements.
- Documentation: Maintain runbooks and architecture decisions to ensure practices are repeatable.
What You'll Bring
- Operations Mindset: Experience keeping production systems reliable with a focus on software engineering practices.
- Tooling Experience: Proven ability to build net-new infrastructure tooling (e.g., Terraform modules, Kubernetes operators, or custom automation).
- Coding Proficiency: Ability to read, debug, and reason about code (primarily Go and Ruby).
- Cloud & K8s: Hands-on experience with major cloud providers (GCP or AWS) and deep knowledge of the Kubernetes ecosystem.
- Observability: Familiarity with metrics, logging, alerting, and SLO/SLI practices.
- Communication: Strong written communication skills and the ability to operate effectively in an asynchronous, distributed environment.
Hiring Process
- Recruiter Screen: Background and alignment discussion.
- Core Technical: Collaborative assessment on system architecture and incident review.
- Hiring Manager Interview: Discussion on ownership, judgment, and growth.
- Peer Technical: Team-specific deep dive.
- Skip-Level Interview: Values alignment and cross-team collaboration.
Benefits
- Flexible Paid Time Off
- Equity Compensation & Employee Stock Purchase Plan
- Growth and Development Fund
- Parental Leave
- Comprehensive health and well-being support
Timezone overlap
UTC+0–+1
Benefits
Open to
UK
Sign in to track applications and earn points.