
About the Role
GitLab is seeking Site Reliability Engineers (SREs) to join our Infrastructure Platforms team. This is a single application for SRE opportunities ranging from Intermediate to Senior Staff levels. We evaluate candidates holistically to match you with the team and level that best aligns with your experience and our current hiring needs.
What You'll Do
- Reliability & Scale: Keep user-facing services and production systems reliable, scalable, and efficient.
- Automation: Build tooling to reduce toil, replacing manual work with repeatable, infrastructure-as-code-driven workflows.
- Kubernetes Operations: Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
- Infrastructure as Code: Write and maintain IaC, shipping changes safely through CI/CD and GitOps.
- Observability: Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early.
- Incident Response: Participate in on-call rotations, triage alerts, and conduct post-incident reviews to drive process improvements.
- Documentation: Maintain runbooks and architecture decisions to ensure practices are repeatable.
What You'll Bring
- Operations Mindset: Experience keeping production systems reliable, combining operational expertise with software engineering practices.
- Tooling Experience: Proven ability to build net-new infrastructure tooling (e.g., Terraform modules, Kubernetes operators, or custom automation).
- Technical Proficiency: Ability to read, debug, and reason about code (primarily Go and Ruby).
- Cloud & K8s: Hands-on experience with major cloud providers (GCP or AWS) and deep knowledge of the Kubernetes ecosystem.
- Observability: Familiarity with metrics, logging, alerting, and SLO/SLI practices.
- Communication: Strong written communication skills and the ability to operate as a 'manager-of-one' in an asynchronous, distributed environment.
- Growth Mindset: A track record of using automation and AI to improve team efficiency.
Hiring Process
Our process is designed to evaluate you once for multiple potential teams:
- Recruiter Screen: Background and alignment discussion.
- Core Technical: Collaborative assessment on system architecture and incident review.
- Hiring Manager Interview: Focus on ownership, judgment, and collaboration.
- Peer Technical: Team-specific depth interview.
- Skip-Level Interview: Values alignment and cross-team strategy.
Note: This position is open to candidates based in the United States and Canada only.
Timezone overlap
UTC-8–-4
Open to
NA
Sign in to track applications and earn points.