GitLab logo
GitLabΒ·Verified

Engineering Manager, Observability - GitLab

Overview

GitLab is seeking an Engineering Manager to lead our globally distributed Observability team within Production Engineering. The team builds and operates the metrics, logging, alerting, and capacity planning platforms that GitLab engineers use to understand GitLab.com and GitLab Dedicated.

In this role, you will help determine how the team collects, stores, queries, and acts on telemetry, balancing reliable signals with scale and cost. You will guide improvements to Prometheus-based metrics pipelines, log ingestion and retention, SLO-driven alerting, and capacity forecasting while participating in incident response and fostering a sustainable on-call environment.

What You'll Do

  • Lead & Develop: Lead, hire, onboard, and develop a distributed engineering team working asynchronously.
  • Set Strategy & Priorities: Collaborate with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams to deliver observability services iteratively.
  • Platform Ownership: Own the reliability, scalability, and cost of metrics, logging, alerting, and capacity planning platforms.
  • Improve Observability: Reduce telemetry gaps and noisy or missing alerts using SLOs, error budgets, and self-service instrumentation.
  • Technical Guidance: Guide technical architecture and decisions around time-series storage, high-cardinality metrics, log pipelines, and distributed tracing.
  • Incident Response: Participate in the Incident Manager On Call (IMOC) rotation coordinating high-severity incident responses.
  • Sustainable Operations: Maintain a healthy on-call rotation through balanced time-zone coverage, clear runbooks, and thorough post-incident follow-ups.
  • AI Integration: Leverage AI tools and agents to support operational workflows and incident triage.

What You'll Bring

  • Experience leading an observability, platform engineering, or SRE team operating at scale within a distributed, asynchronous environment.
  • Deep technical knowledge of metrics systems (such as Prometheus and long-term storage), logging platforms (Elasticsearch or cloud-native solutions), and alerting design.
  • Track record of applying SLOs, error budgets, and capacity forecasts to drive reliability and investment decisions.
  • Hands-on experience operating high-scale SaaS platforms and resolving telemetry bottlenecks, ingestion limits, and alert fatigue.
  • Experience participating in and refining production on-call rotations and incident coordination.
  • Strong communication skills to clearly articulate technical tradeoffs to partners and stakeholders.
  • Familiarity with leveraging AI tools or agents for engineering workflows and incident triage.

What GitLab Offers

  • Comprehensive benefits supporting health, finances, and well-being
  • Flexible Paid Time Off (PTO)
  • Equity compensation and Employee Stock Purchase Plan (ESPP)
  • Growth and Development Fund for continuous learning
  • Paid parental leave
  • Team Member Resource Groups and an inclusive, all-remote culture

Timezone overlap

UTC+8–+12

Culture

Async-friendly

Open to

APAC

Sign in to track applications and earn points.

More roles at GitLab

Similar remote roles