
About the Role
GitLab is seeking a Senior Site Reliability Engineer to join the Observability, Monitoring, and Integrations team within our Monetization section. This is a greenfield opportunity to build and operate telemetry, detection, and reconciliation tooling that protects revenue and customer experience across our core systems, including CustomersDot, Salesforce, and Zuora.
What You’ll Do
- Design & Build: Develop metrics, logs, and traces across the Monetization stack using Prometheus, Grafana, and OpenTelemetry.
- Anomaly Detection: Implement automated detection for billing, data, and event anomalies, routing alerts to relevant feature teams.
- Data Integrity: Develop reconciliation checks across usage and billing pipelines.
- Reliability Engineering: Define and track SLOs/SLIs, write runbooks, and participate in incident management.
- Innovation: Explore AI/ML techniques to predict system anomalies and accelerate resolution.
- Collaboration: Review merge requests and partner with Product, Finance, and Support teams to turn operational needs into reliable tooling.
What You’ll Bring
- Professional experience with Ruby on Rails.
- Strong background in SRE or observability (monitoring, alerting, SLOs, incident handling).
- Experience with Prometheus, Grafana, and OpenTelemetry.
- Experience building anomaly detection or risk management tooling.
- Familiarity with data stores (especially ClickHouse) and event streaming pipelines (e.g., NATS JetStream).
- Experience with billing, financial, or business-critical systems.
- Ability to own projects from concept to production in a remote, asynchronous environment.
- Strong communication skills in English.
How GitLab Supports You
- Flexible Paid Time Off
- Comprehensive health, financial, and well-being benefits
- Equity Compensation & Employee Stock Purchase Plan
- Growth and Development Fund
- Parental Leave
- Team Member Resource Groups
Timezone overlap
UTC+8–+12
Benefits
Open to
APAC
Sign in to track applications and earn points.