
About the Role
Grafana Labs is seeking a Staff Software Engineer (SRE) to enhance the reliability of our Cloud databases, including Mimir, Loki, Tempo, and Pyroscope. You will be embedded within product squads to ensure our SaaS offerings deliver exceptional performance for high-SLA customers.
Key Responsibilities
- Reliability Ownership: Own production reliability for high-SLA and complex customer environments.
- Embedded Collaboration: Partner closely with product engineering squads to influence feature design for scalability and operability.
- Automation: Build tools to eliminate toil, scale reliability practices, and improve alert quality.
- Incident Management: Serve as a primary escalation point, lead incident responses, and conduct thorough post-incident reviews (PIRs).
- SLO Management: Define and evolve per-tenant SLOs and proactively reduce SLO burn.
What We Seek
- 8+ years of engineering experience, with at least 4 years in SRE/CRE or production engineering.
- Strong Kubernetes expertise in AWS, GCP, or Azure.
- Proficiency in infrastructure-as-code (Helm, Terraform, Jsonnet).
- Experience operating multi-tenant systems in production.
- Proficiency in one or more programming languages (e.g., Go, Python, Java).
- Strong technical leadership skills, including mentoring and serving as a force-multiplier.
- Excellent problem-solving skills and experience with blame-free incident response.
Why Join Us?
- Remote-First: Work from anywhere within the UK, Sweden, Spain, or Germany.
- Innovation-Driven: Access to modern AI coding assistants and frontier models to accelerate your workflow.
- Global Culture: Join a 100% remote company with team members across 40+ countries.
- Benefits: Includes equity, bonus, and a global annual leave policy of 30 days.
Timezone overlap
UTC+0–+3
Open to
Europe
Sign in to track applications and earn points.