
Who We Are
AI is changing how software gets built. Code production is becoming a commodity, shifting the focus from writing code to specifying, orchestrating, verifying, and governing change. At Develocity, we build a toolchain observability and intelligence platform used by leading software organizations like Netflix, Airbnb, Spotify, SAP, and major global banks.
We are an AI-native company, embedding context engineering and ground truth directly into our platform to help teams achieve delivery excellence. We have partnered with the Apache Software Foundation, the Commonhaus Foundation, the Micronaut Foundation, and other major OSS projects to bring these values to the open-source community.
Our Values
- Seek to Understand: Everything starts with listening and understanding diverse viewpoints, problems, and motivations before taking action.
- Know the Why: Approach work with purpose, urgency, and thoughtful consideration.
- Innovate & Iterate: Embrace challenges, try new things, and develop creative solutions.
- Own the Outcome: Take initiative, maintain transparency, and assume responsibility for decisions and results.
Who You Are
We're building a new SRE team and looking for founding members to help shape how we operate. As a Lead SRE, you’ll be a technical and operational leader for reliability across Develocity, defining our SRE vision, setting operational standards, and mentoring other engineers.
You will be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, as well as supporting infrastructure like artifact registries.
You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and troubleshoot issues across the stack. You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation.
Responsibilities
- Operate and maintain all Develocity instances and supporting services in production.
- Define and evolve SRE standards, practices, and operating models, including on-call, incident response, postmortems, and SLOs.
- Participate in an on-call rotation, acting as a technical escalation point for complex or high-severity incidents.
- Lead incident response and blameless retrospectives, ensuring measurable reliability improvements.
- Set reliability priorities using risk, customer impact, business goals, SLOs, and error budgets.
- Identify systemic reliability risks and continuously evolve SaaS operations as the platform scales.
- Lead architectural and design reviews to ensure reliability, scalability, and operability.
- Drive automation across deployment, upgrades, monitoring, self-healing, recovery, and operational workflows.
- Build and maintain comprehensive observability for all managed services, including logging, metrics, tracing, and alerting.
- Own disaster recovery, backups, and business continuity planning and execution.
- Partner with engineering leadership to balance feature delivery with operational excellence.
- Mentor and coach SREs, supporting technical growth and strong operational practices.
- Help onboard new SREs and contribute to hiring by defining SRE excellence.
- Communicate clearly with customers during incidents and maintenance windows.
- Optimize performance, resource utilization, and operational costs.
Minimum Qualifications
- 7+ years in SRE, DevOps, or an equivalent role operating production services at scale.
- Experience leading reliability initiatives across multiple teams or services.
- Demonstrated ability to influence technical direction without direct authority.
- Experience designing and operating systems with SLOs and error budgets.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
- Track record of incident management and response in a 24/7 on-call environment.
- Scripting proficiency (Python, Bash) for automation.
- Strong written and verbal English communication skills.
Preferred Qualifications
- Experience as a founding or early SRE establishing practices in a growing SaaS organization.
- Familiarity with Develocity.
- JVM language experience (Java, Kotlin).
- Experience with customer-facing and executive-level incident communications.
What We Offer
- A ground-floor role in a new SRE team where you will shape operational practices.
- Real ownership of production systems used by elite engineering teams.
- Direct interaction with customers during incidents and successes.
- A culture that values automation over heroics.
- In-person meetings, including annual company offsites and team gatherings.
- Work from home in a remote-first environment.
- Competitive salaries and equity grants.
Timezone overlap
UTC+0–+3
Culture
Async-friendly
Open to
Europe
Sign in to track applications and earn points.