
Who We Are
AI is changing how software gets built. Code production is becoming a commodity, shifting focus from writing code to specifying, orchestrating, verifying, and governing change. We build Develocity, a toolchain observability and intelligence platform used by some of the world's leading software organizations—including Netflix, Airbnb, Spotify, SAP, and major global banks—to help software teams achieve delivery excellence.
Who You Are
We're building a new SRE team and looking for founding members to help shape how we operate. You'll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries.
You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure.
Responsibilities
- Operate and maintain all Develocity instances and supporting services.
- Participate in an on-call rotation, owning incident response and troubleshooting issues across the stack.
- Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
- Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
- Work with engineering teams to build reliability into features from the start.
- Run incident response and retrospectives, and ensure we learn from them.
- Own disaster recovery, backups, and business continuity.
- Communicate with customers during incidents and maintenance windows.
- Optimize performance, resource usage, and costs.
Minimum Qualifications
- 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
- Track record of incident management and response.
- Knowledge of SRE best practices (SLAs, SLOs).
- Scripting proficiency (Python, Bash) for automation.
- Experience with 24/7 on-call rotations.
- Strong written and verbal English communication.
Preferred Qualifications
- Experience operating SaaS platforms at scale.
- Familiarity with Develocity.
- JVM language experience (Java, Kotlin).
- Disaster recovery planning and execution experience.
- Customer-facing incident communication skills.
- Experience establishing SRE practices in new or growing teams.
What We Offer
- A ground-floor role in a new SRE team to shape how operations are run.
- Real ownership of production systems used by major software engineering teams.
- A culture that values automation over heroics.
- Annual company offsites and team meetings.
- Remote-first work environment.
- Competitive salaries and equity grants.
Timezone overlap
UTC-8–-5
Culture
Async-friendly
Benefits
Open to
NA
Sign in to track applications and earn points.