
Who We Are
AI is changing how software gets built. Code production is becoming a commodity, shifting focus from writing code to specifying, orchestrating, verifying, and governing change—with the toolchain as the new constraint.
We build Develocity, a toolchain observability and intelligence platform used by leading software organizations like Netflix, Airbnb, Spotify, SAP, and major global banks. Develocity helps software teams achieve delivery excellence through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain (Gradle Build Tool, Apache Maven™, sbt, npm, Python, and Bazel).
As an AI-native company, AI is central to how we work, our product, and our roadmap. We encode a decade of build and test expertise into a context engineering layer: the ground truth AI agents need to change code safely, and the governance to gate that change at machine speed. We have partnered with the Apache Software Foundation, Commonhaus Foundation, Micronaut Foundation, and OSS projects like Spring, Quarkus, Kotlin, JUnit, and AndroidX.
Our Values
- Seek to Understand: Everything starts with listening and striving to understand different viewpoints, problems, and motivations before taking action.
- Know the Why: Approach work with a clear sense of purpose, balancing urgency with thoughtful consideration.
- Innovate & Iterate: Embrace challenges, try new things, and develop creative, bold solutions.
- Own the Outcome: Take initiative, maintain transparency, take responsibility for decisions, and measure success.
Who You Are
We're building a new SRE team and looking for founding members to shape how we operate. You'll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, as well as supporting infrastructure like artifact registries.
You'll work on our internally-built Cloud Application Platform and Kubernetes on AWS. When incidents happen, you'll troubleshoot issues across the stack and collaborate with engineering teams to build reliability into how we ship software. If you love automation and hate doing manual tasks twice, you will fit right in.
Responsibilities
- Operate and maintain all Develocity instances and supporting services.
- Participate in an on-call rotation, owning incident response and troubleshooting across the stack.
- Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
- Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
- Work with engineering teams to build reliability into features from the start.
- Run incident response and retrospectives, ensuring continuous learning.
- Own disaster recovery, backups, and business continuity.
- Communicate with customers during incidents and maintenance windows.
- Optimize performance, resource usage, and costs.
- Help evolve our SaaS operations as we grow.
Minimum Qualifications
- 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
- Track record of incident management and response.
- Knowledge of SRE best practices (SLAs, SLOs).
- Scripting proficiency (Python, Bash) for automation.
- Experience with 24/7 on-call rotations.
- Strong written and verbal English communication.
Preferred Qualifications
- Experience operating SaaS platforms at scale.
- Familiarity with Develocity.
- JVM language experience (Java, Kotlin).
- Disaster recovery planning and execution experience.
- Customer-facing incident communication skills.
- Experience establishing SRE practices in new or growing teams.
What We Offer
- A ground-floor role in a new SRE team where you shape the practices.
- Real ownership of production systems used by elite engineering teams.
- Direct interaction with customers during incidents and successes.
- A culture that values automation over heroics.
- Remote-first environment with home-office support.
- Competitive salaries and equity grants.
- Regular in-person meetings, including annual company offsites and team gatherings.
Timezone overlap
UTC+0–+3
Culture
Async-friendly
Benefits
Open to
Europe
Sign in to track applications and earn points.