
Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end business processes. With built-in governance, auditability, and human oversight, Camunda gives enterprises the control they need to move AI from pilots to production — safely and at scale.
About the Role
We're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'll design and maintain our Kubernetes-based multi-cloud platform, improve our monitoring and observability tools, and collaborate with product and engineering teams to solve real problems in real time. This is a role where you'll own the systems that power Camunda, drive automation that raises the bar for everyone, and mentor others who want to do the same.
What You’ll Be Doing
- Design and maintain our infrastructure – Evolve our Kubernetes-based, multi-cloud platform architecture, ensuring it's available, scalable, and fault-tolerant. Establish configuration best practices and network services that our teams rely on.
- Build observability that matters – Implement and improve monitoring and alerting tools that give both SREs and developers real visibility into system health and performance. Make it easy for teams to understand what's happening across our stack.
- Own your systems end-to-end – Adopt a "you build it, you run it" mentality, participating in on-call rotations and being the go-to person when things need quick fixes. Create runbooks and automation that turn complex problems into manageable processes.
- Ship features and improvements with product teams – Work cross-functionally with product engineering, product management, and support to define and deliver features that move the needle.
- Push automation to the next level – Identify repetitive work and automate it away. Share what you learn with your teammates so everyone gets better at their craft.
- Be the expert others learn from – Help less experienced engineers tackle complex infrastructure challenges. Break down technical problems into clear steps and support others in growing their skills.
What You Bring (Must Haves)
- Deep hands-on experience with Kubernetes – You've built, deployed, and maintained Kubernetes clusters in production environments. You understand how to manage workloads, networking, and storage at scale.
- Infrastructure as code expertise – You're fluent in tools like Terraform (or similar IaC tools) and know how to version, test, and safely deploy infrastructure changes.
- Demonstrated experience in monitoring and observability – You've worked with tools like Prometheus, Grafana, or similar platforms to instrument systems and alert on what matters.
- Strong 3rd-level support and incident response skills – You've responded to production incidents, diagnosed complex issues, and communicated clearly with stakeholders under pressure. You understand root cause analysis and how to prevent issues from happening again.
- A passion for automation and raising the quality bar – You see manual work as a problem to be solved. You care deeply about building systems that are reliable, maintainable, and easy to understand.
- Responsible use of AI tools – You leverage AI for research, code review, documentation, test generation, and automation to improve your effectiveness while retaining human accountability.
Nice-to-Haves
- Experience with major cloud providers (AWS/EKS, Google Cloud Platform/GKE, or similar managed Kubernetes services).
- ArgoCD or GitOps workflows.
- Proficiency in Python, Go, or similar languages.
- Experience with SLOs and alerting frameworks.
Compensation & Benefits
Compensation
Salary ranges are location-based. The Annual Total Target Cash (base salary + 100% variable target, where applicable) includes:
- United States: $149,800.00 to $241,500.00
- United Kingdom: £94,100.00 to £154,700.00
- Singapore: S$186,100.00 to S$279,100.00
- Canada: C$159,100 to C$261,600
- If you’re based elsewhere, you’ll be hired via Remote.com.
Benefits & Perks
- Equity: Offered through our Virtual Stock Option Plan (VSOP).
- Remote & Flexible: Work from anywhere with a home office budget, co-working space support, and flexible time off.
- In-Person Connection: Annual Kickoff events, team offsites, and Camundi Connection Budgets.
- Health & Wellbeing: Locally tailored healthcare, Modern Health for global mental wellbeing, and our Live Well Lifestyle Spending Account (LSA) scaling to €1,000 annually from 2027.
- Financial Security: Retirement and pension plans (with company contributions where relevant), plus life and disability insurance.
- Professional Growth: Up to $/€/£1,000 per year for self-driven learning (courses, certifications, books).
Benefits
Equity, Health, Mental health, Pension, PTO, Learning, Wellness, Home office, Coworking
Open to
Worldwide
Sign in to track applications and earn points.