
Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end business processes. Trusted by over 700 organizations worldwide, including 9 of the top 10 US banks, Camunda helps enterprises boost operational efficiency, accelerate time-to-value, and deliver better customer experiences.
Fully remote and global, we are transforming into an AI-first organisation, built on our own platform. We use Agentic AI to automate, orchestrate intelligent processes, and elevate human contribution across every team.
About the Role
If you’re passionate about pushing the limits of reliability, fault-tolerance, and operational excellence in complex distributed systems, Camunda offers the perfect stage.
As a Senior Software Engineer, Backend, you’ll take ownership of automated reliability testing and chaos engineering at the core of Camunda 8. This role thrives on curiosity, experimentation, and relentless improvement—you’ll break things on purpose in safe environments to strengthen our platform before customers ever feel the impact.
What You’ll Be Doing
- Design, implement, and execute automated chaos experiments and reliability tests, validating Camunda's platform under real-world scenarios.
- Investigate, root-cause, and debug potential failures or performance regressions of our Java-based products.
- Continuously improve existing load testing, chaos engineering, and observability infrastructure using Java, Go, Kubernetes, Prometheus, Grafana, etc.
- Introduce new tooling and approaches for reliability testing, driving measurable improvements to operational experience and user outcomes.
- Collaborate closely with QA and cross-functional engineering teams, sharing learnings and influencing team roadmaps based on experiment findings.
- Champion a pragmatic, autonomous approach to software design and advocate for user-centric reliability.
Defining Success
- After 3 months: Designed and delivered a way to run quick and reproducible load tests to reduce feedback loops for engineers and fit directly into the development lifecycle.
- After 3 months: Designed and implemented a realistic, automated, holistic load testing framework for Optimize in close collaboration with Senior Engineers and Field teams.
What You Bring
- Ability and/or willingness to use our product.
- 5+ years of experience in backend software engineering (Java).
- Proven drive to experiment, learn new tech, and conduct automated reliability and chaos testing in distributed environments.
- Deep enthusiasm for improving performance and fault-tolerance in production systems.
- Autonomous, pragmatic problem-solving skills with an ability to guide and influence others.
Nice to Haves
- Hands-on experience with Go or other languages in a multi-language context.
- Working as an SRE on distributed systems with a strong software engineering mindset.
- Working knowledge of Kubernetes, Helm Charts, Operators, and related production infrastructure.
- Experience running apps in production, with solid skills in monitoring, troubleshooting, and performance analysis.
- Background in chaos engineering, automated reliability, load, or performance testing.
What We Offer
- Remote & Flexible: Work from anywhere with a home office budget, co-working space support, and flexible time off.
- In-Person Connection: Annual kickoff, team offsites, and meetup budgets.
- Health & Wellbeing: Locally tailored healthcare, Modern Health mental wellbeing support, and the Live Well LSA.
- Financial Security: Retirement and pension plans plus life and disability insurance.
- Professional Growth: Up to $/€/£1,000 per year for self-driven learning.
Timezone overlap
UTC-8–+3
Benefits
Equity, Health, Mental health, Pension, PTO, Parental leave, Learning, Home office, Coworking, Wellness
Open to
Europe · UK · NA · US
Sign in to track applications and earn points.