
About the Role
Canonical is looking for a Site Reliability Engineer to perfect enterprise infrastructure DevOps practices. You will work on a model-driven approach to automation, managing hundreds of private cloud and Kubernetes clusters across physical and public cloud estates.
Your work will span the entire stack, from bare-metal networking and kernel-level operations up to Kubernetes and open-source applications.
Responsibilities
- Deploy and maintain OpenStack, Kubernetes, storage solutions, and open-source applications.
- Identify and address incidents, monitor and observe applications, and anticipate potential issues.
- Apply a scientific, metrics-driven approach to automation and operations at scale.
- Collaborate with global teams to align on strategy and execution.
Requirements
- Degree in Software Engineering or Computer Science.
- Fluent in Python with professional software development experience.
- Strong operational experience in Linux environments.
- Experience with Kubernetes deployment or operations.
- Excellent interpersonal skills, curiosity, and accountability.
- Ability to travel internationally twice a year for company events (up to two weeks long).
Bonus Skills
- Familiarity with OpenStack deployment or operations.
- Experience with public cloud or private cloud management.
What We Offer
- Distributed work environment with twice-yearly in-person team sprints.
- Personal learning and development budget of USD 2,000 per year.
- Bi-annual compensation reviews and annual bonuses.
- Comprehensive benefits including holiday leave, parental leave, and Employee Assistance Programs.
- Travel perks including Priority Pass and upgrades for long-haul company events.
Benefits
Open to
Worldwide
Sign in to track applications and earn points.