
Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra - Elastic
About Elastic
Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter.
The Role
Platform Observability & Analytics runs the infrastructure that tells Elastic the truth about its own platform. The observability clusters show Cloud engineers how production is behaving right now, and the analytics pipelines show the business how the platform and the products get used over time.
This role sits on the observability side. We run 200+ hosted deployments across every supported cloud region, ingesting logs, metrics, and traces for all of Elastic Cloud, plus the SLA and SLO monitoring for ESS and Serverless. When Cloud engineering needs to know what production is doing, they're looking at something we run.
What You Will Be Doing
- Project Delivery: Owning end-to-end delivery of moderate-to-high complexity projects on the team’s roadmap, with minimal day-to-day direction.
- Infrastructure as Code: Operating and hardening shared Elastic Cloud infrastructure (ECH, ECE, and ECK) as Infrastructure as Code — writing and reviewing the Terraform, Python, and Go that other engineers depend on.
- On-Call Support: Carrying a 24/7 on-call rotation: responding to incidents, driving them to resolution, and writing clear RCAs/postmortems that lead to lasting fixes.
- Code & Design Reviews: Reviewing others’ code and designs, and being a trusted second set of eyes on production changes to critical infrastructure.
- Mentorship: Mentoring less experienced engineers, and proactively raising risks, ideas, and improvements in team discussions.
- Operational Excellence: Improving runbooks, documentation, and operational processes so the on-call load gets lighter over time.
What You Bring
- Experience: 5+ years of SRE, platform engineering, or infrastructure engineering experience.
- Terraform: Proficiency with Terraform; comfortable owning large, multi-workspace configurations in a team setting.
- Software Engineering: Strong software engineering fundamentals in Python; comfort with Go is a plus.
- Systems Knowledge: Deep Linux systems knowledge and experience operating containerized workloads in production.
- Incident Management: Experience carrying a 24/7 on-call rotation, resolving incidents under pressure, and writing RCAs that hold up under review.
- Execution: A track record of consistently delivering end-to-end projects of moderate-to-high complexity with minimal oversight.
- Security Mindset: Comfort thinking about the security implications of the infrastructure you build, defaulting to a security-conscious mindset.
- Collaboration: A pattern of mentoring less experienced engineers and speaking up with ideas and concerns in team discussions.
- Communication: Clear written and verbal communication — you document what you build and can explain it to both engineers and non-engineers.
- Time Zones: Comfort working across time zones, in both real-time and asynchronous contexts.
Bonus Points
- Experience with the Elastic Stack (Elasticsearch, Logstash, Beats, Kibana) in production.
- Experience with GitOps-style deployment tooling (ArgoCD, Helm) or policy-as-code frameworks (e.g., Kyverno) on Kubernetes.
- Experience with secrets management (Vault) or access-control/bastion tooling (Teleport).
- Experience with configuration management tools (e.g., Puppet, Ansible) at fleet scale.
- Exposure to FedRAMP, GovCloud, or other regulated/compliance-driven infrastructure.
Compensation & Benefits
- Salary Range: $128,300 — $203,000 CAD
- Equity: Eligible to participate in Elastic's stock program.
- Retirement: Registered Retirement Savings Plan (RRSP) with dollar-for-dollar matching up to 6% of eligible earnings.
- Health Coverage: Comprehensive health coverage for you and your family.
- Flexible Schedule: Ability to craft your calendar with flexible locations and schedules.
- Time Off: Generous number of vacation days each year.
- Parental Leave: Minimum of 16 weeks of parental leave.
- Charitable Giving: Up to 40 hours each year for volunteer projects and donation matching up to $2,000.
Timezone overlap
UTC-8–-4
Culture
Async-friendly
Benefits
Equity, Pension, Health, PTO, Parental leave, Commission, Bonus, Wellness
Open to
NA
Sign in to track applications and earn points.