Wikimedia Foundation logo
Wikimedia Foundation·Verified

Senior Site Reliability Engineer, Wikimedia Enterprise - Wikimedia Foundation

Remote-firstFull-timeSenior$117K - $181KAsyncWorldwide#aws#kubernetes#terraformVision

Summary

The Wikimedia Foundation is looking for a Senior Site Reliability Engineer to join our team, reporting to the Sr. Engineering Manager. As the Site Reliability Engineer, you will play a key role in designing, developing, and maintaining reliable, scalable, and highly available infrastructure for our API services. You will contribute heavily to the high impact challenges behind innovating, building, and maintaining Wikipedia’s data feeds for high volume reusers. In this role, you will foster cross-department collaboration with the Wikimedia Foundation SRE teams. You will own reliability targets (SLOs) for critical APIs, balancing performance, cost, and availability through data-driven decisions.

You will be involved in designing and running the infrastructure and services that interact with the base of Wikimedia Foundation’s projects, including Kubernetes clusters, application servers, code collaboration infrastructure, and other developer-facing services. You will participate in incident response and on-call rotation. This role requires frequent collaboration with enterprise and Foundation SRE team members, as well as teams in Security, Release, and Software Engineering.

Wikimedia Enterprise is a revenue-generating product that provides fast, comprehensive, reliable, and secure data ingestion for organizations that wish to repurpose Wikimedia/Wikipedia content in third-party environments. We operate like a startup within the Wikimedia Foundation: building quickly, deploying often, and creating high impact for the global knowledge ecosystem.

Responsibilities

  • Define, track, and improve Service Level Objectives (SLOs), SLIs, and error budgets to ensure reliability targets are met.
  • Build and enhance observability systems (metrics, logs, and distributed tracing) to enable proactive detection and faster troubleshooting.
  • Drive reliability engineering practices, including capacity planning, load testing, and resilience validation (e.g., chaos testing).
  • Improve developer experience (DevEx) by enabling self-service infrastructure and streamlining deployment workflows.
  • Partner with engineering team members to embed reliability best practices early in the development lifecycle.
  • Design, implement, and optimize CI/CD and GitOps workflows using tools such as GitLab and ArgoCD, enabling automated, reliable deployments with support for progressive delivery strategies (canary, blue-green).
  • Implement secure-by-default infrastructure and enforce best practices (IAM, secrets management, encryption).
  • Continuously optimize infrastructure cost and efficiency using FinOps principles while maintaining performance and availability.
  • Establish and track operational metrics such as MTTR, MTTD, and incident frequency to drive continuous improvement.
  • Reduce operational toil by identifying repetitive work and implementing automation-first solutions.
  • Contribute to and evolve internal platform capabilities that standardize infrastructure and improve scalability across teams.
  • Collaborate asynchronously with a globally distributed team.
  • Mentor peers in your areas of technical and operational strength.

Skills & Experience

  • Automation & Configuration Management: Experience with Infrastructure as Code and automation tools (e.g., Terraform, Ansible) and proficiency in at least one programming language (e.g., Python, Go).
  • Cloud Infrastructure: Experience designing, operating, and optimizing cloud-based systems across platforms such as AWS, Azure, or GCP, including scalability, reliability, and cost efficiency.
  • CI/CD & Deployment Practices: Experience building and maintaining CI/CD pipelines and GitOps workflows (e.g., GitLab, ArgoCD), with familiarity in progressive delivery approaches.
  • Incident Management & Reliability Operations: Experience with incident response, on-call practices, and leading postmortems with a focus on continuous improvement.
  • SRE Principles & Observability: Strong understanding of SRE best practices (SLOs, SLIs, error budgets) and experience with observability tools (metrics, logging, distributed tracing e.g., Prometheus, OpenTelemetry).
  • Collaboration & Communication: Ability to work effectively in a distributed, cross-functional environment with strong documentation and communication skills.
  • AWS Infrastructure: Passionate about building and maintaining reliable, scalable infrastructure on AWS.

Qualities That Are Important to Us

  • Proven experience operating highly available, large-scale distributed systems.
  • Ownership mindset: End-to-end responsibility for system reliability, proactively identifying and addressing risks.
  • Bias for automation: Continuously seeking to reduce operational toil through automation.
  • Continuous improvement mindset: Actively learning from incidents through blameless postmortems.
  • Customer and reliability focus: Prioritizing user experience by balancing availability, performance, and cost.
  • Adaptability: Comfortable working in a fast-evolving environment.

Nice to Have

  • Experience managing and troubleshooting event streaming platforms at scale (e.g., Kafka, Kinesis).
  • Hands-on experience with cloud platforms such as AWS and GCP in production.
  • Familiarity with data lake architectures and large-scale data processing frameworks (e.g., Iceberg, Flink, Spark).
  • Experience with continuous profiling and performance optimization tools.
  • Experience working with or contributing to open source projects.
  • Prior participation in the Wikimedia movement.

Compensation & Location Details

The Wikimedia Foundation is a remote-first organization with staff across 40+ countries. For US-based applicants, the anticipated annual pay range for this position is US $116,633 to US $181,243, depending on location and experience. Pay ranges for non-US applicants are adjusted based on the country of hire via local Employer of Record (EOR) partners.

Culture

Async-friendly

Benefits

Vision

Open to

Worldwide

Sign in to track applications and earn points.

More roles at Wikimedia Foundation

Similar remote roles