Wikimedia Foundation logo
Wikimedia Foundation·Verified

Senior Site Reliability Engineer, Data Persistence - Wikimedia Foundation

Remote-firstFull-timeSenior$117K - $181KUSEuropeAPAC+2 more#sre

Senior Site Reliability Engineer, Data Persistence

Company: Wikimedia Foundation

Location: Remote

About the Role

The Wikimedia Foundation is seeking a Senior Site Reliability Engineer to join their globally distributed team. You will be instrumental in supporting and developing the platform that powers Wikipedia, serving millions of users worldwide. The Site Reliability Engineering (SRE) team is responsible for the health and continuous development of Wikimedia's top-10 global website and its underlying infrastructure, all in service of the mission to help everyone share in the sum of all knowledge.

This is a remote-first role within a diverse and globally distributed team that embraces open source principles, publishing all documentation, code, and configuration as open source. If you are passionate about improving the reliability and delivery of a major internet website and thrive in a remote work environment, this opportunity may be for you.

Note: This role requires occasional travel (1-2 times per year) for in-person events and team meetings.

Responsibilities

  • Perform day-to-day operational/DevOps tasks on Wikimedia’s public-facing infrastructure, including deployment, maintenance, configuration, and troubleshooting.
  • Implement and utilize configuration management and deployment tools such as Puppet and Kubernetes.
  • Lead continuous improvement initiatives by automating the installation, configuration, and maintenance of services.
  • Collaborate closely with product teams on the architectural design of new services to ensure scalability and operational efficiency.
  • Participate in a 24/7 on-call rotation, including incident response, diagnosis, and follow-up on system outages or alerts.
  • Work effectively within a global, cross-functional team using asynchronous communication.
  • Mentor peers in areas of technical and operational expertise.

Skills and Experience

  • 6+ years of experience in an SRE/Operations/DevOps role within a team environment.
  • Proficiency in shell scripting and at least one SRE-context scripting language (Python, Go, Bash, Ruby; Python is primarily used).
  • Experience with configuration management tools (Puppet, Ansible; Puppet is used).
  • Experience with distributed caching systems, including their underlying algorithms and performance optimization.
  • Experience with package management on Linux systems (Debian is used).
  • Strong Linux system-level troubleshooting skills.
  • Proven history of automating tasks and processes, identifying process gaps, and capitalizing on automation opportunities.
  • Strong English language skills (verbal and written) and the ability to work independently and collaboratively in a distributed team across multiple time zones.
  • Experience leading and participating in incident response and post-incident review processes, focusing on root cause analysis and preventive measures.

Desired Qualifications

  • Experience with Linux kernel tuning.
  • Experience with monitoring, metrics, and logging infrastructure (e.g., Prometheus, Grafana).
  • Contributions to Free and Open Source software or active participation in an open-source community.
  • Experience with LAMP stack technologies (PHP/HHVM, memcached/Redis); MediaWiki experience is a significant plus.
  • Experience defining and implementing cross-team Service Level Objectives (SLOs).
  • Experience operating on-premise filesystems or object stores at scale (e.g., OpenStack Swift, Ceph).
  • Experience with advanced distributed storage and database systems (e.g., Cassandra, MariaDB).
  • Experience managing backups as an SRE, with some experience with Bacula.

About the Wikimedia Foundation

The Wikimedia Foundation is the nonprofit organization behind Wikipedia and other free knowledge projects. Our mission is to build a world where every human can freely share in the sum of all knowledge. We host Wikipedia, develop software for reading and contributing, support volunteer communities, and advocate for policies that foster free knowledge. As a charitable, not-for-profit organization, we rely on donations.

The Wikimedia Foundation is committed to diversity and inclusion, fostering an equitable and inclusive workplace. We encourage applications from individuals with diverse backgrounds and do not discriminate based on protected characteristics.

Salaries are competitive, equitable, and consistent with our values. The anticipated annual pay range for this position within the United States is US$116,633 to US$181,243, with the final offer determined by factors such as cost of living and location. For applicants outside the US, the pay range will be adjusted accordingly. Salary history is not considered in compensation decisions.

Hiring Locations:

  • US States: Arizona, California, Colorado, Connecticut, District of Columbia, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Maryland, Massachusetts, Michigan, Minnesota, Missouri, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Oregon, Pennsylvania, Puerto Rico, Rhode Island, Tennessee, Texas, Utah, Vermont, Virginia, Washington, West Virginia, Wisconsin, Wyoming.
  • Countries: Brazil, Canada, Colombia, Germany, Ghana, India, Indonesia, Italy, Kenya, Mexico, Morocco, Netherlands, Poland, Singapore, South Africa, Spain, Switzerland, United Kingdom.

Note: Non-US employees are hired through a local third-party Employer of Record (EOR) and must have current work authorization in their location. Some locations require citizenship/permanent residency.

For assistance or accommodation during the application process, please contact recruiting@wikimedia.org or +1 (415) 839-6885.

Open to

US · Europe · APAC · LATAM +1

Sign in to track applications and earn points.

More roles at Wikimedia Foundation

Similar remote roles