
CloudLinux builds Linux infrastructure and security products. You will join our Automation & Management Services cell, working closely with Dmitrii Petrov to solve problems across teams and services: cloud-cost data, infrastructure inventory, network policies and capacity workflows.
Inside the Infrastructure Department, the Platform cell is a small team. We run the observability platform, the company's GitLab and the CI runners behind it, a few smaller engineering services, and the automation the department relies on for provisioning and configuration.
We are looking for a Platform Engineer for the PaaS team to take ownership of an agreed set of platform services: keep them reliable and within agreed service levels, make changes and recovery repeatable, and reduce recurring operational work through automation and self-service.
You get real freedom in how you implement things, and you own the result: you pick the approach, defend it in review, and answer for how the service behaves afterwards.
Most of the job is what site reliability engineering is about: services that run well day after day, and requests from product teams handled properly. Some of it is building. Expect the work to split between platform engineering, incident response, and helping engineers in other teams use what we run.
What you'll do
- Run the observability platform. Keep it healthy, onboard teams, watch cost and capacity, and maintain the alerting that runs on top of it.
- Run GitLab and the CI runner fleet. Upgrades, capacity, access, backups and restore drills.
- Keep the rest of our services healthy, with the monitoring and runbooks a production service needs.
- Deploy new services when they are requested. Research the options, pick a design, and stand the service up from scratch according to good practice: as code, monitored, backed up, documented.
- Work with developers' requests. Access, onboarding, pipeline problems, new exporters and dashboards.
- Run incidents. Diagnose and mitigate impact, restore service safely, then complete the root-cause analysis and post-mortem.
- Ship everything as code, reviewed in merge requests.
- Write for engineers outside the team. Runbooks, onboarding guides, maintenance notices and status updates.
- Work with AI agents. Delegate collection and drafting, review their output, and record what you learn.
Requirements
Must have:
- Senior-level experience in infrastructure, platform or site reliability engineering, including at least one production service you were responsible for keeping up.
- Linux systems administration and debugging on bare metal and virtual machines.
- Kubernetes in production delivered through GitOps, including cluster upgrades.
- Infrastructure as code: Ansible and Terraform or OpenTofu.
- GitLab administration and GitLab CI in production, self-hosted or SaaS.
- Working knowledge of the Prometheus and Grafana ecosystem (PromQL, alerts, dashboards).
- Written technical explanation for engineers outside your team.
- Strong communication and interpersonal skills.
- Advanced use of AI engineering assistants such as Claude and Codex.
- English - upper-intermediate or higher.
Nice to have:
- Alerting design: SLOs, burn-rate alerts, thresholds sized from data.
- MicroVM isolation for CI: Kata Containers, Firecracker or gVisor.
- S3-compatible object storage operations: Ceph RGW or similar.
- AWS with real cost work.
- Self-hosted Sentry, or another Kafka, ClickHouse and Redis-backed application under load.
- Python or Go for exporters and small internal services.
What's in it for you?
- A focus on professional development.
- Interesting and challenging projects.
- Fully remote work with flexible working hours.
- Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leave.
- Compensation for private medical insurance.
- Co-working and gym/sports reimbursement.
- Budget for education.
- The opportunity to receive a reward for the most innovative idea that the company can patent.
Benefits
Open to
Worldwide
Sign in to track applications and earn points.