
About CertifyOS
CertifyOS is building the data infrastructure that powers modern healthcare. Today, healthcare organizations rely on fragmented and outdated provider data, creating administrative work, regulatory risk, and higher costs. We're solving that problem with an API-first platform that automates provider licensing, enrollment, credentialing, and network monitoring.
About the Role
We're looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships automation, and stands behind it in production. You'll work across cloud-native infrastructure on systems processing millions of provider records.
Problems You'll Solve
- Reliability and observability at scale: Maintain uptime, reduce alert fatigue, and build actionable observability across GKE and Cloud Run using meaningful SLIs, error budgets, and data quality signals.
- Scaling infrastructure efficiently: Improve autoscaling behavior, resource utilization, and workload efficiency across cloud-native distributed systems as platform usage grows.
- Incident response and operational maturity: Own incident response processes, root cause analysis, escalation workflows, and runbooks.
- Infrastructure automation and developer velocity: Build and maintain Infrastructure as Code, CI/CD pipelines, and operational tooling to reduce manual work.
- Reliability engineering for data platforms: Instrument data freshness and infrastructure health alongside service uptime.
What We're Looking For
Reliability Engineering Fundamentals
- 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering operating production systems at scale
- Track record of improving reliability end-to-end and building alerting to prove it
- Strong Linux systems administration, incident response, and root cause analysis skills
- Comfort influencing operational standards and mentoring teams
Cloud Infrastructure & Platform Engineering
- Deep hands-on experience with GCP (GKE, Cloud Run, and containerized workloads at scale)
- Experience building and maintaining Infrastructure as Code with Terraform and/or Pulumi
- Fluency across deployment patterns (rolling, blue/green, canary) and rollback strategies
- Experience with autoscaling, resource optimization, and infrastructure security in regulated environments
Observability & Operational Excellence
- Strong understanding of Golden Signals monitoring and making them actionable
- Experience designing SLIs, SLOs, error budgets, alerting strategies, and dashboards
- Hands-on experience with observability platforms like Google Cloud Monitoring, Datadog, Grafana, or Prometheus
Automation & Software Delivery
- Experience building and maintaining CI/CD pipelines using GitHub Actions or similar
- Scripting or programming fluency in Python, Bash, Go, or similar
- Experience with Git workflows and modern software delivery practices
Communication & Compliance
- Strong written and verbal communication skills
- Experience operating systems handling sensitive data or PII in regulated environments
Benefits
- 100% coverage of health, dental, and vision insurance premiums for employees
- Unlimited PTO with at least two weeks off each year to recharge
Timezone overlap
UTC-8–-4
Open to
US
Sign in to track applications and earn points.