PostHog logo
PostHog·Verified

Site Reliability Engineer (US - Central/Eastern time)

Remote-firstFull-timeSeniorUTC-6–-4AsyncUS#kubernetes#aws#terraform

About PostHog

PostHog makes products self-driving. It is the only platform that acts as a co-pilot for developers and AI agents to diagnose bugs, analyze data, and roll out changes autonomously. We are product-led, default alive, and well-funded, with over 450,000 organizations using our platform.

Things We Care About

  • Transparency: We share our roadmap, strategy, and internal metrics openly.
  • Autonomy: Engineers lead product teams and make their own decisions.
  • Shipping Fast: We prioritize small, autonomous teams that own products end-to-end.
  • Time for Building: We are a natively remote company that defaults to async communication with meeting-free Tuesdays and Thursdays.
  • Ambition: We aim to solve big problems and value optimistic, high-upside thinking.
  • Being Weird: We embrace unconventional ideas as a competitive advantage.

Who We're Looking For

We are seeking SREs based in the US (Central/Eastern Time) who enjoy deep ownership of production systems. You should be comfortable with stateful infrastructure, AWS, and automation.

  • Enthusiastic Drivers: Proactive individuals who own projects from start to finish.
  • Optimistic Problem Solvers: Resilient collaborators who iterate their way out of complex challenges.
  • Grown Ups: Low-ego, professional, and respectful team members.
  • Genuine Builders: People who love building software and infrastructure.

What You'll Be Doing

This is not a "keep the lights on" role. You will turn a fast-growing, stateful system into a predictable, well-automated platform.

  • Operate EKS clusters with Karpenter, Cilium, and ArgoCD.
  • Manage multi-account AWS environments, networking, and access control.
  • Maintain Terraform/Terragrunt IaC pipelines.
  • Improve operational tooling for deploys, schema changes, and incident response.
  • Reduce operational load through self-healing automation.
  • Optimize cloud spend and participate in on-call rotations.

Requirements

  • Deep hands-on experience with Kubernetes (EKS) in production at scale.
  • Strong experience operating production infrastructure on AWS (IAM, networking, multi-account).
  • Proficiency in Terraform or Terragrunt with module design experience.
  • Solid understanding of Linux systems (disk, memory, networking).
  • Experience supporting stateful systems (databases, queues, storage).
  • Ability to debug performance and reliability issues in production.

Nice to Have

  • Experience with GitOps (ArgoCD) and CI/CD (GitHub Actions).
  • Experience building AI agent-enabled infrastructure services.
  • Familiarity with multi-region infrastructure and consistency/availability tradeoffs.

Timezone overlap

UTC-6–-4

Culture

Async-friendly

Open to

US

Sign in to track applications and earn points.

More roles at PostHog

Similar remote roles