Knock logo
Knock·

Infrastructure Engineer - Knock

Remote-firstFull-timeSeniorUTC-8–-4USUnited StatesNew York#terraform#kubernetes

About Knock

Knock is on a mission to help products communicate with their users in a more thoughtful way. Building product notifications in-house takes months, often leading to poor user experiences. We believe that—when done right—product notifications help users find value in the products they use every day. That’s why we built Knock.

We're a remote-first (with a NYC base) Series A startup of 20+ employees that believe in the power of great software. We're APIs all the way down at Knock—Stripe for payments, Algolia for search, WorkOS for SSO. We're excited to add Knock to that list and to push forward the API-first movement.

About the role

We're looking for an infrastructure engineer to join our small but growing platform team. The platform team at Knock is responsible for building, scaling, and maintaining the core services and infrastructure that run Knock.

You will have a high degree of ownership and autonomy in improving the Knock platform, starting with our foundational infrastructure. We’re an engineer-led team that obsesses over the reliability and availability of our service.

What you’ll be doing

  • Adopting a Terraform-backed EKS cluster, modernizing & maintaining it for elastic scale, reliability, performance, and security.
  • Troubleshooting Postgres performance, queues of every shape and size, and creating a plan to scale 10x to 100x.
  • Identifying and correcting scaling issues by improving telemetry and traces in Datadog, AWS Cloudwatch, and Honeycomb.
  • Maintaining and improving upon our >99.95% uptime track record.
  • Supporting our product engineering team by improving day-to-day developer experience through canaries, faster cycle times, and blue/green deployments.
  • Participating in on-call rotations on a schedule with the rest of the engineering team.

What we’re looking for

  • 4+ years of experience as a DevOps engineer or similar in a startup or mid-sized company working with complex systems at scale.
  • Experience working in and on production Kubernetes clusters using infrastructure as code (Terraform, Pulumi, or Cloudformation).
  • Experience working on complex AWS deployments (multi-account, complex VPC structures to support EKS).
  • Experience operating and scaling database technologies (Aurora Postgres, Mongo, or ClickHouse).
  • Familiarity operating and scaling queues and streams across SQS, Kinesis, Kafka, or similar.
  • Strong problem-solving skills with a focus on reliability, scalability, and performance.
  • Strong communication skills for a fully distributed, remote-first team.

Timezone overlap

UTC-8–-4

Open to

US · New York · United States

Sign in to track applications and earn points.

More roles at Knock

Similar remote roles