Fingerprint logo
Fingerprint·

Senior Site Reliability Engineer - Fingerprint

Fully remoteFull-timeSenior$152K - $205KWorldwide#Go#AWS#Terraform

About Fingerprint

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence. We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.

Fingerprint is a globally dispersed, 100% remote company. We were named on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.

We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb), Nexus Venture Partners (Postman, Apollo.io, MinIO, Druva), and Uncorrelated Ventures (Redis, Rollbar, Gradle).

About the Role

We're looking for a Senior Site Reliability Engineer to join our Infrastructure team and take ownership of how our platform behaves in production. This is a hands-on engineering role where you will write code and infrastructure, own systems end to end, and be measured by whether the things you own stay fast, available, and predictable as we grow.

You'll work across observability, incident response, capacity and performance, change safety, and the tooling that makes all of it routine. You'll define what "reliable" means for critical paths, instrument them, and partner with product engineering teams to make their services operable by design.

Responsibilities

  • Own the reliability of core production systems end-to-end, setting targets and operating them under real traffic.
  • Define and maintain SLIs and SLOs for critical paths, wiring them into dashboards and alerts, using error budget burn to prioritize fixes.
  • Drive alert quality: raise signal, eliminate noise, and implement anomaly and correctness detection.
  • Take a lead role in incident response—investigate across service boundaries, restore service, and write actionable postmortems.
  • Build secure, resilient, and cost-efficient infrastructure with attention to failure modes (timeouts, retries, backpressure, load shedding, blast radius containment).
  • Conduct capacity and performance analysis using real data (load testing, profiling, headroom planning).
  • Improve change safety through progressive delivery, automated rollbacks, and safe deployment practices.
  • Manage infrastructure as code primarily using Terraform.
  • Design, write, and ship software and developer tooling to reduce toil.
  • Run deliberate failure testing (game days and chaos exercises).
  • Partner with product engineering teams on production readiness reviews for new and high-risk services.
  • Participate in and improve the on-call rotation to reduce pager fatigue.
  • Incorporate security best practices in code and peer reviews.
  • Act as a technical expert for production problems and mentor team members.

Qualifications

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within cloud-based environments (AWS preferred).
  • Proven track record of owning end-to-end production systems.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets.
  • Strong incident response leadership on high-severity, customer-facing incidents.
  • Deep understanding of distributed systems failure modes in high-throughput, low-latency environments (cache/database saturation, cascading failures, retry storms).
  • Solid cloud infrastructure fundamentals: networking, load balancing, containerization (EKS/Kubernetes), and distributed systems.
  • Strong experience managing infrastructure with Terraform or equivalent.
  • Proficiency in Go, Python, or a comparable language to build production-ready software.
  • Fluency with observability tools (Datadog, Prometheus, Grafana, OpenTelemetry).
  • Hands-on experience operating Redis/ElastiCache in production (cluster/shard management, failover, memory eviction policies, scaling strategies).
  • Fluency with software engineering best practices: version control, code reviews, testing, and safe deployments.
  • High level of personal autonomy and experience working with undefined requirements.
  • Pragmatic mindset balancing reliability investments with delivery timelines.
  • Strong written and verbal English communication skills.
  • Experience leveraging AI tools as part of day-to-day engineering workflows.

Work Eligibility & Hiring Notice

Fingerprint is an all-remote company and hires from almost any country, provided teammates are authorized to work from their home location. Fingerprint does not sponsor visas. Due to regulatory and security reasons, there are a small number of restricted countries where Fingerprint cannot employ teammates.

Open to

Worldwide

Sign in to track applications and earn points.

More roles at Fingerprint

Similar remote roles