
About Supabase
Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.
About the Role
Supabase manages millions of Postgres instances and is growing. We are concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, establishing the practices, frameworks, and feedback loops that empower engineering teams to own their reliability.
This role is ideal for someone with a strong vision for SRE who thrives in async, fast-paced environments where influence matters more than authority.
What You'll Own
- SLI/SLO Strategy: Partner with service teams to define meaningful SLIs/SLOs and build error budget policies that drive engineering decisions.
- Operational Readiness: Evolve the Operational Readiness Review (ORR) process for new services and major changes.
- Incident Management: Strengthen the incident-to-improvement pipeline, identifying repeat failure patterns and driving systemic fixes.
- Reliability Expertise: Act as the reliability expert for architecture reviews, failure mode analysis, and resilience design.
- Toil Reduction: Identify and quantify operational toil, advocating for and building automation to eliminate it.
- On-Call Excellence: Help teams design sustainable on-call practices, including alert quality, runbook coverage, and noise reduction.
- Maturity Tracking: Report on org-wide operational maturity and surface systemic gaps.
You Might Be a Good Fit If You
- Have 7+ years of experience in SRE, production engineering, or reliability-focused roles.
- Possess a software engineering mindset—you write code and build tools, not just configure them.
- Have hands-on experience defining and operationalizing SLOs/SLIs at scale.
- Have deep experience with incident response and postmortem facilitation.
- Have worked with large-scale multi-tenant systems (managed database platforms or Postgres experience is a bonus).
- Are proficient with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred, Terraform/CDK acceptable).
- Communicate clearly and persuasively, with the ability to influence without authority.
- Have experience in async or globally distributed teams.
Nice to Have
- Experience with Kubernetes-based platform operations.
- Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tooling.
- Experience building developer-facing reliability tooling (SLO dashboards, ORR frameworks, DORA metrics).
What We Offer
- Fully Remote: We hire globally with a WeWork membership or co-working allowance.
- ESOP: Equity ownership in the company.
- Tech Allowance: Budget to set up your ideal work environment.
- Health Benefits: 100% coverage for employees and 80% for dependents.
- Annual Off-Sites: Company-wide gatherings for connection and collaboration.
- Professional Development: Annual education allowance for courses, books, and conferences.
Culture
Async-friendly
Open to
Worldwide
Sign in to track applications and earn points.