Authzed logo
AuthzedΒ·

Sr. Site Reliability Engineer - Authzed

About AuthZed

We are the creators and maintainers of SpiceDB and the authorization infrastructure that companies around the world depend on to keep their engineering teams focused on what matters most - their own product.

AuthZed is a fully remote company with employees across the US, Canada, and Europe. We bring integrity to all our interactions, fostering confidence in decision making - trusting and respecting each voice on our team, every day.

Company Values:

  • Agency: Everyone should have the capability, freedom, and confidence to bring about changes to our business and product.
  • Collaboration: Success is defined in various dimensions and no single person can be an expert in all of them.
  • Open-mindedness: Without asking questions, testing assumptions, and questioning our pre-existing biases we risk operating within an echo-chamber.

About the Role

As a Site Reliability Engineer, you will play a critical role in ensuring the reliability, availability, and performance of our systems. You will be responsible for designing, implementing, and maintaining scalable infrastructure solutions to support our growing customer base.

What you’ll own:

  • Design, implement, and maintain highly available and scalable infrastructure solutions for our projects, products, and customers.
  • Monitor and analyze system performance, identifying and resolving bottlenecks and issues to ensure optimal performance and reliability.
  • Automate infrastructure deployment and configuration management processes.
  • Continuously improve system reliability, security, and efficiency through proactive monitoring, capacity planning, and performance tuning.
  • Troubleshoot and resolve complex infrastructure and application issues in production and test environments.
  • Collaborate with software engineering teams to design and implement systems that are resilient, scalable, and secure.
  • Participate in on-call rotation and respond to production incidents in a timely manner.
  • Document system configurations, troubleshooting procedures, and operational guidelines.

What you bring:

  • Proven experience as a Site Reliability Engineer or in a similar role.
  • Strong understanding of networking, operating systems, and cloud infrastructure.
  • Experience with Site Reliability Engineering, System Design, and Distributed Computing.
  • Experience in various programming languages (e.g., NodeJS, Java, Python, Ruby, and Go).
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Knowledge of infrastructure-as-code tools like Terraform and Pulumi.
  • Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack).
  • Experience with lower-level implementation details of relational databases (bonus for distributed SQL databases like Google Cloud Spanner or CockroachDB).
  • Experience working with Git, GitHub, and continuous integration/deployment systems.
  • Strong problem-solving, troubleshooting, communication, and collaboration abilities.

Extra shine:

  • Experience with Authorization systems.

Life at AuthZed:

  • Opportunity to work with cutting-edge technology in a rapidly growing sector.
  • Competitive salary and stock options at an early-stage startup.
  • Comprehensive benefits including healthcare (US-based) and other insurance.
  • Fully remote and flexible schedule with twice-yearly travel for team offsites.

Timezone overlap

UTC-8–-4

Open to

NA

Sign in to track applications and earn points.

More roles at Authzed

Similar remote roles