Cribl logo
Cribl·

Senior Site Reliability Engineer - Cribl

Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the world’s biggest enterprises, including half of the Fortune 100, to bridge the gap between AI ambition and infrastructure reality. As the AI Platform for Telemetry, we give customers the choice, control, and flexibility to manage and analyze telemetry for both humans and agents, so they can build what’s next.

We’re seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data. You will join a team of technical engineers committed to shipping high-quality software. Fixing things at the operational side should always be the last resort, so our SREs are involved from conception to design to development and all the way through production and beyond.

As An Active Member Of Our Team, You Will

  • Engage with teams and improve service delivery and reliability across their entire lifecycle
  • Measure and monitor all production systems with an eye towards availability, latency and overall system health
  • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
  • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
  • Help identify and drive down toil with creative innovation and automation
  • Participate in stand-by, on-call, or off-hours duties

What We're Looking For

  • Proven experience designing, implementing, and operating observability systems for complex cloud-based platforms, with deep knowledge of best practices
  • Experience with Configuration Management and Infrastructure as a Code Tools like Terraform (preferred) or Ansible, plus Cloud SDKs
  • Knowledge of cloud platforms (AWS and Azure preferred) and container + orchestration technologies
  • Experience with APM and Observability tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, and Sentry
  • Extensive experience with enterprise scale continuous delivery environments
  • Development experience with JavaScript/Node.js/TypeScript in a Linux/Mac environment
  • Experience with sustainable incident response in a blameless environment
  • Background in Linux Systems Engineering
  • Experience with incident response tools like PagerDuty, FireHydrant, and Blameless
  • Comfortable with a high level of autonomy and working with a distributed team
  • Knowledge of Cloud and application security best practices
  • Strong knowledge of cloud design patterns for scale, data management, and resiliency
  • Strong opinions about business metrics and SLOs

Timezone overlap

UTC-8–-4

Open to

US

Sign in to track applications and earn points.

More roles at Cribl

Similar remote roles