Replicant logo
Replicant·

Senior Site Reliability Engineer - Replicant

At Replicant, we believe AI should work for people, starting with customer service. That’s why we built a platform that helps contact centers resolve more requests, proactively identify issues, and improve agent performance with AI-powered conversation intelligence and AI agents that act like your best reps.

Our AI agents handle millions of calls every month for Fortune 500 companies and high-growth innovators. From processing payments to booking appointments and authenticating users, they help customers get what they need instantly, 24/7. Meanwhile, our real-time conversation insights help contact center leaders coach better and improve every interaction.

We are leading the shift from legacy systems to AI-first service, powered by large language models (LLMs) and designed for enterprise scale, security, and empathy. If you’re excited by the potential of LLMs, voice AI, and building category-defining technology with a kind, ambitious team, you’ll love it here.

Our SRE team builds — not just supports — an AI-native platform, and they own exciting domains like platform and harness engineering, site reliability, cloud infrastructure, CI/CD, DevEx, observability, incident management, and COGS (e.g. cloud spend visibility). We're looking for a Site Reliability Engineer who has opinions about how these domains should work and wants agency in shaping where they go.

What You’ll Do

  • Contribute to patterns, design, and implementation of our domains; help shape the future of platform engineering at Replicant.
  • Build and improve systems that help reduce toil and enable Replicant's production infrastructure to remain available and operable under large-scale, real-time conversational AI traffic.
  • Extend and iterate our agent harness: Unsupervised AI agents are currently used by about 10% of the dev team — help us grow that number. The agent harness includes CI, sandzones, guardrails, and validation (e.g. agent-first eval loops).
  • Own and improve our CI/CD pipelines and surrounding developer tooling: build and test performance, deployment ergonomics, and paved paths for new services.
  • Participate in on-call rotation and incident management to ensure platform uptime and quality (SRE owns the base infrastructure, not the applications; non-business-hours pages are rare).

What You'll Bring

  • 6+ years’ experience in software development enablement roles.
  • Solid experience owning CI/CD platforms end to end — including domains like caching, architecture, and developer self-service.
  • Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning, paired with a defensible opinion on where to avoid using AI tools.
  • Familiarity with Node/TypeScript including making code changes (e.g. exposing new metrics), Python and Terraform for automation, and developing in a Kubernetes/Helm ecosystem.
  • Practical experience with observability: logs/metrics/tracing, monitoring/alerting, incident management process, and tooling.
  • Experience working in fully remote teams.
  • Bonus: Harness engineering experience, production-at-scale experience with GCP, telephony and SIP architectures (FreeSWITCH in particular).

Our Tech Stack

TypeScript/Node and Python running on Kubernetes — primarily on GCP (multi-cloud), with GitLab CI, Helm, Terraform, Datadog, Prometheus, and Grafana.

Perks & Benefits

  • 🌴 In-person connection: Company-wide offsites and smaller team gatherings.
  • 🖥️ Tech & learning stipend: Funding for conferences, books, and courses.
  • 📍 Remote by design: Distributed team with trusted flexibility.
  • 🏋️ Health & wellness: Flexible vacations, paid sabbatical after 5 years, comprehensive benefits, and a physical/mental well-being stipend.
  • 💸 Competitive compensation: Salaries that match your impact.
  • 📈 Equity with upside: Shared ownership in a fast-growing AI company.

Open to

Worldwide

Sign in to track applications and earn points.

More roles at Replicant

Similar remote roles