
About SigNoz
SigNoz is an open-source observability platform that helps modern engineering teams monitor, debug, and optimize their applications with deep visibility into metrics, traces, and logs — all in one place. We're built natively on OpenTelemetry and offer both self-hosted and cloud options, so teams can run observability the way they want, without vendor lock-in.
We are growing fast and building core developer infra products:
- 27,000+ GitHub stars
- 800+ customers
- 7,000+ members in our Slack community
The Role
We're looking for a Senior Site Reliability Engineer (SRE) to own the reliability, scalability, and operability of the SigNoz cloud platform. You'll keep a petabyte-scale observability system fast and dependable — making sure the people who trust us to watch their systems can always trust ours. The platform team handles infra, scalability of SaaS, ingest pipelines, staging environments, automation, and the operational backbone of the product.
This is a deeply hands-on role for someone who understands what actually breaks in production at scale — and enjoys fixing it for good.
What We're Looking For
- Kubernetes at scale: Real fluency with the nuances and gotchas (resource tuning, autoscaling behavior, networking, stateful workloads, upgrades, and failure modes under load).
- ClickHouse: Working knowledge of operating it, tuning queries, and understanding its behavior at scale is a strong plus.
- Golang: Knowledge of Go is a plus (most of our stack and tooling is in Go).
- OpenTelemetry: Familiarity with OpenTelemetry and running large-scale data ingest pipelines is a plus.
What You'll Work On
You'll work with a high-caliber team across areas like:
- Reliability: SLOs/SLIs, error budgets, incident response, and sustainable on-call practices.
- Scaling: Making the ingest path robust to bursts while maintaining data freshness.
- SaaS Scalability: Capacity planning across a petabyte-scale system.
- Data Layer: Operating and tuning ClickHouse for performance and cost.
- Kubernetes Infrastructure: Cluster operations, upgrades, multi-tenancy, and automation.
- Dogfooding: Helping make the observability of SigNoz itself world-class.
- Tooling: Infrastructure-as-code, CI/CD, and tooling for a small team to operate big systems.
What Will Make You Successful
- 5–8 years in SRE, infrastructure, or platform/backend roles operating production systems at scale.
- Deep, practical Kubernetes experience.
- Strong grasp of distributed systems failure modes, performance debugging, and capacity planning.
- Comfortable in code (Go preferred) to automate and fix things.
- Love for open source, ideally with prior contributions to OSS projects.
- Comfortable in a high-ownership, fast-moving, remote-first environment.
- Strong communication skills for writing clear runbooks, tech docs, and explaining trade-offs.
Nice-to-Haves
- Past experience on platform/infra/SRE teams of Series B+ startups.
- Hands-on experience operating ClickHouse, Kafka, or similar high-throughput data systems.
- Experience in observability (monitoring / logging / tracing) and with OpenTelemetry.
Why You'll Love Working at SigNoz
- Work on a globally used open-source project that engineers actually love.
- Huge scope and ownership — your work directly shapes how teams adopt SigNoz.
- Collaborate with a high-caliber team who just can't stop shipping.
- Remote-first, async-friendly culture.
- Opportunity to help define the future of open-source observability.
Timezone overlap
UTC+8–+12
Culture
Async-friendly
Open to
APAC
Sign in to track applications and earn points.