Atlan logo
AtlanΒ·

Senior Software Engineer - Reliability

Who We Are

Atlan is building the context layer for enterprise AI. Enterprises are pouring money into AI and most of it dies in production because the AI does not understand the business context around the data. That is the problem Atlan solves.

Gartner has named context the defining issue in enterprise AI, and calls context graphs the essential infrastructure for AI agents, naming Atlan one of three vendors already building it. We are also the only vendor named a Leader across all four major Gartner and Forrester evaluations for data catalogs, data governance, and metadata management.

Come build the infrastructure that AI runs on.

The Team

You'll join the Reliability team, which runs a multi-agent AI SRE platform that is already live in production. Today, agents investigate incidents, remediate them, and resolve networking tickets across every tenant we run.

Be clear on what this role is not: It is not an incident-command SRE seat. Our view is simple: if a human is fixing production by hand, the system has failed. It is also not a greenfield charter. You'll join a mature, opinionated codebase with documented architecture decisions and eval-gated pull requests, and you'll make it better.

Why now: The platform works, and the hardest problems are wide open. Making agent investigations accurate enough that engineers trust them without re-checking is a genuine frontier problem in agentic reliability. What you build in the next year decides how far autonomous operations can go at Atlan.

What You Will Do

  • Make investigation agents right, not just fast: Push root-cause accuracy toward the point where engineers trust the answer without re-checking it. Treat every wrong diagnosis as a class of failure to remove, not a one-off bug to patch.
  • Grow auto-remediation coverage, safely: Take remediation from a handful of playbooks to dozens. Each one earns autonomy in stages: dry run before execute, human-approved before autonomous.
  • Own the evals that decide when an agent can be trusted: Build and run fault-injection benchmarks and eval harnesses. Set pass marks before the run, report results with honest denominators, and validate the harness itself.
  • Design the gates that let an agent write to production: Implement fail-closed checks, kill switches, approval flows, and blast-radius limits. Start from the question: "What happens when the agent is wrong?"
  • Build new agents where the toil says they belong: Every agent maps to a category of toil it removes, making the next agent cheaper to build.
  • Build the platform for other teams: Provide clean interfaces, guardrails, and safe defaults so teams beyond reliability can trust the platform on their own production.
  • Ship the unglamorous fix when it matters: When something is on fire, stop the bleeding first, then come back and remove the failure class.

What Makes You a Match

  • You've lived operational toil: You've carried a pager, owned incidents end-to-end, or worked an escalation queue long enough to know which pain is worth removing and why.
  • You remove classes, not tasks: You can point to a recurring operational problem you eliminated entirely, explain the category behind it, and articulate why you chose it.
  • You measure by adoption: You track metrics with real numbers and honest denominators. You can talk about what worked, what failed to gain adoption, and what you learned.
  • AI has structurally changed how you work: You've rebuilt core parts of your workflow with AI and shipped AI-native systems that others depend on. If you've built autonomous agents, you design guardrails before capability.
  • Failure-mode humility: You think natively in false-positive rates, rollback paths, and blast radius because you know the operational cost of an agent being wrong in production.
  • You build for external production: Clean interfaces, documented failure modes, and safe defaults guide your work.
  • Systemic rigor: You write deterministic code for routing, filtering, and safety before any generative step. You use code where code is right, treating agents as leverage rather than a substitute for engineering judgment.

Working hours: We work best with overlap roughly between 11:00 AM and 8:00 PM IST.

Why Atlan?

  • Competitive Compensation: Top-of-market base salary, performance-based variable pay, and equity.
  • AI-Native Culture: AI is woven into how we build, think, and work every day.
  • Health & Wellness: Day-1 medical, dental, vision, mental health coverage, and flexible health stipends.
  • Flexible Time Off: Modern leave policies to support your well-being.
  • Global & Remote-First: Work asynchronously with a distributed team across 15+ countries.

Timezone overlap

UTC+5–+6

Culture

Async-friendly

Open to

APAC

Sign in to track applications and earn points.

More roles at Atlan

Similar remote roles