
Who We Are
At Domino, we build solutions that help the largest, highly regulated organizations adopt AI to accelerate mission-critical use cases. Our platform integrates a streamlined model and app development environment, advanced model, agent, and app hosting capabilities, and novel governance capabilities providing regulator-ready AI at scale.
Our customers — like Johnson & Johnson, GSK, Bristol Myers, UBS, FINRA, and the US Navy — use our software to solve critical global challenges. Backed by Sequoia Capital, Coatue Management, NVIDIA, Snowflake, and other leading investors, we operate with the agility and spirit of a startup after a decade in business.
What We Are Building
The Automation Team at Domino acts as a force multiplier for engineering, building the tools and systems that enable teams to ship code confidently and consistently. A core part of this mission is Tempest, an in-house platform that orchestrates realistic, long-duration workloads against live Kubernetes clusters and validates results against real observability data.
When scale testing surfaces bottlenecks, resource misconfigurations, or regressions, we need someone who can profile services, trace root causes through Prometheus and New Relic data, and partner with platform engineers to drive durable fixes at the infrastructure level.
What Your Impact Will Be
In your first year, you will:
- Serve as the technical owner of Tempest, Domino's scale and reliability platform, ensuring it remains reliable, extensible, and aligned with evolving infrastructure needs.
- Diagnose and drive resolution of performance bottlenecks and resource misconfigurations surfaced by scale testing — working directly with platform and infrastructure teams to ship fixes.
- Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on empirical testing across deployment sizes.
- Strengthen observability across scale testing by improving Prometheus and New Relic instrumentation.
- Establish and operationalize scale testing on cloud platforms, ensuring appropriate sizing and configuration guidance.
- Partner with platform teams to enable effective scale and reliability testing across additional cloud providers.
- Increase efficiency by building infrastructure automation that scales operationally.
What We Look For
- Background in SRE, platform engineering, or infrastructure with hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments.
- Strong proficiency in Python and comfort working in a large, modular codebase spanning orchestration, infrastructure automation, and systems integration.
- Experience with observability stacks (Prometheus, Grafana, New Relic, or similar) — writing queries, building dashboards, and analyzing metrics to diagnose performance issues.
- Demonstrated ability to profile services, identify resource bottlenecks, and ship durable fixes with engineering teams.
- Familiarity with performance and load testing methodologies (e.g., Locust, k6, or similar).
- Clear ownership mindset — self-directed, accountable, and skilled at communicating priorities and status in a remote, async environment.
What We Value
- A growth mindset with high-performing creative individuals who dig into complex problems.
- Transparency, truth-seeking, and authenticity at work.
- Continuous improvement — an ethos that everything is a work in progress.
- An environment of teaching, learning, and mutual mentorship.
- A diverse, inclusive team welcoming of all backgrounds.
Timezone overlap
UTC-8–-4
Culture
Async-friendly
Open to
US
Sign in to track applications and earn points.