Docker logo
Docker·

Staff Software Engineer, Billing - Docker

About Docker

Docker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout.

We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.

We're building AI-native development practices into how this team works at a foundational level. That means infrastructure design needs to account for a new kind of collaborator: AI agents that generate, deploy, and operate software. The Staff Engineer on this team will keep systems running, define what safe, observable, AI-assisted infrastructure operations look like in practice, and set the standard for how the broader engineering organization follows.

What You'll Work On

The Billing Platform Engineering team owns the systems that make Docker's commercial model real. You'll work on problems like:

  • Designing infrastructure that makes AI-generated deployments safe to ship and easy to roll back
  • Instrumenting billing systems so that failures—billing miscalculations, entitlement gaps, payment errors—are detected immediately and unambiguously
  • Building infrastructure that scales with usage-based billing workloads without manual intervention
  • Improving the developer experience on this team (local environments, CI/CD pipelines, deployment tooling)

Responsibilities

  • Own and evolve the infrastructure supporting Billing Platform services: compute, storage, networking, CI/CD, and observability
  • Design and maintain IaC (Terraform) for billing system infrastructure on AWS; set module patterns and standards for the team
  • Build and own observability systems (metrics, logging, alerting) with a focus on billing accuracy and payment reliability
  • Define deployment patterns and runbooks that work well in an AI-agent-assisted development workflow: clear rollback procedures, safe promotion gates, automated validation
  • Partner with software engineers on service design—bringing infrastructure constraints and operational requirements into the conversation before code is written
  • Identify systemic risks and drive improvements that span team or organizational boundaries
  • Lead incident response for billing system issues (participation in an on-call rotation may be required outside standard business hours)
  • Mentor engineers across the team and raise technical standards

Qualifications

  • 8+ years in platform, infrastructure, or SRE roles supporting production SaaS systems at scale
  • Deep AWS expertise: ECS or EKS, RDS (Postgres preferred), networking, IAM, cost management
  • Expert-level Terraform; experience designing reusable module patterns and setting organization standards
  • Experience building and owning observability stacks (Datadog, Grafana, or similar) at an organizational level
  • Strong familiarity with CI/CD systems (Jenkins, GitHub Actions, or equivalent)
  • Operational and architectural experience with Kubernetes
  • Track record of identifying systemic risks and driving cross-team or organizational improvements
  • Security-first mindset: threat modeling, blast radius analysis, least-privilege by default, audit trails
  • Strong written English communication skills
  • Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent practical experience

What Sets You Apart

  • Experience with billing, payments, or financial systems infrastructure
  • Thoughtful perspective on infrastructure requirements for AI-agent-assisted software development
  • Proven ability to drive consensus and lead without direct authority

What to Expect

  • First 30 Days: Ship code in your first week using Docker's agent-first development workflow. Get hands-on with Billing Platform infrastructure, shadow on-call, and build a system overview.
  • First 90 Days: Take ownership of infrastructure components, deliver a production improvement, and fully participate in the on-call rotation.
  • One Year Outlook: Serve as the trusted authority on billing infrastructure, leading major improvements to observability, deployment safety, and platform reliability.

Benefits & Perks

  • Remote-first culture with optional access to offices in Seattle and Paris
  • Flexible scheduling
  • Generous PTO, quarterly Whaleness Days, and an end-of-year Whaleness break
  • Home office support & US$100 net/month technology stipend
  • Annual learning & development stipend for conferences, courses, and certifications
  • 16 weeks paid parental leave after 6 months
  • Equity package for full-time employees
  • Comprehensive medical and retirement benefits

Timezone overlap

UTC-8–-4

Open to

US

Sign in to track applications and earn points.

More roles at Docker

Similar remote roles