Deepgram logo
Deepgram·

Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Remote-firstFull-timeSeniorUTC-8–-4US#kubernetes#terraform#awsBonus

Company Overview

Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and production-grade voice agents. With over 200,000 developers and 1,300+ organizations, we are at the forefront of voice-native foundation models.

Company Operating Rhythm

At Deepgram, we maintain an AI-first mindset. We expect team members to actively experiment with advanced AI tools and integrate them into their daily workflows. We move at the pace of AI; this role requires a candidate who is excited to adapt, think on their feet, and learn constantly in a fast-evolving environment.

The Opportunity

We are looking for an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for our AI/ML research and product development. You will architect and run platforms spanning AWS and bare metal data centers, enabling teams to train and deploy complex models at scale using Kubernetes, Terraform, and Slurm.

What You’ll Do

  • Architect & Maintain: Manage core computing platforms using Kubernetes on AWS and on-premise.
  • Infrastructure as Code: Develop and manage infrastructure using Terraform to ensure reproducible, versioned, and automated environments.
  • AI/ML Orchestration: Design and optimize job scheduling systems, integrating Slurm with Kubernetes for GPU resource management.
  • Bare Metal Management: Provision and maintain on-premise server infrastructure for high-performance GPU computing.
  • Networking & Storage: Implement CNI, service mesh, CSI, and S3 solutions for high-throughput, low-latency workloads.
  • Observability: Develop a comprehensive stack for monitoring, logging, and tracing.
  • Collaboration: Partner with AI researchers and ML engineers to build tools that accelerate development cycles.
  • Automation: Automate the lifecycle of single-tenant, managed deployments.

Requirements

  • 5+ years of experience in Platform Engineering, DevOps, or SRE.
  • Proven, hands-on experience with Terraform in production.
  • Expert-level knowledge of Kubernetes architecture and operations.
  • Strong scripting skills (Python, Go, or Bash).
  • Experience with CI/CD systems (e.g., GitLab CI, Jenkins, ArgoCD).

Bonus Points

  • Experience with HPC job schedulers (Slurm) for GPU-intensive workloads.
  • Experience managing bare metal infrastructure (PXE boot, MAAS).
  • Familiarity with FinOps and cloud cost optimization.
  • Knowledge of Kubernetes networking (Calico, Cilium) and storage (Ceph, Rook).
  • Experience in multi-region or hybrid cloud environments.

Timezone overlap

UTC-8–-4

Benefits

Bonus

Open to

US

Sign in to track applications and earn points.

More roles at Deepgram

Similar remote roles