
At Voxel51, our mission is to bring transparency and clarity to the world's data. Our platform, FiftyOne, is where AI work happens, serving as the mission-critical linchpin for managing unstructured data, model development, and AI systems at the world's largest companies. We are fully remote, hiring people based in the United States who are prepared to travel to at least two in-person retreats per year.
About Your Role
As a Senior Infrastructure Engineer, you will shape the architecture and strategy of the systems that power our platform — from individual researchers to enterprise-scale deployments. You'll lead the design of containerized systems, CI/CD pipelines, and deployment solutions across cloud and on-premises environments, while solving the unique challenges of serving unstructured data (images and video) at scale.
You'll partner with enterprise customers, guiding and troubleshooting their production deployments. You'll collaborate across engineering teams to improve developer productivity, and mentor peers while setting infrastructure best practices.
What You Will Do
- Shape the architecture and evolution of Voxel51's infrastructure to support deployments ranging from individual researchers to Fortune 500 enterprises.
- Design, build, and scale deployment systems across cloud (GCP, AWS, Azure) and on-premises environments, ensuring reliability, security, and repeatability.
- Partner with enterprise customers to deliver and support production-grade deployments, guiding them through installation, troubleshooting, and scaling.
- Lead infrastructure initiatives across engineering teams, enabling peers to develop, test, and ship features faster with robust internal tooling and automation.
- Drive best practices in CI/CD, evolving pipelines (GitHub Actions + Google Cloud Build) and introducing new approaches.
- Develop and maintain deployment solutions for Voxel51-hosted environments (GKE) and customer on-prem installations (K8s or Docker Compose).
- Troubleshoot and resolve complex infrastructure issues spanning build failures, runtime failures, and customer deployment challenges.
- Anticipate and prevent failures by designing monitoring, alerting, and predictive solutions.
- Mentor engineers and set technical direction to keep infrastructure ahead of customer needs.
What You Should Bring
- Deep experience with containerized environments (building/debugging container images, Kubernetes, Docker Compose, Helm charts).
- Infrastructure as Code expertise (Terraform, Ansible, or equivalent).
- Scripting and automation skills (Bash or similar).
- Python expertise, including build and environment management, packaging, release management, and dependency debugging.
- CI/CD systems experience, ideally GitHub Actions.
- Cloud infrastructure knowledge, especially GCP (IAM, VPC, load balancing, routing, proxies, firewall rules).
- Database fundamentals, ideally MongoDB or similar NoSQL systems.
- Observability skills (monitoring, logging, tracing, alerting).
- Security best practices (certificates, service accounts, least privilege, role assumptions).
- Troubleshooting ability across complex, distributed systems with customer interaction.
- Strong communication skills for working with enterprise customers and remote-first teams.
Timezone overlap
UTC-8–-4
Benefits
Open to
US
Sign in to track applications and earn points.