
SpyCloud is on a mission to make the internet a safer place by disrupting the criminal underground. Our solutions thwart cyberattacks and protect more than 4 billion accounts worldwide. Cybersecurity is an exciting, evolving space, and being at the forefront of the fight to disrupt cybercrime makes SpyCloud a special place to work.
We are looking for an experienced Cloud Ops Engineer who thrives in a fast-paced and AI-forward environment. You have a passion for innovation, solid design principles, and high-quality development, excelling at designing, improving, and maintaining secure, performant infrastructure while automating processes to ensure efficiency and scalability.
What You'll Do
- Infrastructure Design and Maintenance: Design, improve, and maintain secure, durable, and performant infrastructure to power APIs, AI operations, web applications, and data mining/ETL workflows to meet established SLAs. Collaborate with developers to bring new products and services into production.
- Automation and Monitoring: Automate testing, deployment, and monitoring of all products and services throughout the software development lifecycle. Continuously improve operational processes and apply best practices to ensure scalability, security, and availability.
- Security and Compliance: Proactively meet standards for information security and compliance, such as SOC 2, ISO 27001, and CMMC. Implement and uphold security measures across all infrastructure components.
Requirements
- Professional Experience: At least 5 years of professional experience in a Cloud Ops, Platform Ops, or DevOps role maintaining production infrastructure, preferably supporting a highly available environment for a SaaS or cloud service provider.
- Technical Proficiency:
- Strong working knowledge of AWS services such as EC2, ECS or EKS, Lambda, API Gateway, RDS, DynamoDB, CloudWatch, S3, Code/Build/Pipeline/Deploy, and VPC Lattice.
- Strong working knowledge of Terraform (or similar tools), Ansible, AWS CLI/SDK, and Boto.
- Proficiency with scripting languages such as Python and Bash, alongside Linux environments.
- Strong understanding of system and networking concepts and troubleshooting techniques for bare metal and containerized workloads.
- Experience supporting AI and ML systems, agent orchestration frameworks (e.g., AgentCore), and integrating LLMs into production systems.
- Additional Skills: Experience with release automation, system administration and configuration, and system debugging.
Nice to Have
- Databricks or Snowflake infrastructure experience
- Cost optimization for AI and LLM workloads
- Internal developer platform experience
Timezone overlap
UTC-8β-4
Open to
US Β· Austin Β· United States
Sign in to track applications and earn points.