
SpyCloud is on a mission to make the internet a safer place by disrupting the criminal underground. SpyCloud’s solutions thwart cyberattacks and protect more than 4 billion accounts worldwide. Cybersecurity is an exciting, evolving space, and being at the forefront of the fight to disrupt cybercrime makes SpyCloud a special place to work.
We're looking for a Senior Data Scientist, Applied ML to design, build, and deploy models for critical cybersecurity use cases like incident detection and mitigation, fraud intelligence, and risk scoring.
You'll own the full model lifecycle — from data understanding and preparation through prototyping and deployment in production — and work closely with engineering, product, and research teams to turn complex problems into scalable, reliable systems.
What You'll Do
- Develop, train, and deploy models using real-world structured and unstructured data to power critical security features such as threat detection and alerting, entity resolution and risk scoring, and natural language-based tagging and classification.
- Build the preprocessing and feature engineering pipelines your own models depend on, owning model monitoring, evaluation, and feedback loops.
- Take existing prototypes from research or your own experimentation to production-grade reliability in modern cloud environments like AWS.
- Partner with product managers and domain experts to define success criteria, rapidly prototype MVPs, and contribute to system design and architectural decisions.
- Clearly articulate model design choices, tradeoffs, and outcomes to technical and non-technical stakeholders.
Requirements
- 4+ years of experience building and shipping models in production with direct, hands-on ownership of the data lifecycle.
- Strong background in applied math (linear algebra, optimization, statistics) and machine learning.
- Demonstrated experience leveraging Natural Language Processing (NLP) techniques for text classification, tagging, or entity extraction.
- Proficiency in Python and key ML libraries: PyTorch, TensorFlow, scikit-learn, XGBoost.
- Demonstrated experience building or maintaining data/feature pipelines (e.g., Airflow, Spark, Pandas).
- Comfort with model versioning and monitoring in production (e.g., MLflow, DVC).
- Working experience deploying models into cloud environments or containerized services.
- Strong communication skills to translate complex problems into actionable solutions.
Nice to Have
- Deeper MLOps/DevOps/data engineering exposure (infra-as-code, CI/CD depth).
- Familiarity with cybersecurity datasets or domains (threat intelligence, account takeover, ransomware).
- Exposure to graph analytics, knowledge graphs, or cybersecurity frameworks like MITRE ATT&CK.
- Background working with unstructured data (log files, threat reports, breach datasets).
Timezone overlap
UTC-8–-4
Open to
US · Austin · United States
Sign in to track applications and earn points.