
About Domino Data Lab
At Domino, we build solutions that help the largest, highly regulated organizations adopt AI to accelerate mission-critical use cases. Our platform integrates a streamlined model and app development environment, advanced model, agent, and app hosting capabilities, and novel governance capabilities providing regulator-ready AI at scale. Our customers — like Johnson & Johnson, GSK, Bristol Myers, UBS, FINRA and the US Navy — are using our software to solve some of the most important challenges in the world.
Backed by Sequoia Capital, Coatue Management, NVIDIA, Snowflake and other leading investors, we have been in business for over a decade but are still a small team operating with the spirit of a startup.
About the Role
As a Technical Support Engineer, you're the bridge between our customers and our Engineering organization. You'll own technical support cases end-to-end, triaging issues across Kubernetes infrastructure, ML platform components, authentication, data connectivity, and model deployment, and ensuring every customer gets a clear and timely resolution. You'll also contribute to the knowledge base that helps the whole team scale.
What You'll Do
- Own support cases for enterprise customers across all severity levels, from initial triage through resolution, with clear communication and accurate expectations throughout
- Diagnose and resolve Kubernetes and cloud infrastructure issues: pod failures, resource limits, persistent volumes, RBAC, ingress, and cluster-level diagnostics
- Troubleshoot ML platform problems including workspace and job failures, environment build errors, model deployment issues, and data connector failures
- File detailed, actionable bug reports and enhancement requests in Jira and act as the customer's advocate with Product and Engineering
- Write and review knowledge base articles, how-to guides, and troubleshooting docs to help customers and teammates solve problems faster
- Hand off cases cleanly in a follow-the-sun model across AMER, EMEA, and APAC, ensuring continuity for global enterprise accounts
- Run live troubleshooting sessions with customers via video call and participate in EMEA weekend on-call rotation per team schedule
What We Look For
- 3 to 5 years in enterprise technical support, solutions engineering, or a similar customer-facing technical role at a SaaS or data/AI platform company
- Hands-on Kubernetes experience: pod lifecycle, kubectl, RBAC, namespaces, persistent volumes, and cluster-level troubleshooting
- Strong Linux and command-line proficiency: log analysis, process management, file system navigation, and shell scripting
- Familiarity with Python-based ML workflows: Jupyter, package management, model training and serving
- Experience with cloud platforms (AWS, GCP, or Azure) and containerized application environments
- Methodical troubleshooting skills: forming hypotheses, testing them, and adapting when logs disagree
- Clear written communication for case updates and KB documentation
- Comfort managing multiple open, time-sensitive cases concurrently
- Ability to work effectively asynchronously across time zones in a remote-first, globally distributed team
- Bachelor's degree in computer science, engineering, or a related technical field (or equivalent practical experience)
What We Value
- Diverse teams and inclusive perspectives across all backgrounds
- A growth mindset with high-performing creative problem solvers
- Transparency, authenticity, and speaking the truth
- Continuous improvement in everything we build
- A culture of teaching and continuous learning
Timezone overlap
UTC+0–+1
Culture
Async-friendly
Open to
UK
Sign in to track applications and earn points.