Domino Data Lab logo

Staff Software Engineer, IDSM - Domino Data Lab

Who we are

At Domino, we build software that helps the largest, AI-driven organizations build and operate advanced data science and AI solutions at scale. Our platform integrates a streamlined model development environment, MLOps capabilities, and novel features for collaboration, reuse, and reproducibility — all of which make data science teams more productive, reduce time to value, and ensure compliance. Our customers — like Johnson & Johnson, GSK, Bristol Myers, UBS, FINRA and the US Navy — are using our software to solve some of the most important challenges in the world.

Backed by Sequoia Capital, Coatue Management, NVIDIA, Snowflake, and other leading investors, we have been in business for a decade but still operate with the spirit of a startup.

What we are building

The Model Development Lifecycle Team is building a cutting-edge platform to simplify the entire machine learning journey. From development and training to deployment and management, we empower teams to turn data into actionable insights. Our platform supports:

  • Seamless API Integration: Deploy models as APIs for consistent use across applications, whether on-premises, in enterprise infrastructures, or through third-party hosting.
  • Collaboration and Discoverability: Use our model registry to version, store, and easily find models across the organization.
  • Scalable Training Resources: Leverage advanced tools like GPUs, Ray, and Spark to meet the needs of diverse AI projects.

What your impact will be

In your first year, you will:

  • Collaborate with customers to design solutions for deploying models to platforms like AWS SageMaker and Azure ML.
  • Introduce a “Data Science Catalog” for discovering and summarizing global data science resources within Domino.
  • Integrate model monitoring to provide a holistic view of deployment health and performance.
  • Enhance tagging capabilities across Domino entities to improve discoverability and tracking.
  • Expand LLM hosting capabilities to address customer needs for scale, performance, and logging.

What we look for in this role

  • 8+ years previously in a software engineering individual contributor role.
  • Strong Background in AI/ML: Extensive experience in designing, building, and deploying AI/ML models, with deep understanding of model development, training, optimization, and lifecycle management.
  • Building Scalable Systems: Hands-on experience developing and managing high-performance back-end systems in distributed computing environments.
  • Collaboration Across Teams: Working closely with cross-functional teams to integrate systems with front-end interfaces and third-party services.
  • API Development: Designing and implementing secure, scalable APIs (e.g., RESTful APIs, gRPC).
  • Performance Optimization: Profiling and optimizing back-end performance, especially in cloud environments or with container technologies like Docker and Kubernetes.
  • Testing and CI/CD: Using robust testing frameworks (unit, integration, end-to-end) and setting up CI/CD pipelines.
  • Distributed Computing: Experience with frameworks like Apache Spark, Azure ML, or SageMaker is a plus.
  • Cloud Platforms: Proficiency with cloud providers (AWS, Azure, GCP) and deploying services in these environments.

What we value

  • A growth mindset: high-performing creative individuals who dig into problems and see opportunities for success.
  • Individuals who seek truth, speak truth, and can be their whole selves at work.
  • Continuous improvement: everything is a work in progress, and we can do better at everything.
  • An environment of teaching and learning to equip teammates with tools for success.
  • A diverse and inclusive culture welcoming individuals of all backgrounds.

Timezone overlap

UTC-8–-3

Open to

NA · LATAM

Sign in to track applications and earn points.

More roles at Domino Data Lab

Similar remote roles