Mercury logo
MercuryΒ·

Senior Machine Learning Operations Engineer - Mercury

Fully remoteFull-timeSenior$167K - $208KUTC-8–-4NAUnited StatesSan Francisco+2 more#mlopsEquityRSUs

Mercury's use of machine learning in risk decisioning is growing fast in scope and in stakes. Models increasingly drive real-time decisions about fraud and financial crime, and the Machine Learning Platform (MLP) team exists to build a paved path from a trained model to a reliable production deployment, speeding up iteration, and ensuring granular production observability.

MLP owns the production ML lifecycle: the systems that take a model from registry through deployment, real-time inference, observability, and retraining. Our Data Science colleagues author and train the models. We build the platform that lets them register, deploy, and observe those models in production without carrying the operational burden themselves. We also serve low-latency, highly available scores to the decision engine that depends on them.

As part of this role, you will:

  • Build and operate the real-time inference service that scores models for the risk decision engine, with low latency and high availability as first-class requirements.
  • Own model deployment infrastructure: registry and versioning, CI/CD with performance, bias, and consistency checks, shadow mode, and staged rollouts.
  • Build model observability: availability, latency, and error monitoring, plus drift detection as a retraining trigger.
  • Partner with Risk Data Science to take models from a clean development-to-production handoff through to production operation under MLP ownership.
  • Implement experimentation capabilities such as champion/challenger and canary routing, and explainability outputs like SHAP attributions.
  • Feel a strong sense of product ownership and actively seek responsibility on a brand-new platform team.

The ideal candidate for the role has:

  • 5+ years in machine learning engineering, backend software engineering, MLOps, or a closely related field.
  • Production ML service experience: deploying, serving, and operating models in low-latency, high-availability contexts.
  • Strong backend engineering fundamentals in Python, with API frameworks like FastAPI or Flask.
  • Experience with model deployment and lifecycle tooling: model registries, CI/CD for models, versioning, and staged rollout patterns.
  • Experience building observability and alerting for production services: latency, errors, and model-specific signals like drift.
  • Comfort with the data layer ML depends on: SQL, key-value/low-latency stores (Redis, DynamoDB), and streaming pipelines (Kafka, Kinesis, Redpanda).

Nice to have:

  • Familiarity with a modern data stack (Snowflake, dbt, Dagster, Airflow, or similar).
  • Experience operating in a regulated, audit-sensitive, or compliance-adjacent environment.
  • Exposure to functional languages or willingness to work across a stack that includes Haskell, React, and TypeScript.

Timezone overlap

UTC-8–-4

Benefits

Equity, RSUs

Open to

NA Β· San Francisco Β· United States Β· New York +1

Sign in to track applications and earn points.

More roles at Mercury

Similar remote roles