Ondo Finance logo
Ondo FinanceΒ·

Site Reliability Engineer - Low-Latency Trading Systems

About Ondo Finance

Ondo Finance's mission is to provide institutional-grade, blockchain-enabled investment products and services. We develop decentralized finance technology and create/manage tokenized funds. We are the global leader in tokenized treasuries, tokenized stocks, and ETFs, building the future of institutional-grade financial services onchain.

Founded by veterans from Goldman Sachs Digital Assets Team, we are backed by leading investors including Founders Fund, Coinbase Ventures, Pantera Capital, and Tiger Global. We are fully remote with team members across the U.S.

About the Role

Ondo operates real-time trading systems that run 24/7 across traditional and crypto venues. The platform spans low-latency Rust engines, a fleet of Go services for trading, execution, and PnL accounting, and a multi-region Kubernetes footprint on AWS.

We are looking for an SRE with strong systems programming skills to own the reliability, observability, and performance of this platform. This is a hands-on role where you will read and modify Go and Rust code, debug latency regressions down to the feed handler, run incident response during market hours, and build automation for 24/7 trading systems.

Target Outcomes

  • Own production reliability for real-time trading services: trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines.
  • Operate and evolve multi-region Kubernetes clusters on AWS (EKS), deployed via GitOps (Flux) with SOPS-encrypted secrets.
  • Build and refine observability using Prometheus metrics, Datadog logs and dashboards, and SLOs to catch degradation early.
  • Improve deploy safety with progressive rollouts, config-reload behavior, and guardrails for live trading.

Responsibilities

  • Debug production incidents end-to-end: stale market data feeds, exchange rate limits, WebSocket disconnects, order-lifecycle desyncs, and latency regressions.
  • Harden market data ingestion from providers such as Databento and venue-native feeds (REST and WebSocket), including staleness detection, failover, and replay.
  • Build reconciliation and data-integrity tooling across live gauges, Postgres, and our S3 parquet data lake.
  • Participate in an on-call rotation covering US equity market hours and 24/7 crypto venues.

Requirements

  • 5+ years in SRE, production engineering, or infrastructure roles, with time supporting real-time or latency-sensitive systems.
  • Strong programming ability in Go or Rust, with willingness to work in both application code and infrastructure.
  • Deep, hands-on Kubernetes and AWS experience running stateful, latency-sensitive workloads in production.
  • Strong observability instincts including fluent PromQL, structured-log analysis, and alerting design.
  • Solid Linux internals and networking fundamentals to trace p99 regressions.
  • Sound judgment under pressure and clear written communication.

Tech Stack

  • Languages: Go, Rust, Python
  • Infrastructure: Kubernetes (EKS), Flux, SOPS, AWS (multi-region), S3 parquet lake
  • Observability: Prometheus, Grafana, Datadog
  • Databases: Postgres, CockroachDB, BigQuery
  • Market Data: Databento, venue WebSocket/REST feeds

What We Offer

  • Competitive compensation including salary, future token rights, and/or equity.
  • Full benefits including medical, vision, and dental.
  • Flexible vacation policy (PTO).
  • Remote-first team across many countries with an exceptional, experienced peer group.

Timezone overlap

UTC-8–-4

Open to

US

Sign in to track applications and earn points.

More roles at Ondo Finance

Similar remote roles