Yuno logo
Yuno·

Staff Site Reliability Engineer - Yuno

Who We Are

Yuno is the AI-native operating system of global commerce, powering the financial infrastructure of enterprise merchants, banks, and wallets. Through a single API, Yuno connects them to pay-ins, payouts, fraud prevention, KYC/KYB, and stablecoins globally, so they can operate everywhere. Agnostic by design and connected to 1,000+ payment methods and 460+ integrations in 190+ countries, Yuno optimizes acceptance rates, reduces costs, and strengthens security through specialized AI agents that learn from every transaction. Global brands including McDonald's, NetEase Games, GoFundMe, and Rappi run their payments on Yuno.

About The Role

Yuno is looking for a Staff Site Reliability Engineer to set the technical direction for reliability across our infrastructure — starting with the platform that provisions, deploys, and manages AI agents at scale on AWS, the system powering payments across 190+ countries. The platform is in production and growing, and we need the most senior reliability voice in the room to evolve the architecture and make sure it stays reliable, observable, and ready to scale.

This is not a "maintain what exists" role, and it's not a single-system role. You'll own the reliability strategy — driving architectural decisions, designing event-driven communication, defining how we measure and defend reliability, and setting the standards other engineering teams build on.

How AI Shows Up in This Role

  • AI agent infrastructure — The platform you own is Yuno's AI agent infrastructure, provisioning and deploying AI agents at scale, plus the agents that route payments and prevent fraud. Keeping the AI-native layer reliable is the core of the role.
  • AI-assisted tooling — AI is our default execution layer: you're encouraged to use AI-assisted tooling across automation, runbooks, incident analysis, and root-cause investigations, and to help define how the wider engineering org adopts it.

Your Contribution Will Be

  • Reliability strategy and standards — Define the SLO culture, error-budget policy, and incident practices that scale across engineering teams, turning reliability from firefighting into a measurable, org-wide discipline.
  • Platform architecture and evolution — Drive architectural decisions as the platform matures; act as the deciding voice on choosing technologies, designing systems, and when to evolve the infrastructure.
  • Messaging and event-driven architecture — Design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable, reliable async messaging.
  • Infrastructure and deployment — Own the cloud infrastructure, automate provisioning with IaC, and ensure the platform scales reliably as transaction volume grows.
  • Observability — Build the monitoring, tracing, and alerting that keeps the platform healthy; design dashboards and alerts that explain failures before manual digging is required.
  • Incident leadership and mentorship — Act as the senior escalation point for complex production problems, run blameless postmortems, and raise the reliability bar by mentoring senior and mid-level engineers.
  • Chaos engineering mindset — Conduct continuous fault injection and resilience experiments to surface weaknesses before they turn into incidents, and propose resilience patterns to prevent production failures.

What Success Looks Like

Within your first 6–12 months, you've set the reliability strategy for the platform, driven at least one major architectural evolution (event-driven messaging, streaming reliability, or observability), and engineering teams have adopted the SLO and error-budget framework you defined. You're the person Yuno trusts with the hardest reliability calls.

Skills You Need

Minimum Qualifications

  • Event-driven architecture and messaging systems — Experience designing and owning systems around message queues (Kafka, NATS, RabbitMQ) with an understanding of at-least-once delivery, consumer groups, dead letters, and backpressure; experience migrating a system from synchronous to async.
  • Deep AWS knowledge — EC2, VPC, IAM, S3, and RDS with strong networking fundamentals for inter-service communication over internal VPC.
  • Infrastructure as Code — Terraform or Pulumi, reviewed in PRs rather than clicked in consoles.
  • Kubernetes and Docker — Container lifecycle, resource limits, health checks, and orchestration at scale in production.
  • Observability and SLOs — Datadog fluency or equivalent (dashboards, monitors, APM, distributed tracing), and a track record of defining and operating SLOs, SLIs, and error budgets.
  • Chaos engineering — Hands-on experience with fault injection, game days, or chaos experiments (Gremlin, Chaos Mesh, AWS FIS, or similar).
  • Distributed systems debugging — Experience diagnosing async flows and cascading failures in production; comfortable coding for automation and tooling (Go, Python, or similar).
  • Databases — Solid SQL (PostgreSQL) and NoSQL (MongoDB, Redis) skills, including indexing, replication, and performance tuning.
  • Technical leadership — Proven experience setting reliability standards, influencing architecture across teams, and mentoring engineers.
  • English — Advanced proficiency, written and spoken.

Preferred Qualifications

  • AI / MLOps infrastructure — Running AI workloads in production (model serving, LLM inference, GPU/resource management, and agent evaluation/observability tools like LangFuse, LangSmith, Braintrust, or MLflow).
  • Multi-tenant container platforms — Running customer or user workloads in containers (Replit, Railway, Fly.io, or internal PaaS).
  • Data pipelines and orchestration — Airflow, Prefect, or similar; data warehouses like Databricks, Snowflake, or BigQuery.
  • Incident management — Experience with PagerDuty, Opsgenie, or incident.io.
  • Industry experience — Experience in the payments industry.

Nice to Have

  • ECS experience.
  • s6-overlay for container process supervision.
  • Experience with AI agent framework ecosystems.
  • Spanish proficiency.

What We Offer at Yuno

  • Competitive Compensation
  • Remote Work — you can work from everywhere
  • Home Office Bonus — a one-time allowance to set up your ideal home office
  • Work Equipment
  • Stock Options
  • Health Plan wherever you are
  • Flexible Days Off
  • Language, Professional, and Personal Growth courses

Timezone overlap

UTC+0–+3

Culture

Async-friendly

Open to

Europe

Sign in to track applications and earn points.

More roles at Yuno

Similar remote roles