
Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing.
Many infrastructure platforms simply focus on how you deploy your singular application, and not how these applications function in concert. Questions like "How do you build systems for zero downtime deployment" or "How do you do service-to-service communications" are usually left up to the engineers to define. At Railway, our goal is to be an all-encompassing solution to all these problems. As such, we take special care as we define our networking and platform infrastructure.
"But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing."
— Radia Perlman
About The Role
In this role, you will:
- Build ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry.
- Build scalable, fault-tolerant alerting engines for notifying users in real-time of threshold breaches.
- Craft rich backend observability APIs, working with product to build amazing experiences for instantly understanding application health.
- Provide APIs to access real-time log/metrics streams consumed by the Dashboard and Product teams.
- Build Golang/Rust gRPC services from scratch capable of supporting tens of thousands of users, and the millions to come.
- Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principles of immutable infrastructure (Terraform, Ansible).
- Write Engineering Requirement Documents (ERDs) to take initiatives from idea to defined tasks, implementation, and success monitoring.
- Interface with our TypeScript and GraphQL edge to expose your microservice APIs for internal and external consumption.
This is a high-impact, high-agency role with direct effect on company culture, trajectory, and outcome.
About You
- A strong understanding of distributed systems and a passion for building fault-tolerant, resilient, and scalable services.
- Interest in VictoriaMetrics, ClickHouse, and other tools for building observability stacks from the ground up.
- Solid intuition about system lifecycles and scaling timelines in an early-stage startup.
- The discipline to implement solutions, create monitors for error boundaries, and document requirements clearly.
- Great prioritization skills when dealing with ambiguity in a fast-growing environment.
- A sense of grit to dive into a problem, scale a solution, and refactor/replace it when needed.
- Excellent communication skills to collaborate effectively across a distributed team.
Things to Know
- Globally Distributed: We're distributed across the globe. You'll need to be diligent about setting boundaries, as your workday may overlap with someone else's start.
- High Ownership: We are a lean team serving hundreds of thousands of users. We give team members full ownership of their choices, successes, and learnings.
- Minimal Meetings: We keep meetings to a minimum (a brief sync on Monday and Friday) to protect your focus time.
Benefits & Perks
- Competitive compensation and strong equity grants
- Full health benefits (including coverage for dependents)
- Equipment stipend
- High autonomy, high agency culture with creative and high-leverage problems
Open to
Worldwide
Sign in to track applications and earn points.