
About Railway
Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing.
Many infrastructure platforms simply focus on how you deploy your singular application, and not how these applications function in concert. Questions like "How do you build systems for zero downtime deployment" or "How do you do service-to-service communications" are usually left up to engineers to define. At Railway, our goal is to be an all-encompassing solution to these problems.
“But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing.”
— Radia Perlman
About the Role
This role is focused on building the internal-facing platform that Railway engineers run on. It is a high-impact, high-agency position with direct effect on company culture, trajectory, and outcome.
Key Responsibilities
- Build and maintain our host provisioning stack (PXE boot, Ansible, and burn-in agents) to bring new bare metal online quickly and confidently.
- Evolve our homegrown orchestration engine to manage clusters, containers, and VMs through a single lens.
- Optimize bin-packing algorithms to maximize utilization/performance and minimize costs.
- Own the internal tooling that Railway engineers use to interact with the fleet daily.
- Build internal observability and alerting to detect fleet issues before customers experience them.
- Design and maintain CI pipelines to ship infrastructure code safely.
- Define infrastructure that can be torn down, failed over, and reconstituted using immutable infrastructure principles (Terraform, Ansible).
- Build Golang and Rust gRPC services from scratch capable of supporting millions of users.
- Write Engineering Requirement Documents (ERDs) to take initiatives from idea to execution and monitoring.
What We're Looking For
- Strong understanding of distributed systems and operating fault-tolerant, resilient, and scalable services.
- Hands-on experience with bare-metal provisioning, configuration management, and hardware production readiness.
- Experience building and operating internal tools focused on developer experience.
- Sound intuition on system longevity and pragmatic startup architecture design.
- Clear communication skills and the tact to implement solutions, document requirements, and set up monitor error boundaries.
- High grit and comfort operating in an early-stage startup environment with ambiguity.
Culture & Things to Know
- Globally Distributed: We operate asynchronously across timezones worldwide.
- High Ownership: Small team (~21 people) serving hundreds of thousands of users with high autonomy and responsibility.
- Sustainable Boundaries: Diligence around work-life boundaries is encouraged as global hours overlap.
Benefits & Perks
- Competitive salary and strong equity grants.
- Full health benefits (including dependents).
- Equipment stipend.
- High-autonomy culture with minimal meetings.
Culture
Async-friendly
Open to
Worldwide
Sign in to track applications and earn points.