
About VRChat
VRChat offers a first-of-its-kind, game-changing platform that provides an endless collection of social VR experiences and gives the power of creation to its robust community. With over 250,000 worlds and growing, VRChat’s vision is to allow users to bring their imaginations to life and help shape the metaverse anywhere in the world on any device. VRChat has raised $100M to date with the support of investors Makers Fund, Anthos Capital, and HTC.
Our team includes people from Netflix, Twitter, Meta, Microsoft, Roblox, Google, Amazon, Unity, Spotify, Discord, Uber, eBay, Robinhood, Twitch, Zynga, and TikTok.
Job Overview
We are looking for a Senior/Staff Platform Engineer to help improve the reliability, performance, and scalability of our production platform.
This role focuses on operating reliable infrastructure, improving observability, driving incident response, and using data-driven reliability practices such as SLIs, SLOs, SLAs, error budgets, and DORA metrics. Database experience with MongoDB, Elasticsearch, or Redis is essential.
Help us run and secure the platform that connects our users and empowers them to create their part of the VRChat universe. The role reports to the Head of Platform and works closely with IT and Engineering teams, as well as functional leads, to plan and deploy infrastructure.
Duties & Responsibilities
- Operate and improve production infrastructure with a focus on reliability, security, performance, and cost efficiency.
- Define, measure, and improve reliability using SLIs, SLOs, SLAs, error budgets, and DORA metrics.
- Build and enhance monitoring, alerting, dashboards, logging, and incident response processes.
- Participate in incident management, root cause analysis, postmortems, and follow-up remediation.
- Automate infrastructure and operational workflows using modern Infrastructure as Code (IaC) and scripting tools.
- Collaborate closely with engineering teams to improve service reliability, deployment quality, and operational readiness.
- Turn ambiguous infrastructure, reliability, and operational challenges into clear, scalable, and measurable solutions.
- Engage with backend codebases through code reviews, pull requests, and occasional feature or tooling work to build shared context with product engineering teams.
Experience, Skills & Qualifications
- 8+ years of experience in SRE, DevOps, Platform Engineering, or Infrastructure Engineering.
- Strong track record operating high-availability production systems.
- Experience with cloud or hybrid cloud environments and tools such as Terraform or OpenTofu.
- Strong knowledge of Linux, networking, automation, observability, and incident management.
- Strong communication skills with the ability to collaborate across technical and non-technical stakeholders.
- Operational knowledge of databases such as MongoDB, Elasticsearch, or Redis.
Nice to Have
- Experience with AWS, including core infrastructure services, cost optimization, and multi-account architecture.
- Experience with Kubernetes, including networking, service discovery, ingress, and workload reliability.
- Experience with Cilium or other Kubernetes networking/security solutions.
- Experience supporting large-scale storage systems.
- Experience with CDNs, caching, distributed systems, or real-time platforms.
Benefits
- Work from anywhere (100% remote company)
- Comprehensive health benefits
- 401(k) for US employees & RRSP for Canadian employees
- Stock options
- Generous paid holiday schedule
- Unlimited/flexible vacation time
- Paid parental leave benefits
Benefits
Health, 401k, Equity, PTO, Unlimited PTO, Parental leave, Vision
Open to
Worldwide
Sign in to track applications and earn points.