
About Upsun (formerly Platform.sh)
Upsun is the software factory for AI-human workflows. It is built for todayβs hybrid teams, where AI agents write and test code and humans focus on solving the problems that really matter. Developers, DevOps engineers, and platform teams use Upsun to build, ship, and scale confidently without wrestling with backend infrastructure.
What You Get
- Predictable performance, even at scale
- Secure, compliant environments by default
- Real-time observability and profiling built in
- Cloning, configuration, and provisioning in seconds
- AI-ready features that plug directly into your stack
Upsunners are a remote, global workforce. We are committed to open source and an open, welcoming environment spanning the globe and experience spectrum.
Our Values
- πΏ We make a positive impact.
- β¨ We aim for the stars.
- π We care for each other.
Impact of a Senior Site Reliability Engineer
As a Senior Site Reliability Engineer at Upsun, you will lead the evolution of our cloud application platform from traditional cloud operations into a proactive, automation-driven SRE model. You will own critical engineering workstreams that enhance system reliability, scalability, and operational efficiency across multi-cloud environments. Partnering closely with engineering, product, and platform teams, you will embed reliability and performance into every stage of the software delivery lifecycle. In this role, you will anticipate architectural bottlenecks, drive infrastructure-as-code practices, and establish robust observability standards that ensure long-term system stability and uptime for our global users.
What to Expect
- Drive reliability & observability strategy: Architect and elevate system monitoring, alerting, and logging using Prometheus, Grafana, and ELK Stack, establishing actionable SLIs/SLOs aligned with core business metrics.
- Automate infrastructure & workflows: Eliminate operational toil by designing and implementing resilient, automated solutions using IaC tools like Terraform and Ansible across AWS, GCP, and Azure.
- Scale CI/CD & delivery pipelines: Optimize pipeline architectures for fast, secure, and zero-downtime releases, ensuring infrastructure resilience during high-volume deployment cycles.
- Lead incident response & post-mortems: Guide high-priority incident triage, drive blameless post-mortem analysis, and implement preventative measures to continuously improve system resiliency.
- Cross-functional leadership: Partner with product and software engineering teams to incorporate SRE best practices into product roadmaps.
- Champion technical innovation: Proactively identify performance bottlenecks and evaluate emerging technologies (e.g., eBPF, container orchestration) to optimize platform stability and performance.
- Time distribution: Follow a 4-week rotation balancing engineering and operations to focus on reliability, automation, and scalability through hands-on troubleshooting and engineering innovation.
What You Bring
- Senior SRE & Cloud Expertise: 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps, with proven experience owning reliability for production platforms at scale.
- Software Engineering & Tooling: Strong proficiency in Go or Python to build custom automation tools, custom controllers, or SRE platform components (beyond basic shell scripting).
- Deep Linux Internals Proficiency: Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
- Infrastructure as Code & Cloud Platforms: Deep expertise with cloud providers (AWS, GCP, Azure, or Openstack) with custom tooling built around cloud SDKs, and declarative infrastructure tools (e.g., Terraform) to manage distributed systems.
- Autonomous Ownership & Systems Thinking: Proven ability to anticipate operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal guidance.
- Collaborative Communication: Outstanding cross-functional communication skills with a track record of building alignment and fostering an inclusive engineering culture.
Bonus Points
- Experience with custom-built orchestration, edge, storage, and operational tooling in a dynamic environment.
- Experience with Docker and production Kubernetes cluster management or containerized deployment architectures.
- Familiarity with Platform-as-a-Service (PaaS) architectures or developer-facing cloud platforms.
Hiring Location & On-Call Requirements
- Location: We are currently focused on hiring for this role in Western Australia. Candidates must be legally authorized to work in Australia (visa sponsorship is not available at this time).
- On-Call: This role includes on-call rotation (one week every 4-5 weeks, 02:00 AM β 10:00 AM UTC, including weekends during the assigned week).
Interview Process
- 45 minutes with Talent Acquisition
- 60 minutes with Hiring Manager
- 60 minutes with Team (ICs)
- 60 minutes with Senior Director, SRE
Note: All roles require background checks.
Timezone overlap
UTC+8β+12
Benefits
Equity, PTO, Learning, Equipment, Wellness, Internet, Parental leave, Bonus, Unlimited PTO, Visa
Open to
APAC Β· Australia
Sign in to track applications and earn points.