
About Fundraise Up
Fundraise Up is a global fundraising platform designed to make donating to nonprofits fast, seamless, and accessible. Our technology processes tens of millions of dollars in donations monthly for leading nonprofits like UNICEF and the Alzheimer's Association. We are a distributed team of 160+ product professionals, including 80+ engineers, managing a large ecosystem of checkout widgets, donor portals, admin tools, and internal applications.
Our technology stack includes Node.js (Koa, NestJS), MongoDB, TypeScript, Vue.js, and React, with Kafka and Bull (Redis) for messaging and background jobs, ClickHouse for analytics, and Elasticsearch for search.
About the Role
This is a senior-only role within our DevOps team, which is responsible for the platforms every engineering team relies on daily: CI/CD, observability, logging, and developer tooling. On a busy giving day, ten minutes of downtime can cost around $500k, highlighting the critical importance of this role.
You will own an entire core platform area—observability or CI/CD—from requirements and design through rollout, operations, and mentoring. You will serve as the go-to technical reference and senior escalation point for your area.
What You’ll Do
- Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents). You will drive its architecture, reliability, and roadmap.
- Drive technical initiatives end-to-end: gather requirements, write design documentation, decompose into tasks, implement, deliver to production, and own operational health.
- Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps.
- Design for reliability and scale: evolve the architecture of our platforms, including topology, integration points, scaling approach, and reliability model.
- Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, assist teams with metrics, alerts, and logs, and participate in chat duty for developer support.
- Automate away toil: codify repetitive operations, provisioning, and maintenance.
- Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, and implement systemic fixes. Participate in on-call rotations and raise the bar for on-call practices.
- Mentor less experienced engineers through design discussions, reviews, and pairing; identify and address debt-inducing shortcuts at the review stage.
- Use AI in all aspects of day-to-day work: researching, troubleshooting, and developing.
Requirements
- 6+ years as a DevOps Engineer / SRE (or equivalent responsibilities).
- Track record of owning technical initiatives end-to-end—from requirements and technical design through production delivery. You can showcase initiatives that were yours, not just tasks you completed.
- Confident Linux skills (we use Ubuntu).
- Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works—enough to navigate and extend an existing setup.
- Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery.
- Containers: Docker, image building, registries.
- Ansible.
- Git.
- Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work).
- Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems.
- Experience mentoring less experienced engineers.
- Ownership and attention to detail. Downtime is expensive: during busy events, 10 minutes of downtime can cost around $500k.
Must Have
Solid hands-on experience in two or more of the areas below:
- VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure.
- Log pipelines at scale: Graylog / VictoriaLogs / ELK—collection (fluent bit or similar), retention, sharding, performance.
- Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets.
- Container registries and artifact management: Harbor, Nexus, base images, image policies.
- Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting.
- Grafana: dashboards as code, alerting, performance at scale.
Bonus Points
Experience with any of the following is a plus:
- Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte—deployment, maintenance, resource limits. Building platforms around these tools to improve Quality of Life for Analytics.
- Remote development environments and AI agent execution environments (e.g., Coder/Telepresence).
- Bare-metal Kubernetes: provisioning, networking, scaling.
- Flux and GitOps.
- Terraform.
- Sentry on-premise: operating self-hosted error tracking.
- ClickHouse, MongoDB.
You'll Thrive Here If You…
- Want an area that's truly yours, and the accountability that comes with it.
- Commit to your own estimates, and flag early when something is slipping instead of hoping it won't.
- Ask for help at 2 a.m. rather than quietly going down the rabbit hole alone.
- Explain a complex incident in two minutes, not twenty.
- Disagree openly, make your case with data, then commit and move.
- Treat every manual step as a bug waiting to be automated.
- Get more satisfaction from making other engineers self-sufficient than from being the only one who knows how it works.
- Enjoy a fast-moving team that's rebuilding how it works around AI, and stay curious as the tools change.
Perks & Benefits
- Private medical insurance for the employee and their family
- 20 paid vacation days per year
- 15 paid public holidays per year
- 5 company-paid sick leave days
- English learning courses
- Relevant professional education
- Gym or swimming pool membership
- Home Office Setup Assistance (furniture like office chair, desk, monitor, and other items)
- Co-working access
- Remote working
Open to
Serbia
Sign in to track applications and earn points.