
About uRun
AI inference today is slow, expensive, and stateless. Send a query, wait, get a response, reset. That's fine for batch, but AI is becoming interactive, and interactive means inference has to respond instantly, hold context across a session, and be steerable in real time.
uRun (Universal Runtime) is the layer that makes real-time, stateful inference possible. Our platform lets AI respond instantly, hold context across a session, and be directed as it runs. We prove it through the hardest problem in the stack: real-time AI video generation. We're an infrastructure company building the layer model labs, builders, and research teams ship on top of.
Where You Come In
You'll build the services, APIs, and core application systems that power uRun's runtime—the software layer that turns our real-time inference platform into something product and applied AI teams can actually build on.
This is not a conventional CRUD backend role. The work centers on low-latency, high-throughput systems: real-time interaction, evolving session state, and request handling that stays reliable under heavy compute and concurrency. You'll work closely with product, infrastructure, and applied AI teams in an early-stage environment where the architecture is still being set.
What You'll Be Doing Day-to-Day
- Build and maintain backend services, APIs, and internal platform components that power uRun's real-time inference runtime.
- Design systems for real-time interaction, evolving session state, and scalable request handling across production environments.
- Partner with infrastructure and platform engineers to keep services observable, reliable, and efficient under heavy compute and concurrency.
- Translate experimental AI capabilities into robust, user-facing software, working closely with product and applied AI teams.
- Shape architecture decisions on data flow, service boundaries, performance optimization, and fault tolerance for interactive systems.
- Raise engineering quality through testing, monitoring, code review, documentation, and sound operational practice.
Qualifications & Skills
- 7+ years building and shipping backend software in production.
- Language Proficiency: Strong experience in one or more backend languages—TypeScript/Node.js, Python, or Go.
- Systems Design: Experience designing APIs, service-oriented systems, and distributed application components.
- Cloud & Infrastructure: Solid understanding of cloud infrastructure, containers, and modern deployment workflows.
- Performance & Reliability: Ability to reason about performance, concurrency, reliability, and debugging in complex systems.
- Real-time Systems: Experience with real-time, interactive, streaming, or latency-sensitive systems is central to this role.
Nice to Have
- Experience with WebRTC or WebSockets for real-time communication.
- Familiarity with AI infrastructure, inference-adjacent systems, media pipelines, or event-driven architectures.
- Experience with Kubernetes, observability tooling, and hands-on production operations.
- Early-stage startup experience—owning problems end-to-end and moving quickly with limited scaffolding.
Benefits & Perks
- Competitive Salary & Equity: Meaningful equity grant in an early-stage AI infrastructure company.
- Health Coverage: Full health, dental, and vision insurance.
- Retirement: 401(k) company-supported retirement savings.
- FSA/HSA: Flexible spending accounts for healthcare costs.
- Paid Time Off: Flexible PTO policy.
- Top-tier Tooling: Access to Claude, Codex, Kimi, and leading AI development tools.
- Hardware: MacBook Pro and AirPods provided.
Timezone overlap
UTC-8–-4
Open to
US · San Francisco · United States
Sign in to track applications and earn points.