
About Cribl
Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the world’s biggest enterprises, including half of the Fortune 100, to bridge the gap between AI ambition and infrastructure reality. As the AI Platform for Telemetry, we give customers the choice, control, and flexibility to manage and analyze telemetry for both humans and agents, so they can build what’s next.
We’re one of the fastest‑growing private companies and a leading player in a massive, fast‑moving market. With a global workforce, we’re remote‑first and grounded in a simple idea: software is a people business. Cribl is the place where curious, collaborative people can do their best work, grow fast, and bring their full selves to the herd.
Why You’ll Love This Role
You will help build Cribl Engineering’s AI platform and productivity rails: the shared systems, runtimes, and workflows that let engineers use autonomous AI across the software delivery lifecycle. Instead of one-off prompts or single-feature AI experiments, you will focus on intent engineering, orchestration and harnesses, and production agent infrastructure that other teams can trust and extend.
You will work with a small, high-impact team alongside tech leads, principals, and partner engineering groups. You will ship systems that plan, implement, review, test, and operate work with strong guardrails, then drive adoption until those tools become part of how Cribl builds software. You will also bring a builder mentality: use the product yourself to solve real team and Engineering problems, then turn that lived experience into better systems.
What You Will Do
- Design & Build: Design, build, and operate production agentic workflows and the platform harnesses they run on (orchestration, tool integrations, shared context, extension points for other teams).
- Intent Engineering: Turn goals into clear specifications, rules, constraints, and acceptance criteria that AI systems can execute reliably.
- Evangelize Patterns: Stay current on AI tooling and practices, and evangelize what works across Engineering so partners adopt proven patterns.
- Collaborate: Engage hard in design discussions and healthy debate while the direction is open, then pivot with equal energy into implementation once the team decides.
- Observability & Guardrails: Build tracing, regression detection, human-in-the-loop controls, safe rollout, and operability for agentic systems others depend on.
- Own Agent Runtime: Manage job isolation, scheduling, execution environments, secrets and access, and production operations in the cloud.
- Event-Driven Architectures: Design and operate event-driven architectures where it fits (queues, streams, webhooks, async job fan-out) so agentic systems stay scalable and loosely coupled.
- Continuous Delivery: Keep agentic systems on automated CI/CD paths with progressive delivery and clear rollback.
- Dogfooding: Bring a builder mentality and use the platform to unblock yourself and the team, surface sharp edges, and turn real usage into product and platform improvements.
- Optimize Delivery: Compress software delivery loops by improving how AI helps engineers write, test, review, debug, and validate changes in real repositories and pipelines.
- Partner Across Teams: Understand workflows, ship tools that fit how people work, and drive adoption through playbooks, examples, demos, and enablement.
- Evaluate Tech: Evaluate build vs. buy; stay current on models, agent frameworks, MCP-style tool protocols, and the broader AI tooling ecosystem.
- Track Metrics: Define and track success metrics for adoption, quality, reliability, and satisfaction; use feedback and data to decide where to invest.
- Security & Access: Partner on security and data access so tools are useful while respecting permissions and company policy.
- End-to-End Ownership: Take designed projects from zero to production, owning the path from agreed design through implementation, rollout, and day-two operability.
- On-Call: This position may include stand-by, on-call, or off-hours duties for systems you own.
What We Are Looking For
- Experience: Staff-level (or equivalent) professional software engineering experience building and operating production distributed systems.
- Technical Stack: Strong TypeScript (and modern JavaScript) plus Node.js experience shipping production services. Polyglot comfort is welcome, but TypeScript is the primary stack for this team.
- Engineering Fundamentals: Strong software engineering fundamentals and the ability to ship quickly (design, testing, debugging, APIs/services, and code quality).
- CD Culture: Experience thriving in a continuous deployment culture with automated pipelines, progressive delivery, monitoring, rollback, and a bias toward frequent, reversible production changes.
- AI/LLM Fluency: Hands-on fluency with modern LLM and agentic coding workflows in production or serious internal platforms (not curiosity-only).
- Backend & Integrations: Experience building backend services, integrations, automation, and internal tools; comfort across product surfaces when needed.
- Agent Orchestration: Professional experience with agent orchestration, tool-calling systems, evaluation or guardrail techniques, and connecting agents to reliable backend systems.
- Frameworks: Familiarity with agent frameworks, orchestration layers, and integrating external tools and data sources into LLM-based systems (MCP or equivalent experience is a plus).
- Event-Driven Systems: Experience with event-driven systems (queues, streams, pub/sub, webhooks, or similar async patterns in production).
- Observability: Fluency with metrics, logs, traces, and using them to operate and improve production systems.
- Communication: Clear communication, documentation, and teaching ability; comfort driving adoption, not only writing code.
- Security Mindset: Good judgment around security, permissions, data access, and safe tool rollout.
- Problem Solving: Ability to problem-solve from first principles, make sound trade-offs, and drive work independently through ambiguity.
Nice to Have
- Hands-on Kubernetes in production.
- Terraform or similar infrastructure-as-code for cloud provisioning.
- Temporal or other workflow-platform experience.
- Deeper AWS / cloud-native ops fluency (IAM, networking, running production workloads end to end).
Timezone overlap
UTC-8–-4
Culture
Async-friendly
Open to
US
Sign in to track applications and earn points.