
About Us
Circle is building the world’s leading AI-powered, all-in-one platform for digital businesses. We make it possible for creators, coaches, educators, and businesses to bring together their audience with engaging discussions, live streams, events, chat, courses, and payments — all in one place, all under their own brand.
We’re proud to be a fully remote company of around 270 (and growing!) team members from 30+ countries around the world. We collaborate across time zones, are highly async, and like to document a lot. Twice a year, we bring the whole company together in beautiful places around the world for our company offsites.
About the Role
The AI Quality engineering team at Circle owns the foundation for measuring, diagnosing, and improving the quality of Circle's AI-powered features. This team focuses on building the infrastructure to measure, diagnose, and improve production AI systems, rather than ML research or model training.
We're looking for a Lead Engineer to help us build out the evaluation frameworks, observability tooling, and diagnostic infrastructure that tell us whether our AI Agents are working well, where to improve them, and how to make them faster and more cost-efficient.
This is a hands-on player-coach role where you will also lead and manage the AI Quality engineering team with an ambitious and growing roadmap.
What You'll Be Doing
- Build and own our evaluation infrastructure: Design the CI/CD pipelines, scorers, and datasets that tell us whether Circle's AI agents (planners, tool-callers, and sub-agents) are actually working, from a single tool call to a full multi-turn conversation.
- Diagnose quality bottlenecks: Trace failures across the agent pipeline including plan creation vs. execution, tool selection, and tool trajectory in complex areas like workflows, site builder, and analytics.
- Grow evaluation datasets: Stand up annotation workflows and build out AI-generated and simulated conversations to cover more of the product faster than manual labeling alone.
- Run structured experiments: Evaluate new and open-source models against our production baseline, and chase cost and latency wins through model swaps, caching, and routing by plan complexity.
- Lead and grow the team: Set technical direction and manage day-to-day priorities while staying hands-on in the code.
- Partner closely with AI Core engineering: Work with product AI engineers to ensure changes measurably improve quality.
What You'll Need to Be Successful
- 7+ years of experience building and shipping production software, ideally including LLM-powered agents that take real actions in a product.
- Proficiency in Ruby on Rails / Python or readiness to pick them up quickly (Ruby on Rails is our core production system; Python is a strong plus).
- Experience building evaluation or observability infrastructure for ML/AI systems (e.g., Braintrust, LangSmith).
- Experience designing datasets or annotation workflows for ML/AI evaluation.
- Strong alignment with company values and C2-level proficiency in English.
Timezone overlap
UTC+0–+3
Culture
Async-friendly
Benefits
Open to
Europe
Sign in to track applications and earn points.