Monte Carlo logo
Monte Carlo·

Applied AI Engineer - Monte Carlo

Remote-firstFull-timeSeniorUTC-8–-4NA#python#llm#agentsEquity

About Monte Carlo

Monte Carlo is the agent trust platform that unifies data and agent observability to monitor, troubleshoot, and improve production AI systems. As enterprises prepare to deploy thousands of agents across business-critical use cases, Monte Carlo provides the reliability infrastructure to support them along this AI transformation, from human-guided agents to fully autonomous operations. Founded in 2019 and backed by leading investors, Monte Carlo empowers data and AI teams to ship trusted AI at scale.

The Role

We're building the products that tell enterprises whether their AI agents can be trusted — and we need someone who works end to end, from an ambiguous problem statement through research, prototyping, and production. You'd get the problem, not the spec: research the approaches, prototype, prove what works, build it, and integrate it into the platform alongside our engineering and data science teams.

What You'll Do

  • Take an open problem end-to-end — from research and prototyping through production, killing what doesn't work before it becomes someone's roadmap
  • Design and ship agent-powered features — root-cause analysis, incident triage, monitor generation — and integrate them into the platform with our engineering team
  • Build the eval infrastructure that makes those features safe to change: golden datasets, regression suites, offline and online scoring, and the judgment calls about what "good" means
  • Own retrieval and context pipelines over customer metadata, lineage, and query history, and instrument agent behavior in production — traces, failure taxonomies, cost and latency budgets — to close the loop on quality
  • Partner with data science on detection quality and experiment design, and with PM on what an agent should do versus what it merely can do
  • Set the technical bar for how we build with LLMs — patterns, guardrails, and the internal tooling other engineers reuse

What We're Looking For

  • You've built agents in production. Not integrated a framework. Not worked on a team that had one. Built them — agents with real autonomy and internal loops, where the model uses tools and decides what to do next without a human in the middle, and you kept them running once real users showed up.
  • You've run evals and monitored agents after launch. You've owned an eval framework — golden datasets, regression suites, offline and online scoring — not a folder of one-off scripts, and watched agents in production.
  • Python, plus an ML or data science background. Python is your daily language and you're solid on the backend. You understand models well enough to reason about how they behave.
  • You work from a problem, not a spec. Handed an ambiguous problem statement, you design the experiment, build the smallest version to test it, and take what works into production.
  • You use AI tools every day. Claude or its equivalents are part of how you write code and do research (backend and model layer work, no frontend).
  • You'd rather ship than polish. Most of this work needs a good answer quickly, not a perfect one eventually.

Nice to have: Statistics and hypothesis testing, building and maintaining MCP servers, and experience in the data and cloud space (Snowflake, Databricks, dbt, Airflow).

Why Monte Carlo

  • We created the data observability category and we're doing it again with agent observability
  • Series D, $236M raised, backed by Accel, Redpoint, Notable Capital, ICONIQ Growth, and Salesforce Ventures
  • Customers include HubSpot, Fox, Nasdaq, Toast, and Mercado Libre
  • Remote-first by design since day one, and recognized as a Best Workplace for it
  • Competitive compensation, equity, and a remote-first environment

Timezone overlap

UTC-8–-4

Benefits

Equity

Open to

NA

Sign in to track applications and earn points.

More roles at Monte Carlo

Similar remote roles