
At Dscout, we’re building the most flexible and powerful UX research platform on the market—trusted by the world’s top brands in finance, healthcare, consumer goods, and tech. Our tools help teams deeply understand the humans behind their products, so they can build better ones.
AI native product is fundamentally a different engineering problem than building deterministic software: the same input won't always produce the same output, and "working" means the agent behaves well across the full distribution of real world scenarios, not that it passes a fixed test suite.
We're looking for an Applied AI Engineer with 2-5 years of experience building and shipping AI systems used by enterprise professionals. You're comfortable working with modern LLM-based systems and agentic workflows, and you know how to turn powerful models into reliable product features.
What you'll do
- Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes.
- Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context.
- Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring.
- Investigate why agents underperform across context, knowledge, instructions, tools, routing, guardrails, or workflow design.
- Design and ship targeted behavior improvements, including changes to prompting, context construction, decision logic, tool use, and human-review paths.
- Build backend services, APIs, data models, and feedback pipelines that make agent behavior observable, steerable, and reproducible.
- Run controlled experiments, production replays, or staged rollouts to measure whether changes improve quality and downstream business results.
- Partner with Product, Data Science, and Sales to prioritize high-value problems and define customer and business success.
- Ship with appropriate safeguards for privacy, security, reliability, human oversight, and safe operational rollout.
What you bring
- 2-5 years of software engineering experience, with hands-on experience building or operating LLM-powered features or agents in production.
- Fluency with prompting and context engineering as an engineering discipline.
- Experience building or maintaining evaluation harnesses for AI systems: offline eval sets, LLM-as-judge or human-in-the-loop scoring, regression detection.
- Genuine comfort with non-determinism and variance in production traffic.
- Experience running experiments (A/B, staged rollouts, production replay) to validate outcomes.
- A track record of shipping features real users depended on, and owning what happened after launch.
- A high-agency mindset with comfort investigating ambiguous "why is this underperforming" problems.
- Comfort using AI coding tools (Cursor, Claude Code, Copilot, or similar) as part of your workflow.
Nice to have
- Experience with voice or real-time conversational AI systems.
- Familiarity with LLM observability/tracing tools (e.g., Braintrust, LangSmith, Datadog LLM Observability).
- Experience with agentic orchestration frameworks (LangChain/LangGraph or similar).
- Exposure to MCP-based tooling or agentic data workflows.
Timezone overlap
UTC+8–+12
Benefits
Equity, Bonus, Health, 401k, Unlimited PTO, PTO, Parental leave, Learning
Open to
APAC
Sign in to track applications and earn points.