Elastic logo
Elastic·Verified

AI QA & Evaluation Engineer - Elastic

Remote (region-restricted)Full-timeSeniorUTC+8–+12APACBangaloreIndia#python#llmPTOParental leave

Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter.

What is The Role

We are looking for a skilled QA & Evaluation Engineer to join our team. The role blends strategic QA leadership with hands-on technical validation and structured evaluation to safeguard the accuracy, reliability, compliance, and ethical use of AI models. You will partner across IT and Engineering teams to identify, design, implement and run robust testing frameworks and evaluation rubrics for a portfolio of GenAI solutions that will be used across our organization.

What You Will Be Doing

  • AI Strategy Contribution: Be a primary contributor to our AI strategy, helping validate and test AI infrastructure, custom solutions, and third-party SaaS offerings.
  • Test Strategy & Execution: Design and implement comprehensive test strategies for AI/ML systems, including accuracy, bias, robustness, and regression testing.
  • Rubric-Based Evaluation: Design and implement self-contained evaluation tasks, including prompts, supporting files, and detailed grading rubrics to assess AI performance on functional workflows.
  • Automation & CI/CD: Automate validation suites for agentic/multi-agent systems, integration testing, and CI/CD pipelines for ML models.
  • Data Validation: Validate that AI/ML models are consuming accurate, authorized, and properly structured data sources; ensuring data quality across training and inference.
  • Observation & Reporting: Meticulously observe and document AI agent behaviors, producing crisp, precise summaries and reports on model performance and hallucinations.
  • Output Grounding: Validate prompt engineering outputs from a data accuracy standpoint, ensuring responses are grounded in verified data sources.
  • Refinement & Iteration: Iterate and refine evaluation tasks and rubrics based on feedback and team collaboration to ensure robust benchmarking methodologies.
  • Security & Governance: Ensure all AI data sources and structures meet governance, regulatory, and compliance standards, while implementing best practices for security and data privacy.
  • Cross-Functional Collaboration: Collaborate with teams from different areas including IT Engineering, IT Operations, Data & Integrations, PMO, CRM, Risk & Compliance, and business technology.
  • Industry Alignment: Stay current on the latest work in AI and make technical recommendations to the organization.

What You Bring

  • Proficiency in Python, TypeScript, or other programming languages used in AI and test automation.
  • Proven skill in designing or applying rubric-based evaluation, grading against set criteria, or building structured scoring frameworks.
  • Direct experience with LLM evaluation frameworks and benchmarking tools such as LangSmith, Confident AI, etc.
  • Knowledge of the GenAI stack and solutions including Retrieval Augmented Generation (RAG).
  • Experience with LLMs: Azure OpenAI, Vertex AI, ChatGPT Enterprise or similar.
  • High attention to detail and ability to notice subtle patterns or inconsistencies (such as data hallucinations or logic errors) that others might miss.
  • Advanced written communication skills, especially for documenting nuanced observations and feedback.
  • Experience with Cloud platforms (Azure, GCP, AWS).
  • Thorough understanding of DevOps/automation/CI/CD tools: GitHub, Terraform.
  • Comprehension with AI ethics, risk management, and data governance.

Benefits & Perks

  • Competitive pay based on the work you do and not your previous salary.
  • Health coverage for you and your family in many locations.
  • Ability to craft your calendar with flexible locations and schedules for many roles.
  • Generous number of vacation days each year.
  • Charitable donation matching up to $2000 (or local currency equivalent) for financial donations and service.
  • Up to 40 hours each year to use toward volunteer projects you love.
  • Minimum of 16 weeks of parental leave.

Timezone overlap

UTC+8–+12

Open to

Bangalore · India · APAC

Sign in to track applications and earn points.

More roles at Elastic

Similar remote roles