Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
U

QA Engineer โ€“ Generative AI & Agentic AI

UST
  • ๐Ÿ‡ฎ๐Ÿ‡ณ India
  • On-site
  • Staff / Principal
  • 2 days ago
  • Machine Learning
  • SQL
  • RAG
  • AI
  • Snowflake
  • OpenAI
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for an experienced QA Engineer โ€“ Generative AI & Agentic AI to lead quality assurance activities for GenAI and Agentic AI solutions.

This role goes beyond traditional functional testing. The successful candidate will be responsible for validating the correctness, reliability, quality, and production-readiness of AI-generated outputs, including multi-step agent workflows, Text-to-SQL pipelines, Retrieval-Augmented Generation (RAG) systems, and LLM-based applications.

The ideal candidate will have hands-on, production-grade experience testing GenAI/Agentic AI applications end-to-end and should have successfully supported at least one production release of a GenAI or Agentic AI solution.

Key Responsibilities
  • Design, develop, and execute comprehensive test strategies for GenAI and Agentic AI applications, covering functional, non-functional, integration, and AI output-quality testing.
  • Test multi-step agentic workflows, validating agent orchestration, tool/function calls, task completion, chained decision-making, reasoning consistency, and outputs across individual agent hops.
  • Evaluate AI-generated outputs for correctness, relevance, coherence, hallucination, consistency, and reliability.
  • Test Text-to-SQL workflows, including validation of:
    • Natural language to SQL generation accuracy
    • SQL syntax and logic
    • Schema/table/column alignment
    • Query execution
    • Result accuracy against Snowflake data models
  • Validate RAG pipelines, including:
    • Retrieval relevance
    • Context grounding
    • Chunking effectiveness
    • Context utilization
    • Faithfulness of generated responses to retrieved information
  • Apply GenAI evaluation metrics and methodologies such as:
    • Embedding similarity
    • ROUGE
    • BLEU
    • LLM-as-a-Judge
    • Other relevant GenAI quality and accuracy metrics
  • Design and maintain LLM-as-a-Judge evaluation frameworks and scoring rubrics for scalable and automated assessment of GenAI outputs.
  • Create and maintain golden datasets, evaluation datasets, regression suites, and reusable AI testing frameworks.
  • Develop test harnesses/scripts for batch evaluation and automated quality validation of LLM-generated outputs.
  • Own QA activities for end-to-end production releases, including test planning, execution, regression testing, release quality assessment, sign-off, and post-production monitoring.
  • Partner closely with Data Science, Platform Engineering, Product, and other stakeholders to establish acceptance criteria, quality gates, and go/no-go release benchmarks.
  • Identify LLM-specific edge cases and failure modes, including hallucinations, prompt injection risks, bias, inconsistent responses, and agent/tool failures.
  • Document and track defects through resolution and provide clear evidence of AI application quality.
  • Maintain traceability of test coverage, evaluation results, quality metrics, and release reports for stakeholders and audit requirements.
Mandatory Skills & Experience

GenAI / Agentic AI Testing

Strong, demonstrable hands-on experience with:

  • Generative AI / LLM application testing
  • Agentic AI and multi-step agent workflows
  • Agent orchestration and tool/function-call validation
  • Text-to-SQL testing
  • Retrieval-Augmented Generation (RAG) testing
  • LLM-as-a-Judge evaluation
  • Prompt and LLM behavior validation
  • Hallucination and response-grounding assessment
  • Context-window and context-handling concepts
  • AI output-quality evaluation beyond traditional pass/fail testing

Platforms

  • Dataiku: Strong hands-on experience is highly preferred and considered the first preference
  • Exposure to platforms such as OpenAI workspace, Leena AI, or similar GenAI/Agentic AI platforms

Data / Database

  • Strong hands-on working experience with Snowflake
  • Ability to understand and validate data models, schemas, generated SQL queries, and query results

AI Evaluation

Practical experience with at least some of the following:

  • Embedding Similarity
  • RO
Provide your feedback on BizChat

QA Engineer โ€“ Generative AI & Agentic AI ยท UST

Auto apply with Likeremote