QA Engineer โ Generative AI & Agentic AI
- ๐ฎ๐ณ India
- On-site
- Staff / Principal
- 2 days ago
- Machine Learning
- SQL
- RAG
- AI
- Snowflake
- OpenAI
We are looking for an experienced QA Engineer โ Generative AI & Agentic AI to lead quality assurance activities for GenAI and Agentic AI solutions.
This role goes beyond traditional functional testing. The successful candidate will be responsible for validating the correctness, reliability, quality, and production-readiness of AI-generated outputs, including multi-step agent workflows, Text-to-SQL pipelines, Retrieval-Augmented Generation (RAG) systems, and LLM-based applications.
The ideal candidate will have hands-on, production-grade experience testing GenAI/Agentic AI applications end-to-end and should have successfully supported at least one production release of a GenAI or Agentic AI solution.
Key Responsibilities- Design, develop, and execute comprehensive test strategies for GenAI and Agentic AI applications, covering functional, non-functional, integration, and AI output-quality testing.
- Test multi-step agentic workflows, validating agent orchestration, tool/function calls, task completion, chained decision-making, reasoning consistency, and outputs across individual agent hops.
- Evaluate AI-generated outputs for correctness, relevance, coherence, hallucination, consistency, and reliability.
- Test Text-to-SQL workflows, including validation of:
- Natural language to SQL generation accuracy
- SQL syntax and logic
- Schema/table/column alignment
- Query execution
- Result accuracy against Snowflake data models
- Validate RAG pipelines, including:
- Retrieval relevance
- Context grounding
- Chunking effectiveness
- Context utilization
- Faithfulness of generated responses to retrieved information
- Apply GenAI evaluation metrics and methodologies such as:
- Embedding similarity
- ROUGE
- BLEU
- LLM-as-a-Judge
- Other relevant GenAI quality and accuracy metrics
- Design and maintain LLM-as-a-Judge evaluation frameworks and scoring rubrics for scalable and automated assessment of GenAI outputs.
- Create and maintain golden datasets, evaluation datasets, regression suites, and reusable AI testing frameworks.
- Develop test harnesses/scripts for batch evaluation and automated quality validation of LLM-generated outputs.
- Own QA activities for end-to-end production releases, including test planning, execution, regression testing, release quality assessment, sign-off, and post-production monitoring.
- Partner closely with Data Science, Platform Engineering, Product, and other stakeholders to establish acceptance criteria, quality gates, and go/no-go release benchmarks.
- Identify LLM-specific edge cases and failure modes, including hallucinations, prompt injection risks, bias, inconsistent responses, and agent/tool failures.
- Document and track defects through resolution and provide clear evidence of AI application quality.
- Maintain traceability of test coverage, evaluation results, quality metrics, and release reports for stakeholders and audit requirements.
GenAI / Agentic AI Testing
Strong, demonstrable hands-on experience with:
- Generative AI / LLM application testing
- Agentic AI and multi-step agent workflows
- Agent orchestration and tool/function-call validation
- Text-to-SQL testing
- Retrieval-Augmented Generation (RAG) testing
- LLM-as-a-Judge evaluation
- Prompt and LLM behavior validation
- Hallucination and response-grounding assessment
- Context-window and context-handling concepts
- AI output-quality evaluation beyond traditional pass/fail testing
Platforms
- Dataiku: Strong hands-on experience is highly preferred and considered the first preference
- Exposure to platforms such as OpenAI workspace, Leena AI, or similar GenAI/Agentic AI platforms
Data / Database
- Strong hands-on working experience with Snowflake
- Ability to understand and validate data models, schemas, generated SQL queries, and query results
AI Evaluation
Practical experience with at least some of the following:
- Embedding Similarity
- RO
QA Engineer โ Generative AI & Agentic AI ยท UST