Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
AP

QA / Automation Engineer— Agentic AI

ATech Placement
🇺🇸 United States
Hybrid
3 weeks ago
  • AI
  • Large Language Models
  • Python
  • Playwright
  • pytest
  • CI/CD
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

QA Automation Engineer – Agentic AI Testing

Position Overview

We are seeking aQA Automation Engineer to help build and validate a next-generationagentic AI application within a growing AI-driven HCM software ecosystem.

This isnot a traditional QA automation role. The application uses large language models (LLMs) and AI agents to dynamically determine application behavior and workflows, creating testing challenges that cannot be solved through conventional scripted automation alone.

The team is building anew Python-based testing and evaluation framework from the ground up specifically for agentic AI systems. This includes LLM-as-judge evaluations, agent state checks, behavioral validation, and testing approaches designed for non-deterministic AI outputs.

Python and hands-on AI/LLM testing experience are the primary requirements for this role. Playwright is used for front-end automation but is a secondary/supporting skill rather than the core focus.

Key Responsibilities

  • Design, build, and maintain automated testing and evaluation solutions foragentic and LLM-driven applications.
  • Helpbuild a new Python-based AI testing and evaluation framework from the ground up, including defining patterns, standards, and approaches as the platform evolves.
  • DevelopLLM-as-judge evaluators and other automated methods for assessing AI-generated outcomes.
  • Validateagent states, conversational flows, tool interactions, outputs, and expected behaviors across multi-step agentic workflows.
  • Develop testing strategies fornon-deterministic LLM behavior where conventional scripted pass/fail assertions are insufficient.
  • Evaluate AI responses for factors such asaccuracy, relevance, consistency, groundedness, and appropriate behavior.
  • Test agent behavior across expected paths, edge cases, failure scenarios, and variable outputs.
  • UsePlaywright for front-end/UI automation where appropriate while maintaining a separate Python-based evaluation layer for AI behavior.
  • Partner closely with software engineers and AI developers to identify risks, troubleshoot failures, and improve product quality.
  • Help establishtesting standards, evaluation methodologies, and automation practices for an evolving agentic AI platform.
  • Take shared ownership of quality within a collaborative engineering model rather than operating within a traditional siloed QA organization.

Required Qualifications

  • Strong hands-onPython programming and automation experience.
  • Experience testingAI/LLM-based, conversational AI, or agentic systems.
  • Experience with or strong knowledge ofLLM evaluation techniques, including LLM-as-judge, state/behavior validation, or similar approaches.
  • Experience validatingnon-deterministic AI behavior where expected outcomes cannot always be represented through conventional scripted assertions.
  • Ability todesign and build testing or evaluation frameworks, rather than solely executing within an established automation framework.
  • Strong understanding of software testing, automation principles, debugging, and root-cause analysis.
  • Ability to design test scenarios for complex, multi-step application and agent workflows.
  • Strong problem-solving skills and ability to operate effectively in anearly-stage, greenfield environment where standards and processes are still being defined.
  • Comfortable working in adistributed quality model, with QA responsibility shared closely across engineering rather than owned solely by a formal QA function.

Preferred Qualifications

  • Experience withPlaywright for front-end automation.
  • Experience withPyTest, Behave, or similar Python testing frameworks.
  • Experience testingAI agents, tool/function calling, multi-turn conversations, or complex agentic workflows.
  • Experience evaluating LLM outputs forgroundedness, hallucination, relevance, correctness, consistency, or completeness.
  • Experience testingguardrails, failure scenarios, and AI validation controls.
  • Experience validating agent tool usage and downstream system interactions.
  • API and database testing experience.
  • Experience integrating automated testing and evaluation intoCI/CD pipelines.
  • Experience withinHCM, HR technology, payroll, or related enterprise software.

Team & Environment

This is anearly-stage, greenfield AI initiative where engineering conventions, testing standards, and evaluation frameworks are actively being developed.

The team operates using anAI-SDLC model, with small engineering pods working closely with AI agents rather than relying on traditional, highly structured development and QA workflows. Quality ownership is collaborative and distributed across the team.

The broader product is more than a chatbot. It is an expandingagentic HCM software ecosystem, with a conversational interface serving as one entry point into a growing set of AI-driven HR and workforce capabilities.

Ideal Candidate

The strongest candidate will be aPython-focused QA Automation Engineer with genuine hands-on AI/LLM testing experience who is excited about helping define how agentic applications should be tested.

You should understand that testing an AI agent requires more than confirming whether a predefined workflow completes successfully. You will help determinewhether an agent behaved appropriately, reached the correct state, used the right tools, produced an acceptable result, and performed reliably across inherently variable interactions.

This role is particularly well suited for someone who enjoysbuilding frameworks rather than simply using them, solving testing problems that do not yet have established answers, and working closely with engineers to shape how agentic AI applications are evaluated and validated at scale.

QA / Automation Engineer— Agentic AI · ATech Placement

Auto apply with Likeremote