TT
AI Ops Lead / AIOps Technical Delivery Lead
Talent Technical Services, Inc
🇺🇸 United States
On-site
Staff / Principal
4 weeks ago
$41.66 – $45.45 / hour
- AI
- AIOps
- AI/ML
- GCP
- Ansible
- Terraform
- Large Language Models
- RAG
- Model Context Protocol
- MCP
- Vertex AI
- Dynatrace
- Splunk
- CloudWatch
- AWS
- Devops
- IaC
- PowerShell
- CI/CD
- ETL
- Snowflake
- Integration Testing
- Regulatory Compliance
- Vulnerability Management
4 weeks ago
Job Description – AI Ops Lead
- Job Title: AI Ops Lead / AIOps Technical Delivery Lead
- Location: Louisville, KY
- Duration: 6 months
- Experience Required: 8–10 years
- Primary Skills: AIOps, AI/ML, GCP, Automation, Ansible, Terraform
- Role Focus: Agentic AI, AIOps, Observability, Automation, Cloud Infrastructure
Education
- Master’s degree preferred.
- Bachelor’s degree required in:
- Computer Science
- Information Technology
- Systems Engineering
- Related technical field
Required Experience
- 8+ years of experience leading complex:
- Technical projects
- Product delivery
- Cloud infrastructure rollouts
- Experience working within a global enterprise IT environment.
- 2+ years of experience leading AI/AIOps platform deployments.
- Strong experience with enterprise-scale automation and cloud initiatives.
- Experience working with global delivery models and cross-functional technical teams.
AI & Agentic AI Skills
- Deep understanding of AI concepts, including:
- Large Language Models (LLMs)
- Agentic AI
- Multi-agent reasoning
- Tool-calling architectures
- Retrieval-Augmented Generation (RAG)
- Model Context Protocol (MCP)
- Experience designing or implementing Agentic AI architectures.
- Experience deploying LLM-based solutions in production environments.
- Experience with Google Cloud Vertex AI.
- Experience with agentic AI frameworks and dynamic agent registries.
- Strong understanding of AI tool-calling interfaces and integrations.
- Ability to translate complex enterprise environments into contextual schemas for AI consumption.
Cloud & Observability Skills
- Hands-on experience with enterprise observability and monitoring platforms such as:
- Dynatrace
- Splunk
- CloudWatch
- Strong experience with cloud environments, particularly:
- Google Cloud Platform (GCP)
- AWS
- Experience with hybrid and multi-cloud environments.
- Knowledge of GCVE and hybrid cloud environments.
- Experience with telemetry ingestion and monitoring standardization.
- Experience implementing automated alerting and self-healing workflows.
DevOps / SRE / Automation
- Strong understanding of Site Reliability Engineering (SRE) principles.
- Experience with automated testing.
- Infrastructure as Code (IaC) experience.
- Strong experience with:
- Ansible
- Terraform
- PowerShell
- Shell scripting
- Experience with CI/CD pipeline automation.
- Experience overseeing data pipelines and ETL validation.
- Familiarity with Snowflake and enterprise data validation processes.
- Experience developing automated remediation and self-healing solutions.
Key Responsibilities
- Serve as the primary technical delivery lead for enterprise-wide AIOps and automation initiatives.
- Orchestrate Agentic AI and AIOps platform deployments.
- Drive milestone tracking, code integration, testing, and production deployment.
- Coordinate deployment of:
- Multi-agent reasoning systems
- Dynamic agent registries
- Tool-calling interfaces
- Agentic AI frameworks
- Utilize Google Cloud Vertex AI, MCP, and related agentic technologies.
- Lead enterprise observability and self-healing automation programs.
- Oversee telemetry ingestion and monitoring standardization across multi-cloud infrastructure.
- Drive automated alerting and self-healing workflows.
- Develop and deploy vulnerability tracking workflows.
- Establish automated security remediation frameworks across application and infrastructure portfolios.
- Ensure compliance with enterprise security and governance standards.
- Coordinate engineering, application support, data, and security teams.
- Manage resource dependencies and competing priorities.
- Drive accountability across multiple technical teams.
- Track project milestones, delivery health, and operational improvements.
- Establish KPIs covering:
- Delivery health
- Token consumption
- Operational MTTR improvements
- Financial savings
- Build executive reporting frameworks to measure the business value and cost savings of AI and automation initiatives.
- Provide concise status updates to executive leadership.
- Anticipate technical bottlenecks and independently drive resolution.
- Translate complex technical information into clear executive-level communication.
Governance & Security
- Establish appropriate guardrails for:
- Data security
- AI safety
- Regulatory compliance
- Enterprise governance
- Ensure AI and automation solutions follow organizational security standards.
- Develop governance practices for enterprise AI/AIOps deployments.
- Support vulnerability management and automated remediation initiatives.
Leadership & Managerial Skills
- Ability to work directly with clients in an onsite environment.
- Strong client engagement and stakeholder management skills.
- Proven technical leadership abilities.
- Experience managing global delivery teams.
- Strong decision-making and problem-solving skills.
- Excellent communication and executive presence.
- Ability to manage complex, high-visibility programs.
- Strong thought leadership capabilities.
- Experience with reporting, governance, and executive communications.
- Ability to operate with a high degree of autonomy.
Key Resume Search Keywords
- GCP
- AIOps
- AI/ML
- Agentic AI
- LLM
- RAG
- MCP
- Automation
- GCVE
- Hybrid Cloud
- AWS
- Vertex AI
- Dynatrace
- Splunk
- CloudWatch
- Ansible
- Terraform
- PowerShell
- Shell Scripting
- SRE
- DevOps
- Observability
- Self-Healing Automation
Pre-Screening Questions
- Explain an Agentic AI architecture you have designed or implemented. What problem did it solve?
- What is the difference between traditional AI workflows and Agentic AI systems?
- How have you used LLMs in production environments?
- What is RAG (Retrieval-Augmented Generation), and where have you applied it?
- What experience do you have with MCP (Model Context Protocol) or tool-calling architectures?
- Describe your experience with Google Cloud Vertex AI.
- How have you implemented AIOps or automated self-healing solutions?
- Describe your experience with enterprise observability platforms such as Dynatrace, Splunk, or CloudWatch.
- How have you used Ansible and Terraform in enterprise automation?
- Describe an automation initiative where you achieved measurable operational or financial savings.
Role Classification
- Role Description: AI Ops Lead
- Essential Skills: AI Ops Lead
- Desirable Skills: Ansible, Terraform
- Keyword: AIOps / AI Automation
- Skills: Digital – Ansible | Digital – Terraform
- Experience Required: 8–10 years
AI Ops Lead / AIOps Technical Delivery Lead · Talent Technical Services, Inc