Senior DevOps Engineer, AIOps
- AI
- AIOps
- Devops
- CI/CD
- Kubernetes
- Helm
- Python
- FastAPI
- Node.js
- React.js
- Microservices
- GitLab CI/CD
- Artifactory
- PostgreSQL
- Temporal
- OpenTelemetry
- Datadog
- Grafana
- Secrets Management
- TLS
- RBAC
- Docker
- Linux
- Bash
- IaC
- Terraform
- Ansible
- SQL
- TCP/IP
- DNS
- Load Balancing
- Incident Response
- LangGraph
- Model Context Protocol
- MCP
- ClickHouse
- Redis
- Valkey
- Prometheus
- OpenShift
NVIDIA is powering the world’s most advanced AI factories, where resilient infrastructure is essential to keep accelerated computing environments running at scale. The Agentic AIOps team is building a mission-critical observability and prediction platform - delivered as both a high-scale SaaS solution and a robust on-premises deployment for NVIDIA’s largest enterprise customers.Â
As a Senior DevOps Engineer, you’ll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIA’s AI infrastructure.Â
What You'll Be Doing:Â
- Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations.Â
- Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints.Â
- Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory.Â
- Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity.Â
- Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations.Â
- Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows.Â
- Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening.Â
- Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes.Â
Â
What We Need to See:Â
- Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience.Â
- 5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices.Â
- Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting.Â
- Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible.Â
- Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices.Â
- Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting.Â
- Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting.Â
- Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment.Â
Â
Ways To Stand Out From the Crowd:Â
- Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost.Â
- Familiarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage.Â
- Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana.Â
- Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers.Â
- Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain.Â
Â
With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are passionate about building mission-critical systems at the frontier of AI infrastructure, we want to hear from you.Â
Â
#LI-Hybrid
Senior DevOps Engineer, AIOps · Nvidia