A
AI Platform Engineer
Apolis
- πΊπΈ United States
- On-site
- Senior
- 10 hours ago
- $60 β $62 / hour
- AI
- AI/ML
- Machine Learning
- CI/CD
- MLOps
- GCP
- GitHub Actions
- Jenkins
- GitLab CI
- Vertex AI
- Kubeflow
- MLflow
- IaC
- Terraform
- Docker
- Kubernetes
- GKE
- Prometheus
- Grafana
- IAM
- Secrets Management
- Devops
- Cloud Run
- VPC
- BigQuery
- ArgoCD
- Python
- Bash
- Gemini
- GitOps
- FinOps
10 hours ago
AI Platform Engineer
Location: Charlotte, NC β Onsite
Experience: 7+ Years
About the Role
We are seeking a seniorAI Platform Engineer to build, operate, and scale the infrastructure and tooling that supportAI/ML and Generative AI workloads. This role will focus on cloud platform engineering, CI/CD automation, MLOps, infrastructure automation, security, and observability to enable reliable and scalable delivery of AI solutions.
Key Responsibilities
- Design, build, and maintain scalableGCP cloud infrastructure supporting AI/ML and application workloads.
- Architect and manageCI/CD pipelines for model training, deployment, and application releases using Cloud Build, GitHub Actions, Jenkins, GitLab CI, or similar tools.
- Build and maintainMLOps pipelines for model versioning, training, deployment, and monitoring using Vertex AI Pipelines, Kubeflow, and MLflow.
- ImplementInfrastructure as Code (IaC) using Terraform or similar tools for consistent and auditable environment provisioning.
- Containerize and orchestrate applications and AI services usingDocker and Kubernetes/GKE.
- Establish monitoring, logging, alerting, and observability for AI/ML services usingCloud Monitoring, Cloud Logging, Prometheus, and Grafana.
- Partner with AI/ML engineers and data scientists to productionize models and improve the transition from experimentation to production.
- Implement and maintaincloud security, IAM, secrets management, and compliance across environments.
- Automate build, testing, deployment, rollback, and operational processes to reduce manual effort.
- Troubleshoot infrastructure and platform issues and leadroot-cause analysis and resolution for production incidents.
Required Qualifications
- 7+ years of experience in Platform Engineering, DevOps, Cloud Infrastructure, or related roles.
- Strong hands-on experience withGoogle Cloud Platform (GCP), including Vertex AI, GKE, Cloud Run, IAM, VPC/networking, and BigQuery.
- Strong experience designing and managingCI/CD pipelines using Cloud Build, Jenkins, GitHub Actions, GitLab CI, ArgoCD, or similar technologies.
- Strong proficiency inTerraform and Infrastructure as Code practices.
- Strong scripting or programming skills usingPython, Bash, or Go.
- Hands-on experience withDocker and Kubernetes.
- Solid understanding ofMLOps and the machine learning lifecycle, including training, model versioning, deployment, and monitoring.
- Experience withobservability and monitoring tools such as Cloud Monitoring, Prometheus, Grafana, and ELK/EFK.
- Strong understanding ofcloud security, IAM, secrets management, and enterprise security best practices.
Preferred Qualifications
- GCP Professional certification such asProfessional Cloud DevOps Engineer, Professional Cloud Architect, or Professional Machine Learning Engineer.
- Experience supportingGenerative AI and LLM platforms, including Vertex AI, Gemini Enterprise, and Model Garden.
- Experience implementingGitOps workflows using ArgoCD or Flux.
- Experience working inregulated, highly secure, or enterprise-scale environments.
- Knowledge ofGCP cost optimization and FinOps practices
AI Platform Engineer Β· Apolis