Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Lead HPC Kubernetes Engineer

EPAM Systems
πŸ‡ΊπŸ‡¦ Ukraine
Remote
Staff / Principal
15 hours ago
  • Kubernetes
  • AWS
  • GCP
  • EKS
  • GKE
  • Node.js
  • CI/CD
  • OCI
  • AI/ML
  • Terraform
  • IaC
  • EC2
  • Python
  • Devops
  • Elastic
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are looking for aLead HPC Kubernetes Engineer to help our customer develop and manage several HPC clusters across AWS, CoreWeave, GCP, and other providers, supporting several thousand GPUs today and scaling to 10x in 2026 and beyond. This role is Kubernetes-heavy, requiring you to operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine.

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers
  • Take responsibility for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth
  • Provision HPC infrastructure through a CI/CD system across AWS, CoreWeave, GCP, and OCI, with additional providers to be expanded in the near future
  • Manage job scheduling to allocate GPU compute across training and inference workloads
  • Define and maintain SLIs/SLOs
  • Build monitoring and alerting systems
  • Participate in severity escalation response and author post-incident reviews
  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams

Requirements

  • 5+ years in infrastructure engineering, cloud platforms, or HPC
  • Expertise in Kubernetes with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
  • Proficiency in Terraform for writing and reviewing infrastructure-as-code daily
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
  • Skills in Python for tooling and automation
  • Background in Site Reliability Engineering and DevOps practices
  • English proficiency at B2 level or higher

Nice to have

  • Familiarity with Google Kubernetes Engine and Google Cloud Platform
  • Knowledge of Amazon Elastic Kubernetes Service

Lead HPC Kubernetes Engineer Β· EPAM Systems

Auto apply with Likeremote