Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
ES

Senior CORA HPC DevOps Engineer

EPAM Systems
🇱🇻 Latvia | 🇱🇹 Lithuania
Remote
Senior
21 hours ago
  • Devops
  • Machine Learning
  • MLOps
  • AWS
  • GxP
  • CI/CD
  • CloudWatch
  • Prometheus
  • Node.js
  • Docker
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

We are seeking aSenior CORA HPC DevOps Engineer to drive the scaling, reliability, and automation of HPC CORA — the designated high-performance computing (HPC) and machine learning operations (MLOps) platform for Science, Innovation & Labs.

EPAM Collaborative Omics Research Acceleratorâ„¢ (CORAâ„¢) for AWS HealthOmics provides life science and healthcare organizations with a user-friendly, secure, and GxP-ready environment for analyzing and correlating omics data, leveraging AWS HealthOmics features such as multiomic and multimodal analysis, population sequencing, and fully managed bioinformatics computation. Designed for bioinformaticians, lab and bench scientists, and clinical providers, EPAM CORA delivers a collaborative and extensible environment that democratizes access to high-performance omics computations.

Responsibilities

  • Guide scientists and data teams in navigating and utilising the CORA user interface (UI) effectively, enabling them to run self-service workloads without direct infrastructure friction
  • Advise users and manage infrastructure capacity across capacity blocks (reserved GPU/CPU capacity) versus on-demand usage, optimising cost, quotas, and resource availability for heavy workloads
  • Maintain automated CI/CD pipelines for infrastructure provisioning and platform service deployments
  • Provide Tier-2/3 support and troubleshooting for technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas
  • Collaborate with developer experience teams to improve platform documentation
  • Partner with engineering teams to monitor GPU utilisation (e.g., via CloudWatch/Prometheus) to minimise idle time, right-size clusters, and optimise multi-node job scheduling

Requirements

  • 3+ years of experience in HPC DevOps or similar engineering roles
  • Practical knowledge of MPI, OpenMP, and multi-node GPU communication protocols (NCCL, GPUDirect)
  • Proven background in managing AWS GPU instance families (P-series, G-series, Trainium/Inferentia) and allocating block compute for large-scale ML training and inference pipelines
  • Hands-on expertise in capacity planning and reservation using AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management to ensure compute availability
  • Skills in deploying containerised environments tuned for HPC and GPU pass-through using Apptainer/Singularity, Docker, or Enroot
  • Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, or BeeGFS
  • Proficiency in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
  • Experience deploying or scaling HPC workloads on cloud infrastructure utilising EFA, ParallelCluster, and parallel storage (FSx for Lustre)
  • B1+ English level proficiency

Nice to have

  • Familiarity with AWS Compute services
  • Background in building and maintaining CI/CD workflows

Senior CORA HPC DevOps Engineer · EPAM Systems

Auto apply with Likeremote