Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

DevOps & Infrastructure Engineer (HPC/GPU)

hybridizer
🇺🇸 United States | 🇫🇷 France
Hybrid
10 months ago
  • Devops
  • CI/CD
  • Windows
  • Linux
  • Kubernetes
  • GitHub Actions
  • PCIe
  • CUDA
  • Docker
  • Helm
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Your Mission

You will own the infrastructure that powers our R&D and helps our customers deploy our technology on-premise. You will move beyond standard cloud DevOps into the world of High-Performance Computing (HPC).

  • Think: Design a robust CI/CD strategy that handles cross-platform compilation (Windows/Linux) and execution on specific hardware targets (NVIDIA A100, AMD MI250, Consumer GPUs). Architect solution templates for our customers who need to deploy Hybridizer-generated binaries on their own private clouds.

  • Implement:

    • Set up and maintainKubernetes clusters (both on-premise and cloud) withGPU Passthrough and Multi-Instance GPU (MIG) configurations.

    • DevelopGitHub Actions pipelines that seamlessly dispatch heavy test suites to self-hosted runners equipped with specific GPU accelerators.

    • ConfigureDockerHub registries and secure container lifecycles for our compiler images.

  • Build:

    • Hardware Tuning: Assemble and fine-tune physical servers. This includes managing PCIe topology, cooling profiles, and power constraints to ensure consistent benchmarking results.

    • Driver Ecosystem: Manage the complex matrix of NVIDIA drivers, CUDA toolkits, and ROCm versions across our fleet, ensuring compatibility with our compiler’s output.


What You Bring to the Table

You are a DevOps engineer who loves hardware. You understand that "the cloud" is just someone else's computer, and sometimes you need to manage that computer yourself.

  • Core DevOps: Strong mastery ofDocker andKubernetes. You know how to write custom Helm charts and manage stateful sets.

  • GPU Infrastructure: You have hands-on experience withNVIDIA Container Toolkit orROCm integration in containers. You understand concepts like PCIe passthrough, IOMMU groups, and GPU orchestration.

  • CI/CD Automation: Expert inGitHub Actions. You can write complex workflows with matrix strategies and self-hosted runners.

  • System Administration: You are comfortable with Linux kernel tuning, driver installation (dkms), and diagnosing hardware bottlenecks.

  • Customer Facing: You have the communication skills to assist clients. You can explain how to expose a GPU to a Docker container to a sysadmin who might not be an expert in HPC.

  • Adaptability: You are ready to work with a mix of consumer and data-center grade hardware (e.g., configuring a server with 4x RTX 5090s or managing a DGX station).

DevOps & Infrastructure Engineer (HPC/GPU) · hybridizer

Auto apply with Likeremote