DevOps & Infrastructure Engineer (HPC/GPU)
- Devops
- CI/CD
- Windows
- Linux
- Kubernetes
- GitHub Actions
- PCIe
- CUDA
- Docker
- Helm
Your Mission
You will own the infrastructure that powers our R&D and helps our customers deploy our technology on-premise. You will move beyond standard cloud DevOps into the world of High-Performance Computing (HPC).
Think: Design a robust CI/CD strategy that handles cross-platform compilation (Windows/Linux) and execution on specific hardware targets (NVIDIA A100, AMD MI250, Consumer GPUs). Architect solution templates for our customers who need to deploy Hybridizer-generated binaries on their own private clouds.
Implement:
Set up and maintainKubernetes clusters (both on-premise and cloud) withGPU Passthrough and Multi-Instance GPU (MIG) configurations.
DevelopGitHub Actions pipelines that seamlessly dispatch heavy test suites to self-hosted runners equipped with specific GPU accelerators.
ConfigureDockerHub registries and secure container lifecycles for our compiler images.
Build:
Hardware Tuning: Assemble and fine-tune physical servers. This includes managing PCIe topology, cooling profiles, and power constraints to ensure consistent benchmarking results.
Driver Ecosystem: Manage the complex matrix of NVIDIA drivers, CUDA toolkits, and ROCm versions across our fleet, ensuring compatibility with our compiler’s output.
What You Bring to the Table
You are a DevOps engineer who loves hardware. You understand that "the cloud" is just someone else's computer, and sometimes you need to manage that computer yourself.
Core DevOps: Strong mastery ofDocker andKubernetes. You know how to write custom Helm charts and manage stateful sets.
GPU Infrastructure: You have hands-on experience withNVIDIA Container Toolkit orROCm integration in containers. You understand concepts like PCIe passthrough, IOMMU groups, and GPU orchestration.
CI/CD Automation: Expert inGitHub Actions. You can write complex workflows with matrix strategies and self-hosted runners.
System Administration: You are comfortable with Linux kernel tuning, driver installation (dkms), and diagnosing hardware bottlenecks.
Customer Facing: You have the communication skills to assist clients. You can explain how to expose a GPU to a Docker container to a sysadmin who might not be an expert in HPC.
Adaptability: You are ready to work with a mix of consumer and data-center grade hardware (e.g., configuring a server with 4x RTX 5090s or managing a DGX station).
DevOps & Infrastructure Engineer (HPC/GPU) · hybridizer