Senior HPC Support Engineer & Deployment Lead
- InfiniBand
- Kubernetes
- C#
- CUDA
- PCIe
- Python
- .NET
Your Mission
You will act as the escalation point for our most challenging technical hurdles, ensuring that our compiler technology runs flawlessly on the world's most powerful hardware.
Think: Lead the architectural strategy for customer rollouts. You will analyze client infrastructure—evaluating power, high-speed interconnects (Infiniband/RoCE), and software environments—to plan successful cluster deployments. You will drive complex customer issues to resolution by diagnosing root causes that sit between hardware, the OS, and our application layer.
Implement:
Execute hands-on deployments of Kubernetes clusters (on-prem and cloud) tailored for GPU acceleration.
Dive deep into code and systems to detail, reproduce, and resolve issues. You will set up test environments using C#, CUDA, and ROCm to mimic customer failures.
Work directly with the latest silicon (NVIDIA H100, AMD MI300) and interconnects to ensure our software utilizes the hardware correctly.
Build:
The Knowledge Base: You will author detailed technical solutions, white papers, and "known issue" documentations. Your work will empower the rest of the team and our users to solve problems faster.
Feedback Loops: Collaborate closely with the Engineering and R&D teams. You will translate field data into clear bug reports and feature requests, helping to shape the future stability of the product.
What You Bring to the Table
You are a "System Doctor." You have the computer science fundamentals to understand code, but your expertise lies in making that code run reliably on physical systems.
Experience: You have a BS/MS in Computer Science, Electrical Engineering, or related field, with8+ years of experience in system software development and hardware support. You have a proven track record in customer-facing roles.
HPC & Hardware Fluency: You have a deep understanding of GPU architectures and how they interact with the rest of the system. You are comfortable dealing with high-speed interconnects, PCIe topology, and driver stacks.
Software Ecosystem: You possess strong computer science fundamentals. You are an expert inPython and scripting for automation, but you are also comfortable navigatingC#/.NET environments and theCUDA/ROCm ecosystems.
Containerization: You have practical experience deploying and debuggingKubernetes clusters in production environments.
Communication: Excellent interpersonal skills are non-negotiable. You can remain calm under pressure, communicate complex technical details to stakeholders, and manage customer expectations effectively.
Senior HPC Support Engineer & Deployment Lead · hybridizer