Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Senior HPC Support Engineer & Deployment Lead

hybridizer
🇺🇸 United States | 🇫🇷 France
Hybrid
Staff / Principal
10 months ago
  • InfiniBand
  • Kubernetes
  • C#
  • CUDA
  • PCIe
  • Python
  • .NET
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Your Mission

You will act as the escalation point for our most challenging technical hurdles, ensuring that our compiler technology runs flawlessly on the world's most powerful hardware.

  • Think: Lead the architectural strategy for customer rollouts. You will analyze client infrastructure—evaluating power, high-speed interconnects (Infiniband/RoCE), and software environments—to plan successful cluster deployments. You will drive complex customer issues to resolution by diagnosing root causes that sit between hardware, the OS, and our application layer.

  • Implement:

    • Execute hands-on deployments of Kubernetes clusters (on-prem and cloud) tailored for GPU acceleration.

    • Dive deep into code and systems to detail, reproduce, and resolve issues. You will set up test environments using C#, CUDA, and ROCm to mimic customer failures.

    • Work directly with the latest silicon (NVIDIA H100, AMD MI300) and interconnects to ensure our software utilizes the hardware correctly.

  • Build:

    • The Knowledge Base: You will author detailed technical solutions, white papers, and "known issue" documentations. Your work will empower the rest of the team and our users to solve problems faster.

    • Feedback Loops: Collaborate closely with the Engineering and R&D teams. You will translate field data into clear bug reports and feature requests, helping to shape the future stability of the product.


What You Bring to the Table

You are a "System Doctor." You have the computer science fundamentals to understand code, but your expertise lies in making that code run reliably on physical systems.

  • Experience: You have a BS/MS in Computer Science, Electrical Engineering, or related field, with8+ years of experience in system software development and hardware support. You have a proven track record in customer-facing roles.

  • HPC & Hardware Fluency: You have a deep understanding of GPU architectures and how they interact with the rest of the system. You are comfortable dealing with high-speed interconnects, PCIe topology, and driver stacks.

  • Software Ecosystem: You possess strong computer science fundamentals. You are an expert inPython and scripting for automation, but you are also comfortable navigatingC#/.NET environments and theCUDA/ROCm ecosystems.

  • Containerization: You have practical experience deploying and debuggingKubernetes clusters in production environments.

  • Communication: Excellent interpersonal skills are non-negotiable. You can remain calm under pressure, communicate complex technical details to stakeholders, and manage customer expectations effectively.

Senior HPC Support Engineer & Deployment Lead · hybridizer

Auto apply with Likeremote