Senior High-Performance Computing (HPC) Systems Engineer (On-Site)
- VMware
- HPE
Not enough detail in this posting to match
Senior High-Performance Computing (HPC) Systems Engineer (On-Site)
Shelton, CTPosition Overview
Our client is seeking a Senior High Performance Computing (HPC) Systems Engineer to provide on-site operational and technical support for a large-scale HPC environment. This individual will serve as a technical expert responsible for maintaining compute infrastructure, storage systems, cluster management, and overall system health while also supporting operational planning and project coordination.
Key Responsibilities
-
Provide on-site technical support for an HPC environment utilizing SGI 8600 computer systems.
-
Administer and support HPCM cluster management.
-
Support the DMF storage environment, including system monitoring, health checks, upgrades, and coordination with the DMF team.
-
Maintain and troubleshoot CDU cooling systems supporting the HPC infrastructure.
-
Perform hardware diagnostics and coordinate repairs to ensure maximum system availability.
-
Work closely with the site administrator to maintain operational stability and resolve infrastructure issues.
-
Facilitate weekly operational meetings and assist with project planning and coordination.
-
Provide technical leadership and act as the primary point of contact for infrastructure-related issues.
-
Support ongoing operational improvements and other infrastructure initiatives as needed.
Required Qualifications
-
Bachelor's degree in computer science, Information Technology, Engineering, or a related field (or equivalent experience).
-
8+ years of professional IT infrastructure experience (11+ years without a degree).
-
Strong experience supporting enterprise High-Performance Computing (HPC) environments.
-
Experience administering HPCM cluster management.
-
Experience supporting enterprise storage environments, preferably DMF.
-
Experience with enterprise hardware diagnostics and infrastructure support.
-
Experience maintaining mission-critical compute environments.
-
Strong troubleshooting and problem-solving skills.
-
Excellent communication and organizational skills.
Preferred Technical Experience
-
SGI 8600 Compute Systems
-
HPCM Cluster Management
-
DMF Storage Environment
-
VMware vSphere
-
VMware vMotion
-
VMware vCenter
-
HPE Integrated Lights-Out (iLO)
-
HPE Insight Control
-
Operations Orchestration
-
Emerson Aperture
-
Enterprise server and storage infrastructure
-
Data center cooling systems (CDU)
Preferred Skills
-
Ability to work independently in a production data center environment.
-
Experience coordinating infrastructure projects and operational activities.
-
Comfortable leading technical discussions and collaborating across multiple teams.
-
Strong customer-facing and consulting experience.
-
Experience supporting highly available, mission-critical systems.
Senior High-Performance Computing (HPC) Systems Engineer (On-Site) ยท TEC Group Inc.