M
Senior Infrastructure Engineer
Mphasis
πΊπΈ United States
On-site
Senior
1 week ago
- AI
- Linux
- Python
- Bash
- InfiniBand
- Test Automation
- CI/CD
- Kubernetes
- Slurm
1 week ago
Contractor 2 β AI Cluster Validation Engineer (Tester & Programmer)
Key Responsibilities
- AI Cluster Infrastructure Validation: Build, deploy, and maintain hardware and software test environments for AI cluster validation, qualification, and performance benchmarking.
- Execute test plans, reproduce complex issues, collect and analyze logs, and perform first-level root cause analysis across hardware, software, networking, and system components.
- Develop and enhance automated test frameworks, scripts, and tools to improve validation efficiency, coverage, and repeatability.
- Collaborate with software, hardware, networking, and system engineering teams to investigate issues, validate fixes, and improve overall cluster stability, scalability, and performance.
- Conduct functional, performance, stress, and reliability testing for AI/HPC cluster solutions.
- Document test methodologies, configurations, results, troubleshooting procedures, and operational best practices.
Basic Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- 1+ years of experience in deploying, testing, validating, or supporting data center hardware and software systems, with expertise in one or more of the following areas:
Server systems, Networking, Storage
- Strong understanding of Linux operating systems, system administration, and troubleshooting.
- Familiarity with data center cluster management platforms, distributed computing environments, and infrastructure validation methodologies.
- Programming or scripting experience with Python, Bash, or similar languages.
- Strong analytical, debugging, problem-solving, and troubleshooting skills.
- Proven ability to quickly learn new technologies and adapt in a fast-paced engineering environment.
- Self-motivated team player with strong communication and collaboration skills.
Preferred Qualifications
- Experience operating, maintaining, or supporting data center, cloud, or laboratory environments.
- Familiarity with AI/HPC clusters, GPU-based systems, and high-speed interconnect technologies such as InfiniBand, RoCE, NVLink, and Ethernet fabrics.
- Experience developing test automation tools, validation frameworks, or CI/CD pipelines.
- Knowledge of cluster orchestration and management technologies, including Kubernetes, Slurm, virtualization platforms, or cloud infrastructure.
- Experience with performance analysis, benchmarking, workload characterization, and system optimization.
- Understanding of storage technologies, distributed file systems, and AI workload deployment environments.
Senior Infrastructure Engineer Β· Mphasis