Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
M

Senior Infrastructure Engineer

Mphasis
πŸ‡ΊπŸ‡Έ United States
On-site
Senior
1 week ago
  • AI
  • Linux
  • Python
  • Bash
  • InfiniBand
  • Test Automation
  • CI/CD
  • Kubernetes
  • Slurm
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Contractor 2 – AI Cluster Validation Engineer (Tester & Programmer)

Key Responsibilities

  • AI Cluster Infrastructure Validation: Build, deploy, and maintain hardware and software test environments for AI cluster validation, qualification, and performance benchmarking.
  • Execute test plans, reproduce complex issues, collect and analyze logs, and perform first-level root cause analysis across hardware, software, networking, and system components.
  • Develop and enhance automated test frameworks, scripts, and tools to improve validation efficiency, coverage, and repeatability.
  • Collaborate with software, hardware, networking, and system engineering teams to investigate issues, validate fixes, and improve overall cluster stability, scalability, and performance.
  • Conduct functional, performance, stress, and reliability testing for AI/HPC cluster solutions.
  • Document test methodologies, configurations, results, troubleshooting procedures, and operational best practices.

Basic Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • 1+ years of experience in deploying, testing, validating, or supporting data center hardware and software systems, with expertise in one or more of the following areas:

Server systems, Networking, Storage

  • Strong understanding of Linux operating systems, system administration, and troubleshooting.
  • Familiarity with data center cluster management platforms, distributed computing environments, and infrastructure validation methodologies.
  • Programming or scripting experience with Python, Bash, or similar languages.
  • Strong analytical, debugging, problem-solving, and troubleshooting skills.
  • Proven ability to quickly learn new technologies and adapt in a fast-paced engineering environment.
  • Self-motivated team player with strong communication and collaboration skills.

Preferred Qualifications

  • Experience operating, maintaining, or supporting data center, cloud, or laboratory environments.
  • Familiarity with AI/HPC clusters, GPU-based systems, and high-speed interconnect technologies such as InfiniBand, RoCE, NVLink, and Ethernet fabrics.
  • Experience developing test automation tools, validation frameworks, or CI/CD pipelines.
  • Knowledge of cluster orchestration and management technologies, including Kubernetes, Slurm, virtualization platforms, or cloud infrastructure.
  • Experience with performance analysis, benchmarking, workload characterization, and system optimization.
  • Understanding of storage technologies, distributed file systems, and AI workload deployment environments.

Senior Infrastructure Engineer Β· Mphasis

Auto apply with Likeremote