Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
LI

AI/HPC System Engineer

Leadstack Inc
Location not stated
1 month ago
  • AI
  • IaC
  • LLM APIs
  • Linux
  • AWS
  • Azure
  • GCP
  • Kubernetes
  • Slurm
  • AI/ML
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
LeadStack Inc. is an award-winning, one of the nation’s fastest-growing, certified minority-owned (MBE) staffing services provider of contingent workforce. As a recognized industry leader in contingent workforce solutions and Certified as a Great Place to Work, we’re proud to partner with some of the most admired Fortune 500 brands in the world.

Description:
We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development workloads. This role deploys, automates, and maintains GPU clusters across on-premise and cloud environments, delivering reliable, scalable, and cost-efficient compute for engineering and R&D teams.

Responsibilities:

• GPU/HPC infrastructure: Build, configure, and operate GPU and HPC clusters across compute, storage, and networking; support capacity planning, performance tuning, and optimization for AI training, inference, and compute-intensive workloads
• Hybrid cloud infrastructure: Deploy and maintain compute environments spanning on-premise and public cloud, and contribute to modernization and scaling initiatives for HPC/AI infrastructure
• Automation and observability: Implement infrastructure-as-code, provisioning automation, monitoring, and alerting, and drive improvements in resource utilization and efficiency
• AI platform support: Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by internal engineering teams
• Operations and collaboration: Troubleshoot and resolve infrastructure issues, document standards and runbooks, and work with relevant stakeholders to support day-to-day IT operations

Requirements:
Qualifications: • Bachelor's degree in Computer Science, Engineering, or a related technical field • 3+ years of hands-on experience in IT infrastructure, cloud, platform engineering, or HPC • Hands-on experience with Linux-based infrastructure and public cloud environments such as AWS, Azure, or GCP • Experience deploying or operating GPU/HPC environments, including workload scheduling or orchestration platforms such as Kubernetes or Slurm • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization • Solid understanding of compute, storage, networking, and container technologies; experience with AI/ML infrastructure or workloads is a plus • Strong collaboration and communication skills, with the ability to work across engineering and IT teams

To know more about current opportunities at LeadStack, please visit us at https://leadstackinc.com/careers/
Should you have any questions, feel free to call me on or send an email on _____________________

AI/HPC System Engineer · Leadstack Inc

Auto apply with Likeremote