Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TC

Linux Admin

TechDigital Corporation
Location not stated
Remote
11 months ago
  • AI
  • Linux
  • CUDA
  • AI/ML
  • CI/CD
  • GitHub Actions
  • Prometheus
  • Grafana
  • Network Security
  • VLANs
  • SOC2
  • ISO 27001
  • Devops
  • AWS
  • Azure
  • GCP
  • VPC
  • IAM
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Key Responsibilities
● Infrastructure Management: Provision, deploy, and maintain scalable, secure, and high-availability cloud infrastructure on platforms such as Cloud to support
AI workloads.
● System Management: Administer and maintain Linux-based servers and clusters optimized for GPU compute workloads, ensuring high availability and performance.
● GPU Infrastructure: Configure, monitor, and troubleshoot GPU hardware (e.g., NVIDIA GPUs) and related software stacks (e.g., CUDA, cuDNN) for optimal performance in AI/ML and HPC applications.
● Troubleshooting: Diagnose and resolve hardware and software issues related to GPU compute nodes and performance issues in GPU clusters.
● High-Speed Interconnects: Implement and manage high-speed networking technologies like RDMA over Converged Ethernet (RoCE) to support low-latency, high-bandwidth communication for GPU workloads.
● CI/CD Pipelines: Build and optimize continuous integration and deployment (CI/CD) pipelines for testing GPU-based servers and managing deployments using tools like GitHub Actions.
● Monitoring & Performance: Set up and maintain monitoring, logging, and alerting systems (e.g., Prometheus, Victoria Metrics, Grafana) to track system performance, GPU utilization, resource bottlenecks, and uptime of GPU resources.
● Security and Compliance: Implement network security measures, including firewalls, VLANs, VPNs, and intrusion detection systems, to protect the GPU compute environment and comply with standards like SOC 2 or ISO 27001.
Required Qualifications
● Experience: 3+ years of experience in DevOps, Site Reliability Engineering (SRE), or cloud infrastructure management, with at least 1 year working on GPU-basedcomputer environments in the cloud.
● Linux Administration: Strong knowledge of Linux system administration for managing network services and tools in a GPU compute environment.
● High-Speed Interconnects: Experience with high-performance networking technologies like RoCE, or 100GbE Ethernet in compute-intensive environments.
● GPU-Specific Networking: Proficiency with NVIDIA GPU networking technologies, such as Mellanox ConnectX adapters, and configuring Netplan to support their drivers and firmware.
● Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, GCP).
● Networking & Security: Knowledge of networking concepts (VPC, subnets) and security best practices (IAM, encryption, firewall configurations).

Linux Admin · TechDigital Corporation

Auto apply with Likeremote