Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Engineer, Storage and Data Protection

AHEAD
🇮🇳 India
Hybrid
3 weeks ago
  • Incident Management
  • Change Management
  • Ceph
  • AI
  • Linux
  • Slurm
  • Kubernetes
  • InfiniBand
  • Disaster Recovery
  • Machine Learning
  • AI/ML
  • MinIO
  • Terraform
  • Ansible
  • Helm
  • GitOps
  • Prometheus
  • Grafana
  • Python
  • Bash
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
The High-Performance Computing Storage Engineer is primarily responsible for the overall health and maintenance of storage technologies in our managed services customer's environments. Our Storage Engineers are a valued member of the Managed Services Infrastructure Practice responsible for Tier 3 incident management, service request management and change management infrastructure support for all Managed Services customers.  

Key Responsibilities

     

    • Provide enterprise-level operational support to Managed Services customers for incident, problem, and change management activities 

    • Administer parallel and distributed filesystems such as Lustre, GPFS, BeeGFS, Ceph, Weka, or Vast 

    • Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference 

    • Build and maintain automation for storage provisioning, monitoring, alerting, quota management, and lifecycle operations 

    • Plan and perform maintenance activities 

    • Assess customer environments for performance and design issues and propose resolutions 

    • Work across technical teams to troubleshoot complex infrastructure issues 

    • Create and maintain detailed documentation 

    • Serve as a subject matter expert and escalation point for storage technologies 

    • Work with vendors to resolve storage issues 

    • Communicate with customers and internal team with transparency 

    • Support data movement workflows including ingest, replication, caching, tiering, and archiving 

    • Troubleshoot storage, Linux, network, and I/O bottlenecks across storage clusters and fabrics 

    • Partner with infrastructure, platform, and research teams to support production AI/HPC workloads 

    • Evaluate new storage architectures and technologies for scalability, resilience, and cost efficiency 

    • Communicate with customers and internal team with transparency 

    • Participate in on-call rotation 

Required Qualifications

     

    • 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering 

    • Bachelor’s degree or equivalent Information Systems or related field. Unique education, specialized experience, skills, knowledge, training, or certification may be substituted for education 

    • Strong experience with Linux systems administration 

    • Hands-on experience configuring, managing, and tuning distributed or parallel filesystems 

    • Experience tuning storage for performance-sensitive workloads 

    • Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes 

    • Familiarity with high-speed interconnects such as InfiniBand or RDMA 

    • Ability to troubleshoot complex issues across storage, compute, and networking layers 

    • Understanding of data protection mechanisms, including data replication, backup strategies, and disaster recovery in HPC environments 

    • Experience with machine learning or data science workflows in HPC environments 

    • Managed Services or consulting experience 

    • Strong background with customer service 

    • High level problem-solving and communication skills 

    • Strong oral and written communications skills 

    • Managed Services or consulting experience 

  •  

     

Preferred Qualifications 

     

    • Experience supporting storage solutions for GPU clusters and AI/ML workflows 

    • Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway 

    • Experience with Terraform, Ansible, Helm, or GitOps workflows 

    • Knowledge of observability platforms such as Prometheus and Grafana 

    • Experience with multi-petabyte environments, caching architectures, and storage isolation in multi-tenant systems 

    • Experience with machine learning or data science workflows in HPC environments 

    • Scripting or programming experience with Python and Bash 

    • Related Storage certifications are a bonus 

Engineer, Storage and Data Protection · AHEAD

Auto apply with Likeremote