Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
MD

Senior SRE & Infra Engineer (GPU Cluster Platform Reliability & Infrastructure Engineer)

Macpower Digital Assets Edge Private Limited
πŸ‡ΊπŸ‡Έ United States
Hybrid
Senior
5 days ago
$100,000 – $200,000 / year
  • Prometheus
  • Grafana
  • Incident Management
  • Ansible
  • Terraform
  • Python
  • REST API
  • Kubernetes
  • GCP
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
This hybrid role spans across platform reliability and infrastructure engineering. You'll be instrumental in ensuring high availability, fault tolerance, and performance across internal research and external customers' GPU cluster environments. Responsibilities include automating GPU cluster onboarding, enhancing monitoring, logging, and security systems, and developing new backend features.
Required Skills and Certifications:
  • Proven experience with monitoring tools (e.g., Prometheus, Grafana) and incident management practice.
  • Strong skills in infrastructure automation with Ansible, Terraform, or similar.
  • Deep understanding of logging frameworks, alerting systems, and proactive monitoring solutions.
  • Proficiency in Python for developing automation scripts, REST APIs, and backend support tools.
  • Hands-on experience with Kubernetes and cloud platforms (GCP preferred).
  • Knowledge of high-performance networking and real-time systems.

Senior SRE & Infra Engineer (GPU Cluster Platform Reliability & Infrastructure Engineer) Β· Macpower Digital Assets Edge Private Limited

Auto apply with Likeremote