MD
Senior SRE & Infra Engineer (GPU Cluster Platform Reliability & Infrastructure Engineer)
Macpower Digital Assets Edge Private Limited
πΊπΈ United States
Hybrid
Senior
5 days ago
$100,000 β $200,000 / year
- Prometheus
- Grafana
- Incident Management
- Ansible
- Terraform
- Python
- REST API
- Kubernetes
- GCP
5 days ago
Required Skills and Certifications:
- Proven experience with monitoring tools (e.g., Prometheus, Grafana) and incident management practice.
- Strong skills in infrastructure automation with Ansible, Terraform, or similar.
- Deep understanding of logging frameworks, alerting systems, and proactive monitoring solutions.
- Proficiency in Python for developing automation scripts, REST APIs, and backend support tools.
- Hands-on experience with Kubernetes and cloud platforms (GCP preferred).
- Knowledge of high-performance networking and real-time systems.
Senior SRE & Infra Engineer (GPU Cluster Platform Reliability & Infrastructure Engineer) Β· Macpower Digital Assets Edge Private Limited