MD
Datacenter Observability and Site Reliability Engineer
Macpower Digital Assets Edge Private Limited
🇮🇳 India
On-site
5 days ago
- Grafana
- Loki
- Mimir
- Kubernetes
- Python
- Bash
- Prometheus
- Docker
- Terraform
- AWS
- GCP
- Azure
5 days ago
- Experience: 8 to 12 years.
- Notice Period: Immediate to 30 days preferred.
Key Must-Have Skills:
- 5+ years in Observability Engineering.
- Expertise in Grafana, Loki, Mimir, and Alloy agent.
- Strong understanding of infrastructure metrics (e.g., GPU, CPU, Kubernetes).
- Proficiency in scripting languages ( Python, Go, Bash).
- Prior exposure to tools such as Prometheus, ELK, Docker, and Terraform.
- Flexibility to work with Korean stakeholders and time zones.
Role Highlights:
- Design and manage the observability stack across large-scale data center infrastructure.
- Build scalable telemetry systems, dashboards, alerts, and reports.
- Apply SRE best practices to ensure system reliability and performance.
- Troubleshoot real-time issues and contribute to ongoing system optimization.
Good to Have:
- Previous experience working with Korean stakeholders.
- Familiarity with cloud platforms like AWS, GCP, or Azure.
Datacenter Observability and Site Reliability Engineer · Macpower Digital Assets Edge Private Limited