
System Reliability Engineer (Data Centre)
Centre for Strategic Infocomm Technologies
🇸🇬 Singapore
On-site
2 months ago
- Incident Management
- Encompass
2 months ago
Not enough detail in this posting to match
Responsibilities
- Oversee and manage IT operations within the data centre, including day-to-day monitoring, incident management, and problem management
- Lead the end-to-end incident management lifecycle that encompass immediate troubleshooting, root cause identification, and resolution implementation to restore services, followed by comprehensive post-incident analysis
- Develop and maintain documentation on IT infrastructure, operations, and procedures within the data centre
- Perform capacity planning to ensure IT infrastructure is scalable for future demands
- Collaborate and coordinate with Data Centre Facilities teams on matters related to power, cooling, and physical infrastructure
- Design and implement robust observability platform alongside network monitoring tools for performance monitoring and real-time alerting of IT devices and networks
- Implement and manage remote management tools for out-of-band access and control of IT devices and servers
- Define, implement, and track SRE metrics, including SLO, SLI, and error budgets to improve data centre IT reliability
Requirements (Minimum Qualifications)
- Background in Computer Science, Computer or Electrical Engineering, Information Technology or a related field
- Good technical knowledge in IT infrastructure, including servers, storage, networking, and cloud technologies
- Proficient in IT management software and tools
- 2 years of working experience in IT operations is preferred
- Fresh graduates are welcomed to apply Â
System Reliability Engineer (Data Centre) · Centre for Strategic Infocomm Technologies