U
Production Support Engineer
UST
๐ต๐ญ Philippines
On-site
1 month ago
- Incident Management
- triage
- Dynatrace
- Splunk
- Grafana
- ServiceNow
- AWS
- Microservices
- Linux
- Prometheus
- CloudWatch
- Devops
- Python
1 month ago
We are looking for a Production Support Engineer to monitor and maintain the health of applications, infrastructure, APIs, and services in a 24/7 production environment. The role involves real-time monitoring, incident management, troubleshooting, and collaboration with engineering teams to ensure system reliability and performance.
Key Responsibilities
- Monitor system health across applications, infrastructure, APIs, and services using observability and monitoring tools.
- Review and responds to dashboards and metrics in real time.
- Create, expand, and maintain dashboards to improve observability coverage.
- Perform initial triage using logs, traces, and metrics; identify symptoms and potential root causes.
- Execute runbooks/SOPs for common production issues (e.g., restarts, validation checks, health checks).
- Create incidents, document findings, and collaborate with engineering teams for issue resolution.
- Coordinate with on-call responders during high-priority events.
- Perform routine health checks and proactive monitoring tasks every shift.
- Provide clear communication and shift-to-shift handoff notes.
- Participate in incident management activities and support service restoration efforts.
Required Skills
- 2โ5 years of experience in Production Support or Operations Support roles.
- Hands-on experience with observability and monitoring tools such as Dynatrace, Splunk, Grafana, or similar platforms.
- Ability to analyze logs, metrics, traces, and s to identify production issues.
- Strong troubleshooting and analytical skills with the ability to investigate and resolve issues efficiently.
- Experience executing operational runbooks and support procedures.
- Knowledge of ITSM tools such as ServiceNow.
- Strong communication and documentation skills.
- Ability to work in fast-paced production environments.
- Willingness to work in 24/7 rotational shifts.
- Eagerness to learn new tools and technologies.
Preferred Skills
- Basic understanding of AWS or other cloud platforms.
- Familiarity with microservices, APIs, Linux fundamentals, and networking concepts.
- Exposure to monitoring tools such as Prometheus and CloudWatch.
- Understanding of incident management, problem management, and SRE/DevOps practices.
- Basic scripting knowledge (Shell, Python, or similar).
Shift
- 24/7 rotational shift schedule
- Includes night shifts, weekends, and holidays
- Requires participation in on-call rotations and high-priority incident handling
Production Support Engineer ยท UST