Production Reliability Engineer
- ๐บ๐ธ United States
- On-site
- Manager or above
- 10 hours ago
- ServiceNow
- Jira
- Change Management
- Incident Response
- Linux
- SQL
- Incident Management
- Bash
- Python
- Anthropic Claude
Indent: SF_OP_209943-1 / SF_OP_209943-2 / SF_OP_209943-3
Role: Production Reliability Engineer
Location: Remote
Max Rate: $75-78/hr
Skills Required:Application Production Support, ServiceNow ticketing, Jira, Incident/Problem/Change Management, (8-12 Years)
Job Description
Production Reliability Engineering services that support the stability, availability, security, and performance of the production environment.
Scope and deliverables will include:
- Performing established daily, weekly, and monthly production-support activities
- Managing assigned ServiceNow (SNOW) incidents, service requests, problem records, and change tickets within established service-level expectations
- Managing Jira work items and maintaining accurate status, documentation, and resolution details
- Troubleshooting application, operating-system, and database-related production issues
- Participating in incident response, root-cause analysis, and the development of corrective and preventive actions
- Planning, coordinating, executing, and validating monthly operating-system patching
- Following established change-management, testing, approval, maintenance-window, and rollback procedures
- Documenting patching results, compliance status, exceptions, failures, and remediation plans
- Developing and maintaining operational scripts, automation, runbooks, procedures, and support documentation
- Providing regular status reports covering completed work, open items, risks, and upcoming activities
The proposed contractor resource should meet the following minimum technical requirements:
- At least five years of Linux system administration experience
- At least five years of experience supporting SQL-based database environments
- Demonstrated experience in application production support
- Strong knowledge of IT operations, incident management, problem management, and change management
- Experience developing operational automation using Bash and/or Python
- Experience using Anthropic Claude or similar approved AI tools to automate operational workflows, documentation, analysis, or repetitive support activities
- Ability to troubleshoot issues across applications, operating systems, databases, and infrastructure
- Experience working with ServiceNow and Jira
- Strong written communication and technical-documentation skills
Performance should be measured against defined outcomes, including ticket responsiveness, SLA compliance, completion of scheduled tasks, patching success and compliance rates, automation deliverables, documentation quality, and timely escalation of risks and service-impacting issues.
Production Reliability Engineer ยท PeopleNTech LLC