PI
Reliability Engineer 181585
PeopleSERVE, Inc.
๐บ๐ธ United States
Hybrid
2 months ago
- Angular
- Python
- JavaScript
- AWS
- Ruby
- MVC
- Jenkins
- CI/CD
- Ansible
- Bootstrap
- HTML
- CSS
- OpenStack
- PostgreSQL
- Incident Management
- Kubernetes
- EKS
- AKS
- Power BI
- Tableau
- Node.js
- Java
- Docker
- Docker Compose
- Git
- Datadog
- Splunk
- Azure
- Prometheus
- Grafana
- OpenSearch
- OpenTelemetry
- IaC
- IAM
- Terraform
2 months ago
We are looking for a systems thinking, reliability engineer who has helped teams scale through production insight, data and backup recovery, operational automation, developer guidance, real-time metrics, automation, automation, automation.
- Strong background in several of the following: Go, Angular, Python, JavaScript, AWS, RESTful services, Ruby, MVC, Jenkins CI/CD, Configuration Automation (Chef, Ansible).
- Preferred background in: Bootstrap, HTML/CSS, Shell Scripting, messaging frameworks (MQ), Service Oriented/Micro-service Architectures, OpenStack, Relational Databases (PostgreSQL).
- Comfortable working in both Public and private cloud environments.
- Crafting scalable solutions and automation to monitor the health and establish signals to drive understanding of our Container Platform environments.
- Strengthening operational processes with Fidelity support and incident management teams for our cloud ecosystem
- Working with Fidelity and cloud service provider product teams and driving ongoing reliability improvements in their Kubernetes service offerings.
- Anticipating, discovering through ongoing interaction with, and prioritizing client / partner needs to serve as their voice and guide execution of the team.
The Expertise You Have
- Bachelor's Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
- Production experience running Cloud and on-prem Storage workloads at scale
- Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
- Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
- Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau),
- Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
- 5 -7 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
- Experience building and deploying Docker images including Docker Compose
- Hands-on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
- Experience with distributed version control systems, Git preferred
- Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
- Practical experience in building cloud hosted and native applications for the enterprise. Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
- Experience in incident/crisis management and supporting critically important applications
The Skills You Bring
- Hands on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.)
- Ability to automate with various scripting languages (Python, Shell scripting, etc.)
- Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)
Additional Value in Backup & Recovery:
- Advance enterprise resiliency through improved recovery capabilities.
- Reduce recovery time via automation.
- Enable rehoused recovery into new datacenters.
- Strengthen platform reliability through data protection design.
Reliability Engineer 181585 ยท PeopleSERVE, Inc.