PL
Senior Site Reliability Engineer
PeopleNTech LLC
πΊπΈ United States
On-site
Senior
4 months ago
- CloudWatch
- Grafana
- Prometheus
- Datadog
- Terraform
- Ansible
- CloudFormation
- CI/CD
- GitHub Actions
- GitLab CI
- Jenkins
- Kubernetes
- ECS
- EKS
- Canary Releases
- Incident Response
- Incident Management
- ISO 27001
- SOC2
- IAM
- Devops
- AWS
- EC2
- RDS
- VPC
- IaC
- Python
- Bash
- PowerShell
- DNS
- Load Balancing
- CKA
- HIPAA
- GDPR
4 months ago
Position βSenior Site Reliability Engineer
Experience:7+ years
# of positions:2
C2C - $68/hr USD
Location: Remote
Job Description:
JOB DESCRIPTION
Reliability & Performance
- Design and implement monitoring, alerting, and reliability tooling usingCloudWatch, Grafana, Prometheus, Datadog, or ELK.
- Analyze production performance, capacity, and error budgets to maintain agreed SLIs/SLOs.
- Implement automated health checks, scaling rules, and self-recovery mechanisms to minimize manual intervention.
- Drive root cause analysis (RCA) and post-incident reviews, ensuring permanent fixes and documentation.
- Build automation for deployment, configuration, and infrastructure management using Terraform, Ansible, or CloudFormation.
- Develop and maintain CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins.
- Manage and optimize containerized and serverless workloads (Kubernetes, ECS, EKS, Lambda).
- Implement automated rollbacks, blue/green deployments, and canary releases.
- Participate in 24/7 on-call rotation for critical systems and lead incident management for your domain.
- Reduce mean time to detection (MTTD) and mean time to recovery (MTTR) through proactive automation and observability.
- Develop runbooks and operational playbooks for global SRE teams.
- Embed security practices into automation and deployment processes.
- Ensure systems adhere to ISO 27001 and SOC 2 requirements through continuous compliance monitoring.
- Manage IAM policies, secrets, and network configurations securely and efficiently.
- Partner with developers to design for operability, scalability, and resilience from day one.
- Contribute to cross-team reliability reviews and platform improvement initiatives.
- Champion DevOps and reliability culture across Amtech's engineering organization.
QUALIFICATIONS
- 6+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering roles.
- Strong background in AWS (EC2, ECS/EKS, RDS, Lambda, S3, IAM, VPC).
- Proficiency with Infrastructure-as-Code and automation (Terraform, Ansible, CloudFormation).
- Experience with observability tools (Prometheus, Grafana, CloudWatch, ELK, or Datadog).
- Scripting and automation skills (Python, Bash, Go, or PowerShell).
- Solid understanding of networking, DNS, and load balancing.
- Strong troubleshooting, incident management, and root cause analysis skills.
- Excellent communication and collaboration abilities in a cross-functional, distributed environment.
- Certifications such as AWS Certified SysOps Administrator, SRE Foundation, or CKA.
- Experience with chaos engineering or resilience testing tools.
- Familiarity with SLO/SLI error budget management.
- Exposure to multi-region, multi-account, or hybrid architectures.
- Background supporting SaaS platforms or regulated environments (SOC 2, HIPAA, GDPR).
Senior Site Reliability Engineer Β· PeopleNTech LLC