TC
Site Reliability Engineer
TechDigital Corporation
Location not stated
Hybrid
1 month ago
- Python
- scikit-learn
- TensorFlow
- PyTorch
- Pandas
- NumPy
- Power BI
- Tableau
- SQL
- IaC
- Incident Response
- CI/CD
- Disaster Recovery
- Linux
- Unix
- PowerShell
- AWS
- Azure
- GCP
- Docker
- Kubernetes
- Prometheus
- Grafana
- Splunk
- Dynatrace
- Datadog
- Jenkins
- GitHub Actions
- GitLab CI
- Azure DevOps
- Terraform
- Ansible
- CloudFormation
- Devops
1 month ago
Key Responsibilities
• Monitor, maintain, and improve the reliability, availability, and performance of production systems.
• Design and implement monitoring, alerting, logging, and observability solutions.
• Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
• Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
• Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
• Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
• Perform capacity planning, performance tuning, and scalability assessments.
• Support CI/CD pipelines and deployment automation initiatives.
• Implement high-availability, disaster recovery, and failover strategies.
Required Skills
• Strong experience with Linux/Unix administration.
• Proficiency in scripting languages such as Python, Shell, or PowerShell.
• Hands-on experience with cloud platforms (AWS, Azure, or GCP).
• Experience with containerization technologies such as Docker and Kubernetes.
• Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
• Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
• Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
• Strong troubleshooting, debugging, and problem-solving skills.
• Understanding of networking, security, and distributed systems concepts.
Experience
• 5–10+ years of overall IT experience.
• 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.
Site Reliability Engineer · TechDigital Corporation