Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
TC

Site Reliability Engineer

TechDigital Corporation
Location not stated
Hybrid
1 month ago
  • Python
  • scikit-learn
  • TensorFlow
  • PyTorch
  • Pandas
  • NumPy
  • Power BI
  • Tableau
  • SQL
  • IaC
  • Incident Response
  • CI/CD
  • Disaster Recovery
  • Linux
  • Unix
  • PowerShell
  • AWS
  • Azure
  • GCP
  • Docker
  • Kubernetes
  • Prometheus
  • Grafana
  • Splunk
  • Dynatrace
  • Datadog
  • Jenkins
  • GitHub Actions
  • GitLab CI
  • Azure DevOps
  • Terraform
  • Ansible
  • CloudFormation
  • Devops
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management

Key Responsibilities
• Monitor, maintain, and improve the reliability, availability, and performance of production systems.
• Design and implement monitoring, alerting, logging, and observability solutions.
• Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
• Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
• Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
• Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
• Perform capacity planning, performance tuning, and scalability assessments.
• Support CI/CD pipelines and deployment automation initiatives.
• Implement high-availability, disaster recovery, and failover strategies.

Required Skills
• Strong experience with Linux/Unix administration.
• Proficiency in scripting languages such as Python, Shell, or PowerShell.
• Hands-on experience with cloud platforms (AWS, Azure, or GCP).
• Experience with containerization technologies such as Docker and Kubernetes.
• Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
• Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
• Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
• Strong troubleshooting, debugging, and problem-solving skills.
• Understanding of networking, security, and distributed systems concepts.
Experience
• 5–10+ years of overall IT experience.
• 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.

Site Reliability Engineer · TechDigital Corporation

Auto apply with Likeremote