Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

Sr. Site Reliability Engineer

Toters
๐Ÿ‡ฑ๐Ÿ‡ง Lebanon
On-site
Senior
40 months ago
  • Incident Management
  • Grafana
  • X-ray
  • Node.js
  • Devops
  • EC2
  • CloudWatch
  • ECS
  • IAM
  • CI/CD
  • Sentry
  • Elastic Stack
  • Kubernetes
  • Terraform
  • Ansible
  • ITIL
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

The Company

Toters is an on-demand e-commerce and delivery platform and operates a service that enables customers to get anything in their city at the highest level of convenience.

At Toters, technology is at the heart of everything we do. We have product teams that are working hard every day to create products that make our customers' lives easier. Our engineers are also continuously creating solutions to make our processes more efficient, all in an effort to get to our customers fast and at the best cost. If you are interested in working in a high growth startup environment, and look to be part of a team that will potentially change the way customers shop in the Middle East, apply now.


About the Role

We are looking for aSenior Site Reliability Engineer who will play a critical role in ensuring high availability, performance, and resilience across our production systems. You will be at the heart of operational excellence, leading high-impact incident responses, building proactive monitoring systems, and engineering automation that prevents outages before they happen. If you love solving complex distributed system challenges and thrive in high-pressure environments, this role is for you.


Key Responsibilities

Incident Management & Reliability

  • Act asIncident Commander during major outages, leading real-time diagnosis, communication, and recovery.
  • Own and improve the end-to-endincident management lifecycle, including post-incident reviews and action plans.
  • Driveroot cause analysis and proactive reliability improvements to prevent recurrence.

Monitoring & Observability

  • Design and maintainmetrics, alerts, and dashboards usingPrometheus,Grafana, andNew Relic.
  • ImplementSLIs/SLOs to monitor service health and drive availability targets (99.99%+ uptime).
  • Integratelog management anddistributed tracing with tools likeELK Stack andAWS X-Ray.

Automation & Tooling

  • Developautomation scripts and internal tooling inPython or Node.js to reduce manual ops and accelerate recovery (MTTR improvement).
  • Build self-healing infrastructure usingIaC andautomation pipelines.
  • Optimize on-call workflows, escalation policies, and runbooks usingPagerDuty.

Cloud Infrastructure

  • Operate and improve infrastructure hosted onAWS, ensuring reliability, cost efficiency, and scalability.
  • Collaborate with backend and platform teams to embedSRE best practices across engineering.


Key Qualifications

  • 4+ years of experience inSite Reliability Engineering, DevOps, or Platform Engineering.
  • Proven success managingproduction incidents and participating inon-call rotations.
  • Strong hands-on experience withPrometheus, Grafana, andPagerDuty.
  • Proficient inPython or Node.js for automation and tooling.
  • Experience withAWS services (EC2, CloudWatch, ECS/Lambda, IAM, etc.).
  • Solid understanding ofLinux systems, networking, and CI/CD pipelines.


Nice to Have

  • Experience asIncident Commander in mission-critical environments.
  • Knowledge ofNew Relic,Sentry,ELK Stack, orDatadog.
  • Background implementingSLIs/SLOs/Error Budgets (Google SRE model).
  • Familiarity withDocker, Kubernetes, Terraform, or Ansible.
  • Certifications such as:
    • AWS Solutions Architect Associate/DevOps Engineer
    • ITIL Foundation or relevant reliability certifications.

Sr. Site Reliability Engineer ยท Toters

Auto apply with Likeremote