Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
SP

Site Reliability Engineer

Sphere Partners
๐ŸŒ Worldwide
Remote
Mid level
2 months ago
  • AI
  • Incident Response
  • Disaster Recovery
  • Devops
  • AWS
  • Datadog
  • CloudWatch
  • New Relic
  • Incident Management
  • PagerDuty
  • Jira
  • Confluence
  • Looker
  • Claude
  • Cursor
  • triage
  • Python
  • Bash
  • IaC
  • Terraform
  • CloudFormation
  • Kubernetes
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Contract Details

  • Notice - 1-2 weeksย 
  • 1099 contract / contractor / B2B
  • Term - 3-6 months with possible extension

Our client is a fast-growing global fintech company building a modern digital payments platform that enables fast, secure, and compliant international money transfers. Serving millions of customers across multiple markets, the company operates a highly available cloud-native infrastructure where reliability, scalability, and operational excellence are core business priorities.

We are looking for an experiencedSite Reliability Engineer to help strengthen platform reliability, improve production operations, and drive automation across the engineering organization.

About the Role

As aSite Reliability Engineer, you will own and improve the reliability, availability, and operational health of a large-scale cloud platform. You'll collaborate closely with Software Engineers, Infrastructure, Customer Operations, and Product teams while helping evolve production support processes and operational standards.

The role combinestraditional SRE responsibilities withmodern AI-assisted engineering practices, leveraging AI tools to improve incident response, documentation, operational workflows, and engineering productivity.

This position is ideal for someone who enjoys solving complex production challenges, improving observability, automating repetitive operational work, and building scalable reliability practices.

Responsibilities:

  • Act as the primary technical escalation point for critical production incidents, providing hands-on support during high-severity outages.

  • Improve platform reliability by reviewing new product launches, infrastructure changes, and production readiness before release.

  • Design, implement, and optimize monitoring, alerting, and observability solutions across cloud infrastructure and applications.

  • Analyze operational metrics, recurring alerts, and incident trends to reduce alert fatigue and improve overall system health.

  • Lead incident investigations and post-mortems, ensuring root causes are identified and preventative actions are implemented.

  • Collaborate with Engineering, Infrastructure, Customer Operations, and external support teams to coordinate incident response and customer communications.

  • Participate in capacity planning, peak traffic readiness, disaster recovery exercises, and system performance reviews.

  • Develop and maintain operational runbooks, documentation, and incident response procedures.

  • Improve internal reliability tooling and automate operational workflows using modern AI-assisted development tools.

  • Drive continuous improvements in operational excellence through automation, standardization, and proactive reliability initiatives.

Requirements

  • 4+ years of experience as a Site Reliability Engineer, DevOps Engineer, Production Engineer, or a similar infrastructure-focused role.

  • Strong experience supporting production systems running on AWS.

  • Hands-on experience with monitoring and observability platforms such asDatadog, AWS CloudWatch, New Relic, or similar.

  • Experience with incident management platforms such asPagerDuty.

  • Strong understanding of production incident management, root cause analysis, and post-incident review processes.

  • Experience working with ticketing and documentation platforms such asJira andConfluence.

  • Familiarity with operational dashboards and reporting tools (Looker or similar BI platforms).

  • Experience building operational documentation, runbooks, and support processes.

  • Comfortable working outside regular business hours when critical production incidents require senior engineering support.

  • Experience using AI-assisted engineering tools (Claude, GitHub Copilot, Cursor, or similar) to improve engineering workflows, automate documentation, incident triage, reporting, or operational tasks.

  • Strong scripting or automation skills (Python, Bash, or similar) are considered a plus.

Nice to Have

  • Experience working in fintech, payments, financial services, or other high-availability environments.

  • Experience with Infrastructure as Code (Terraform, CloudFormation, or similar).

  • Familiarity with Kubernetes and containerized environments.

Site Reliability Engineer ยท Sphere Partners

Auto apply with Likeremote