Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
IW

SRE ARCHITECT

Info Way Solutions LLC
๐Ÿ‡บ๐Ÿ‡ธ United States
On-site
Staff / Principal
5 months ago
  • IaC
  • Terraform
  • Ansible
  • CI/CD
  • Dynatrace
  • Grafana
  • Splunk
  • Incident Management
  • Incident Response
  • Devops
  • Python
  • Java
  • Bash
  • AWS
  • GCP
  • Azure
  • Kubernetes
  • Docker
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Job Title : SRE Architect
Job Summary
We are seeking a highly experienced Site Reliability Engineering (SRE) Architect to lead the design, implementation, and governance of highly reliable, scalable, and resilient distributed systems. This role requires a strategic thinker with deep technical expertise who can drive SRE best practices, define reliability standards, and ensure production stability across complex cloud and hybrid environments.
Key Responsibilities
Architectural Strategy
  • Design and implement scalable, resilient, and high-performance infrastructure across cloud and hybrid environments
  • Establish architectural standards for reliability and fault tolerance
SRE Governance
  • Define and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Collaborate with stakeholders to align reliability goals with business objectives
Automation & Toil Reduction
  • Drive Infrastructure-as-Code (IaC) adoption using tools like Terraform and Ansible
  • Lead automation initiatives to reduce manual operational effort ( "toilโ€)
  • Enhance CI/CD pipelines and implement self-healing systems
Observability & Monitoring
  • Design and implement observability frameworks including monitoring, logging, and distributed tracing
  • Utilize tools such as Dynatrace, Grafana, and Splunk for proactive system monitoring
Incident Management & Chaos Engineering
  • Lead incident response, root cause analysis (RCA), and postmortems
  • Implement chaos engineering practices to improve system resilience
Mentorship & Leadership
  • Mentor junior SREs and DevOps engineers
  • Promote SRE culture, best practices, and operational excellence across teams

Required Skills & Experience
  • Experience: 10โ€“12+ years in SRE, DevOps, Software Engineering, or System Administration
  • Programming/Scripting: Proficiency in Go, Python, Java, or Bash
  • Cloud Platforms: Strong experience with AWS, GCP, or Azure
  • Infrastructure as Code (IaC): Hands-on expertise with Terraform, Ansible
  • Containerization: Deep understanding of Kubernetes and Docker
  • Observability Tools: Experience with Dynatrace, Grafana, Splunk
  • Strong troubleshooting, analytical, and problem-solving skills

Preferred Qualifications
  • Experience in large-scale distributed systems
  • Exposure to enterprise environments and high-availability systems
  • Strong communication and stakeholder management skills

SRE ARCHITECT ยท Info Way Solutions LLC

Auto apply with Likeremote