Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
M

Senior Site Reliability Engineer

Mindlance
  • πŸ‡ΊπŸ‡Έ United States
  • Hybrid
  • Senior
  • 12 hours ago
  • CI/CD
  • Kubernetes
  • Devops
  • AWS
  • Azure
  • EKS
  • AKS
  • Fargate
  • IaC
  • Terraform
  • GitLab CI/CD
  • AI
  • Incident Response
  • IAM
  • Secrets Management
  • DevSecOps
  • Vulnerability Management
  • VPC
  • EC2
  • ECS
  • RDS
  • VLANs
  • Python
  • Bash
  • Linux
  • Windows
  • AWS Lambda
  • API Gateway
  • Amazon EventBridge
  • SAST
  • DAST
  • NIST
  • ISO 27001
  • SOC2
  • CloudWatch
  • Prometheus
  • Grafana
  • SIEM
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Overview
Client US is seeking an experienced and pragmatic Senior Site Reliability Engineer to own the reliability, design, implementation, and continuous improvement of the infrastructure that powers restaurant technology β€” from the cloud platforms and CI/CD (Continuous Integration/Continuous Deployment) pipelines we build on, to the Kubernetes-based edge systems deployed in restaurants, to the networks, MDM platforms, and automation tooling that keeps everything running.

This is a generalist role at a senior level. The right candidate brings strong cloud and reliability engineering fundamentals but is equally comfortable working across edge infrastructure, enterprise networking, mobile deployments, and operational automation. You will define what "reliable" means for our systems, measure it, and engineer solutions to continuously improve it. You'll work closely with the Global Reliability Engineering (GRE), DevOps, and security teams, and you'll be expected to shift fluidly between strategic architecture and hands-on execution.

Key Responsibilities
**Reliability & Observability**

- Define and own Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for restaurant technology systems; use error budgets to balance reliability with velocity.
- Build and maintain monitoring, alerting, and observability platforms that provide meaningful signal β€” not noise.
- Lead blameless post-incident reviews; drive root cause analysis and ensure permanent corrective actions are implemented.
- Proactively identify reliability risks across the stack before they become incidents.
- Establish and track reliability metrics; report on system health to engineering leadership.

**Architecture & Infrastructure**
- Design and implement scalable, secure, and highly available cloud architectures primarily on AWS, with working knowledge of Azure.
- Architect and manage containerized workloads using Kubernetes (EKS, AKS), including edge Kubernetes deployments in restaurant environments.
- Design and implement serverless and container-based solutions, including AWS Fargate and other managed services.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Own architecture across the full restaurant technology stack β€” cloud, edge, networking, and device management β€” not just the cloud layer.

**DevOps, Automation & Tooling**
- Build and optimize CI/CD pipelines using GitLab CI/CD and modern DevOps practices.
- Build internal tools, automations, and middleware integrations that eliminate repetitive operational work.
- Use AI-assisted development to accelerate scripting, troubleshooting, and documentation.
- Champion a culture of engineering solutions over repeated manual fixes β€” if something is done twice, it should be automated.

**Restaurant & Edge Technology**
- Design and support Kubernetes-based edge systems deployed in restaurant locations.
- Support mobile application deployments and troubleshoot deployment issues across restaurant endpoints.
- Manage and optimize Mobile Device Management (MDM) platforms covering the restaurant device fleet.
- Configure and troubleshoot enterprise networking β€” primarily switches and restaurant-facing network infrastructure.
- Lead and participate in incident response for restaurant technology systems, including on-call coverage and post-incident review.
- Reduce mean time to detection (MTTD) and mean time to resolution (MTTR) through better tooling, runbooks, and automation.

**Security & Governance**
- Establish and enforce cloud governance, security policies, and architectural standards.
- Implement cloud security best practices: IAM strategy, network segmentation, encryption, and secrets management.
- Conduct security architecture reviews; identify vulnerabilities, misconfigurations, and compliance gaps.
- Integrate security into CI/CD pipelines (DevSecOps β€” Development, Security, and Operations), including automated scanning, policy validation, and vulnerability management.

**Leadership & Collaboration**
- Collaborate with engineering, DevOps, and security teams to ensure secure-by-design solutions across cloud and restaurant tech.
- Provide technical leadership and mentorship to engineering teams.
- Create and maintain documentation, runbooks, and architectural decision records.
- Continuously evaluate emerging technologies and recommend improvements.

Required Qualifications
- 6 years in IT infrastructure, with 3 years focused on site reliability engineering, cloud architecture, or platform engineering.
- Hands-on experience with AWS (VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3).
- Working experience with Microsoft Azure.
- Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.
- Experience with GitLab CI/CD and CI/CD pipeline design.
- Solid experience with Terraform for infrastructure provisioning.
- Experience with enterprise networking β€” switch configuration, VLANs, network troubleshooting.
- Familiarity with Mobile Device Management (MDM) platforms.
- Experience with automation and scripting (Python, Bash, Go, or equivalent).
- Proven ability to build internal tooling and API integrations, not just configure managed services.
- Experience defining and operating against SLOs, SLIs, and error budgets.
- Comfortable working in Linux command-line environments; Windows familiarity a plus where restaurant endpoints require it.

Preferred Qualifications
- Experience designing serverless architectures (AWS Lambda, Fargate, API Gateway, EventBridge).
- Experience with DevSecOps tooling (SAST β€” Static Application Security Testing, DAST β€” Dynamic Application Security Testing, container scanning, IaC scanning).
- Familiarity with security frameworks (CIS, NIST, ISO 27001, SOC 2).
- AWS and/or Azure certifications.
- Experience with monitoring and observability tools (CloudWatch, Prometheus, Grafana, or SIEM solutions).
- Background in restaurant, retail, or distributed edge technology environments.
- Experience using AI-assisted development tools for scripting, troubleshooting, and documentation.

Key Competencies
- Generalist mindset β€” comfortable moving between cloud, edge, networking, and device management in the same week.
- Reliability-first thinking β€” treats toil reduction, error budgets, and post-incident learning as core engineering disciplines, not afterthoughts.
- Bias toward permanent fixes and automation over repeated manual intervention.
- Ability to balance strategic architecture with hands-on execution.
- Strong communication and cross-functional collaboration skills.
- Curious, proactive, and detail-oriented approach to systems design and operations.

β€œMindlance is an Equal Opportunity Employer and does not discriminate in employment on the basis of – Minority/Gender/Disability/Religion/LGBTQI/Age/Veterans.”

Senior Site Reliability Engineer Β· Mindlance

Auto apply with Likeremote