Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
L

Specialist - Software Engineering

LTM
Location not stated
11 months ago
  • Kubernetes
  • AWS
  • Azure
  • GCP
  • EKS
  • AKS
  • GKE
  • Microservices
  • IaC
  • Terraform
  • Helm
  • CloudFormation
  • Prometheus
  • Grafana
  • Datadog
  • CI/CD
  • Devops
  • RBAC
  • Python
  • Bash
  • Linux
  • Jenkins
  • GitLab CI
  • ArgoCD
  • CKA
  • Istio
  • Linkerd
  • GitOps
  • Disaster Recovery
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Role-Kubernetes and Cloud SRE

Type-FTE/Perm

Mode-5 Days from Basildon

Site Reliability Engineer SRE Kubernetes Cloud

Position Summary

We are seeking a highly skilled Site Reliability Engineer SRE with deep expertise in Kubernetes

and cloud technologies AWS Azure or GCP The SRE will be responsible for designing deploying

automating and supporting highly available scalable and secure containerized applications in

cloudnative environments You will work closely with development operations and security teams

to ensure the reliability performance and efficiency of our production systems

Key Responsibilities

Design deploy and manage Kubernetes clusters onpremises andor cloudmanaged

such as EKS AKS GKE to support scalable microservices architectures

Automate infrastructure provisioning and application deployment using Infrastructure

as Code IaC tools such as Terraform Helm or CloudFormation

Monitor troubleshoot and optimize system performance using observability tools

Prometheus Grafana ELK Datadog etc

Implement and manage CICD pipelines to ensure rapid repeatable and reliable software

delivery

Ensure system reliability availability and security through proactive monitoring incident

response and root cause analysis

Develop and maintain runbooks dashboards and documentation for operational

procedures and system architectures

Participate in oncall rotations and respond to production incidents ensuring minimal

downtime and fast recovery

Collaborate with development and operations teams to drive DevOps and SRE best

practices including capacity planning scaling and cost optimization

Continuously improve automation tooling and processes to reduce manual work and

increase system reliability

Required Skills Experience

3 years experience as an SRE DevOps Engineer or similar role supporting largescale

productiongrade environments

Expertise in Kubernetes deployment scaling upgrades troubleshooting networking

RBAC etc

Handson experience with at least one major cloud provider AWS Azure or GCP

Proficiency in scriptingprogramming Python Bash Go etc

Experience with IaC tools Terraform Helm CloudFormation ARM etc

Strong knowledge of Linux systems administration and networking concepts

Familiarity with monitoring logging and ing tools Prometheus Grafana ELKEFK

Datadog etc

Experience with CICD tools Jenkins GitLab CI ArgoCD etc

Understanding of security best practices in cloud and containerized environments

Excellent troubleshooting and problemsolving skills

Strong communication and collaboration skills

Preferred Qualifications

Certified Kubernetes Administrator CKA or similar certification

Experience with service mesh Istio Linkerd ingress controllers and API gateways

Experience in a multicloud or hybrid cloud environment

Familiarity with GitOps practices and tools ArgoCD Flux

Experience with disaster recovery backup and business continuity planning

Education

Bachelors degree in Computer Science Engineering or related field or equivalent

experience

This role is ideal for engineers who are passionate about automation reliability and modern

cloudnative architectures and who thrive in fastpaced collaborative environments

Specialist - Software Engineering · LTM

Auto apply with Likeremote