Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
IW

DevOps and Site Reliability Engineer (SRE): Platform Engineering

Info Way Solutions LLC
  • 🇺🇸 United States
  • On-site
  • 1 day ago
  • Devops
  • Azure
  • IaC
  • Configuration Management
  • CI/CD
  • Kubernetes
  • GitOps
  • Cassandra
  • OpenSearch
  • Debezium
  • Power BI
  • SFTP
  • FTP
  • Python
  • Node.js
  • AKS
  • Terraform
  • Ansible
  • Helm
  • Argo
  • Azure DevOps
  • GitHub Actions
  • Bash
  • Elasticsearch
  • Kafka Connect
  • Prometheus
  • Grafana
  • Fluent Bit
  • Fluentd
  • AWS
  • EKS
  • GCP
  • GKE
  • Argo Workflows
  • GitHub
  • cosign
  • OpenTelemetry
  • Disaster Recovery
  • CKA
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

DevOps and Site Reliability Engineer (SRE): Platform Engineering
Phoenix, AZ

About the role

We are looking for a DevOps Engineer / SRE to join a new Platform, DevOps and SRE team. We need someone who has run customer-facing production systems at scale and been accountable for their uptime. You should have a platform engineering mindset: you automate repeated work, spot operational problems before they become incidents, and enjoy helping development teams ship faster and more safely.

What you will do

- Build and run Azure infrastructure using infrastructure as code and configuration management.

- Build and maintain CI/CD pipelines with security and quality gates.

- Deploy and operate workloads on Kubernetes using GitOps practices.

- Operate data services such as Kafka, Cassandra and OpenSearch.

- Support data integration infrastructure: change data capture with Debezium, Azure Data Lake storage, Power BI connectivity, and SFTP/FTP file transfers with partners.

- Set up monitoring, logging and alerting, and define reliability targets with development teams.

- Automate operational tasks in Python.

- Troubleshoot complex production problems across infrastructure, Kubernetes and applications, and lead blameless postmortems.

- Work with development teams to improve deployment, release and operational processes.

- Take part in an on-call rotation.

Must-have

- [5]+ years in DevOps, SRE or platform engineering roles

- Proven experience managing production, customer-facing, high-scale environments: owning uptime, handling incidents as primary on-call, running releases and changes in live systems, and planning capacity for peak traffic. Be ready to describe the scale, for example requests per second, cluster or node count, data volume, or user base.

- Strong hands-on Microsoft Azure: networking (VNets, private endpoints), identity (Entra ID, managed identities), AKS

- Terraform for infrastructure as code, including writing reusable modules and managing state

- Ansible for configuration management

- Kubernetes in production (AKS preferred), including Helm charts

- GitOps practice with Argo CD or Flux

- CI/CD pipelines in Azure DevOps and/or GitHub Actions

- Python for automation; Bash

- Hands-on operation of at least one of Kafka, Cassandra, or OpenSearch/Elasticsearch in production

- Working knowledge of data integration infrastructure:

- Azure Data Lake Storage Gen2: access control, private endpoints, lifecycle policies

- Power BI: gateways, workspace and capacity administration, secure connectivity to data sources

- Debezium change data capture with Kafka Connect

- SFTP/FTP file transfer, including Azure Storage SFTP, automation and key management

- Monitoring and observability with Prometheus and Grafana; log pipelines (Fluent Bit or Fluentd)

Nice-to-have

- Experience with other clouds (AWS EKS or GCP GKE) or multi-cloud design

- Workflow orchestration (Argo Workflows) or a developer portal (Backstage)

- Supply-chain security: GitHub Advanced Security, Cosign, SBOMs, OPA or Kyverno policies

- Progressive delivery with Argo Rollouts; autoscaling with KEDA

- OpenTelemetry; SLOs and error budgets

- Disaster recovery and failover testing, or chaos engineering

- Certifications: CKA/CKAD, HashiCorp Terraform Associate, Azure (AZ-104 / AZ-400)

- Retail or high-traffic e-commerce environments, including peak events such as holiday seasons

Core Technologies

Azure · AKS · Terraform · Ansible · Python · Kubernetes · Helm · Argo CD · GitOps · GitHub Actions · Azure DevOps · CI/CD · Kafka · Debezium · Cassandra · OpenSearch/Elasticsearch · Azure Data Lake · Power BI · SFTP/FTP · Prometheus · Grafana · Fluent Bit · SRE · Platform Engineering

DevOps and Site Reliability Engineer (SRE): Platform Engineering · Info Way Solutions LLC

Auto apply with Likeremote