Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
PG

Site Reliability Engineer (SRE) / Platform Engineer

Perfict Global, Inc.
Location not stated
4 months ago
  • OpenShift
  • Kubernetes
  • Datadog
  • Prometheus
  • Grafana
  • Terraform
  • Ansible
  • Azure
  • GitOps
  • CI/CD
  • ArgoCD
  • Jenkins
  • GitHub Actions
  • Vault
  • Incident Response
  • NGINX
  • Secrets Management
  • Bash
  • Python
  • IaC
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Job Title : Site Reliability Engineer (SRE) / Platform Engineer
Location: Reston, VA (Hybrid — 2 days onsite / 3 days remote)
Type:3- 6 month contract to hire


Roles & Responsibilities
  • Operate, tune, and optimize OpenShift/Kubernetes clusters (scheduling, ingress, upgrades, quotas, policies).
  • Stand up and/or refineobservability (Datadog, Prometheus, Grafana)—dashboards, alerts, SLOs, runbooks.
  • Map currenthybrid topology and critical delivery pipelines; identify toil and prioritize automation (Terraform/Ansible).
  • Begin supportingAzure environments (compute, networking, storage, data services) used by analytics teams.
  • DriveGitOps-first workflows; harden CI/CD withArgoCD/Jenkins/GitHub Actions and policy-as-code guardrails.
  • Implement or enhanceplatform services (Vault, Kafka/AMQ, ingress, service mesh) for dev and data teams.
  • Lead incident response and postmortems; institutionalize RCA, blameless learning, and continuous improvement.
  • Advance thehybrid service model—migrations, integrations, reliability/latency tuning, cost and performance optimization.

Day-to-Day Responsibilities
  • Operate and optimizeOpenShift/Kubernetes clusters, ingress (e.g., Nginx), and container networking/service mesh.
  • ManageAzure services (compute, VNet, storage, data services) supporting analytics workloads.
  • Build and maintainautomated infrastructure withTerraform, Ansible, and GitOps workflows.
  • Implement and evolveobservability (Datadog, Prometheus, Grafana): metrics, traces, logs, alerting, SLOs, runbooks.
  • Design, harden, and supportdelivery pipelines withArgoCD/Jenkins/GitHub Actions.
  • Provideplatform tooling and enablement for application developers, data engineers, and operations teams.
  • Ensuresecurity and access management (HashiCorp Vault, secrets management, least privilege).
  • Lead incident response, coordinate cross-functional resolution, and drive corrective actions and platform improvements.
  • Script or develop tools inBash, Python, or Go to eliminate toil and improve developer experience.

Tech You'll Work With
  • Kubernetes / OpenShift
  • Azure (compute, networking, storage, and data services)
  • Automation & IaC: Terraform, Ansible, GitOps
  • Observability: Datadog, Prometheus, Grafana
  • Networking & Ingress: Nginx, service meshes, container networking
  • Messaging: Kafka, AMQ
  • Secrets & Access: HashiCorp Vault
  • CI/CD: ArgoCD, Jenkins, GitHub Actions
  • Scripting/Coding: Bash, Python, Go
Must-Have Qualifications
  • 5+ years hands-on operating and managingKubernetes and OpenShift clusters.
  • Strong experience withMicrosoft Azure (compute, networking, storage,and data services).
  • Proven skills inautomation and Infrastructure-as-Code (Terraform, Ansible, GitOps).
  • Proficiency withobservability tooling (Datadog, Prometheus, Grafana).
  • Scripting/coding ability inBash, Python, or Go.
Preferred / Stand-Out Skills
  • Experiencebridging on-prem and cloud in a hybrid service model (migration, integration, optimization).
  • Expertise withKafka/AMQ,HashiCorp Vault, andArgoCD/Jenkins/GitHub Actions.
  • Backgroundleading incident response and postmortems with strong RCA and continuous improvement practices.
Work Model & Team
  • Hybrid: 2 days onsite inReston, VA; 3 days remote.
  • You'll be part of theIT organization, collaborating daily withdevelopers, data engineers, infrastructure operations, and security.
How to Succeed In This Role
  • You're ahands-on engineer who thrives in regulated, high-impact environments.
  • You favorautomation over repetition, andobservability over guesswork.
  • You collaborate openly, communicate clearly, andleave systems better than you found them.


Site Reliability Engineer (SRE) / Platform Engineer · Perfict Global, Inc.

Auto apply with Likeremote