IW
Senior Site Reliability Engineer (SRE) – Kubernetes Platform (FedRAMP High /IL5) - G11
Info Way Solutions LLC
Location not stated
Remote
Senior
2 months ago
- Kubernetes
- FedRAMP
- CI/CD
- Incident Response
- Devops
- EKS
- AKS
- GKE
- Linux
- IaC
- Terraform
- GitHub Actions
- GitLab CI
- Jenkins
- ArgoCD
- Python
- Prometheus
- Grafana
- OpenTelemetry
- NIST
- STIGs
- RMF
- Istio
- Linkerd
- GitOps
- FIPS
- CKA
2 months ago
IL5) - G11
Level: L3
Location: Remote, prefer PST hours
Visa Requirements: UC Citizenship
About the Role
We build technology that simply works—reliable, secure, and easy to use. We're
looking for a Senior Site Reliability Engineer (SRE) - Technical Leader to help us
design, operate, and scale a Kubernetes-based platform supporting highly
regulated environments, including FedRAMP High and DoD IL5.
This role sits at the intersection of software engineering and infrastructure. You'll
work closely with engineers across the stack to ensure our platform is resilient,
observable, compliant, and developer-friendly—without slowing teams down.
What You'll Do
• Design, build, and operate production-grade Kubernetes platforms in
regulated environments
• Improve system reliability through automation, thoughtful design, and
continuous iteration
• Define and drive SLOs, SLIs, and error budgets to guide reliability decisions
• Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
• Implement robust observability (metrics, logs, traces) to make systems
understandable and actionable
• Reduce operational toil by automating repetitive processes and improving
workflows
• Partner with security and compliance teams to meet FedRAMP High and IL5
requirements without sacrificing developer velocity
• Support ATO processes, including documentation, controls implementation,
and audit readiness
• Participate in on-call rotations supporting customer requests and paging
alerts
• Participate in incident response, blameless postmortems, and continuous
improvement efforts
• Help shape a platform that engineers enjoy using
What You Bring
• 10+ years of experience in SRE, DevOps, or infrastructure engineering
• Strong experience running Kubernetes in production (EKS, AKS, GKE, or
upstream)
• Hands-on experience working in FedRAMP High and/or DoD IL5
environments
• Solid understanding of cloud infrastructure, Linux systems, and networking
fundamentals
• Experience with Infrastructure as Code (Terraform preferred)
• Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
• Proficiency in scripting or programming (Python, Go)
• Experience building or operating observability platforms (Prometheus,
Grafana, OpenTelemetry, ELK)
• Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs,
RMF)
Nice to Have
• Experience with service mesh technologies (Istio, Linkerd)
• Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
• Experience with GitOps workflows
• Exposure to multi-cluster or hybrid cloud architectures
• Knowledge of FIPS-compliant systems or DoD Cloud SRG
• Relevant certifications (CKA, CKS, cloud provider certs, Security+)
How We Work
• We value simplicity, transparency, and collaboration
• We believe in blameless culture and learning from incidents
• We focus on building tools and platforms that empower other engineers
• We balance reliability, security, and developer experience—not one at the
expense of the others
What Success Looks Like
• Our platform is reliable, scalable, and easy to operate
• Engineers can deploy confidently in high-compliance environments
• Observability provides clear, actionable insights
• Operational overhead is minimized through automation
• Compliance requirements are met seamlessly as part of the platform
Senior Site Reliability Engineer (SRE) – Kubernetes Platform (FedRAMP High /IL5) - G11 · Info Way Solutions LLC