SL
Site Reliability Engineer with strong experience in AWS, Kubernetes and Terraform
Staffworxs LLC
- πΊπΈ United States
- On-site
- 12 hours ago
- AWS
- Kubernetes
- Terraform
- Azure
- EKS
- AKS
- Fargate
- IaC
- Devops
- CI/CD
- GitLab CI/CD
- AI
- Incident Response
- IAM
- Secrets Management
- DevSecOps
- Vulnerability Management
- VPC
- EC2
- ECS
- RDS
- VLANs
- Python
- Bash
- Linux
- Windows
- AWS Lambda
- API Gateway
- Amazon EventBridge
- SAST
- DAST
- NIST
- ISO 27001
- SOC2
- CloudWatch
- Prometheus
- Grafana
- SIEM
12 hours ago
Job Details:
Role :- Site Reliability Engineer with strong experience in AWS, Kubernetes and Terraform
Location :- Louisville, KY (Hybrid) (on site Tuesday and Thursday )
Duration: 12 Months Contract
Job Description:
Key Responsibilities
**Reliability & Observability**
- Define and own Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for restaurant technology systems; use error budgets to balance reliability with velocity.
- Build and maintain monitoring, alerting, and observability platforms that provide meaningful signal β not noise.
- Lead blameless post-incident reviews; drive root cause analysis and ensure permanent corrective actions are implemented.
- Proactively identify reliability risks across the stack before they become incidents.
- Establish and track reliability metrics; report on system health to engineering leadership.
**Architecture & Infrastructure**
- Design and implement scalable, secure, and highly available cloud architectures primarily on AWS, with working knowledge of Azure.
- Architect and manage containerized workloads using Kubernetes (EKS, AKS), including edge Kubernetes deployments in restaurant environments.
- Design and implement serverless and container-based solutions, including AWS Fargate and other managed services.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Own architecture across the full restaurant technology stack β cloud, edge, networking, and device management β not just the cloud layer.
**DevOps, Automation & Tooling**
- Build and optimize CI/CD pipelines using GitLab CI/CD and modern DevOps practices.
- Build internal tools, automations, and middleware integrations that eliminate repetitive operational work.
- Use AI-assisted development to accelerate scripting, troubleshooting, and documentation.
- Champion a culture of engineering solutions over repeated manual fixes β if something is done twice, it should be automated.
**Restaurant & Edge Technology**
- Design and support Kubernetes-based edge systems deployed in restaurant locations.
- Support mobile application deployments and troubleshoot deployment issues across restaurant endpoints.
- Manage and optimize Mobile Device Management (MDM) platforms covering the restaurant device fleet.
- Configure and troubleshoot enterprise networking β primarily switches and restaurant-facing network infrastructure.
- Lead and participate in incident response for restaurant technology systems, including on-call coverage and post-incident review.
- Reduce mean time to detection (MTTD) and mean time to resolution (MTTR) through better tooling, runbooks, and automation.
**Security & Governance**
- Establish and enforce cloud governance, security policies, and architectural standards.
- Implement cloud security best practices: IAM strategy, network segmentation, encryption, and secrets management.
- Conduct security architecture reviews; identify vulnerabilities, misconfigurations, and compliance gaps.
- Integrate security into CI/CD pipelines (DevSecOps β Development, Security, and Operations), including automated scanning, policy validation, and vulnerability management.
**Leadership & Collaboration**
- Collaborate with engineering, DevOps, and security teams to ensure secure-by-design solutions across cloud and restaurant tech.
- Provide technical leadership and mentorship to engineering teams.
- Create and maintain documentation, runbooks, and architectural decision records.
- Continuously evaluate emerging technologies and recommend improvements.
Required Qualifications
- 6+ years in IT infrastructure, with 3+ years focused on site reliability engineering, cloud architecture, or platform engineering.
- Hands-on experience with AWS (VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3).
- Working experience with Microsoft Azure.
- Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.
- Experience with GitLab CI/CD and CI/CD pipeline design.
- Solid experience with Terraform for infrastructure provisioning.
- Experience with enterprise networking β switch configuration, VLANs, network troubleshooting.
- Familiarity with Mobile Device Management (MDM) platforms.
- Experience with automation and scripting (Python, Bash, Go, or equivalent).
- Proven ability to build internal tooling and API integrations, not just configure managed services.
- Experience defining and operating against SLOs, SLIs, and error budgets.
- Comfortable working in Linux command-line environments; Windows familiarity a plus where restaurant endpoints require it.
Preferred Qualifications
- Experience designing serverless architectures (AWS Lambda, Fargate, API Gateway, EventBridge).
- Experience with DevSecOps tooling (SAST β Static Application Security Testing, DAST β Dynamic Application Security Testing, container scanning, IaC scanning).
- Familiarity with security frameworks (CIS, NIST, ISO 27001, SOC 2).
- AWS and/or Azure certifications.
- Experience with monitoring and observability tools (CloudWatch, Prometheus, Grafana, or SIEM solutions).
- Background in restaurant, retail, or distributed edge technology environments.
- Experience using AI-assisted development tools for scripting, troubleshooting, and documentation.
Key Competencies
- Generalist mindset β comfortable moving between cloud, edge, networking, and device management in the same week.
- Reliability-first thinking β treats toil reduction, error budgets, and post-incident learning as core engineering disciplines, not afterthoughts.
- Bias toward permanent fixes and automation over repeated manual intervention.
- Ability to balance strategic architecture with hands-on execution.
- Strong communication and cross-functional collaboration skills.
- Curious, proactive, and detail-oriented approach to systems design and operations.
Staffworxs is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive workplace for all employees, regardless of race, color, religion, gender, sexual orientation, national origin, age, disability, or veteran status.
Site Reliability Engineer with strong experience in AWS, Kubernetes and Terraform Β· Staffworxs LLC