AL
SRE DevOps
Artech LLC
🇺🇸 United States
On-site
2 weeks ago
$55 – $58 / hour
- Devops
- AWS
- Kubernetes
- Python
- Linux
- CI/CD
- Incident Response
- IaC
- Terraform
- CloudFormation
- Disaster Recovery
2 weeks ago
Title: SRE DevOps
Locations: Sunnyvale, CA /Austin, TX
Duration: 6 Months
Pay Range: $55 - $58/Hour on W2/C2C (All inclusive)
Note: All submissions must have LinkedIn id of Candidate
•We are looking for a skilled Site Reliability Engineer (SRE) to design, build, and maintain highly available, scalable, secure, and reliable cloud infrastructure and applications.
•The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Python, Linux, and cloud-native technologies.
•You will work closely with Development, DevOps, Security, and Operations teams to improve system reliability, automation, observability, and operational efficiency.
•Key Responsibilities Design, implement, and maintain highly available and scalable infrastructure on AWS. Deploy, manage, and troubleshoot containerized applications using Kubernetes.
•Develop automation tools, scripts, and operational utilities using Python. Build and maintain CI/CD pipelines for reliable and repeatable application deployments.
•Monitor system health, availability, performance, and capacity.
•Implement observability using metrics, logs, traces, dashboards, and alerting.
•Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews. Define and improve SLIs, SLOs, and SLAs. Automate repetitive operational tasks and reduce manual intervention.
•Perform Kubernetes troubleshooting, including pods, deployments, services, ingress, networking, storage, and resource management.
•Optimize AWS infrastructure for performance, reliability, security, and cost.
•Implement infrastructure as code using tools such as Terraform or CloudFormation.
•Establish and maintain backup, disaster recovery, and business continuity mechanisms.
•Work with development teams to improve application reliability and production readiness.
•Participate in an on-call rotation and respond to production incidents when required.
•Continuously identify opportunities to improve system resilience and operational processes.
SRE DevOps · Artech LLC