DevOps Cloud Engineer
- πΊπΈ United States
- Hybrid
- Mid level
- 3 hours ago
- $50 β $54 / hour
- Devops
- AWS
- CI/CD
- IaC
- Linux
- Kubernetes
- Python
- Bash
- IAM
- Load Balancing
- Disaster Recovery
- Terraform
- Ansible
- Docker
- CloudWatch
- Prometheus
- Grafana
- Splunk
- Incident Response
- Secrets Management
- Incident Management
- Vulnerability Management
DevOps Cloud Engineer
Location: Austin, TX β Onsite
Experience: 11+ Years
Job Overview
We are seeking an experienced DevOps Cloud Engineer to design, implement, operate, and support highly available and scalable infrastructure across AliCloud and AWS environments. The ideal candidate will have strong expertise in SRE, cloud infrastructure, CI/CD, Infrastructure as Code, automation, Linux, Kubernetes, monitoring, security, and production operations.
Key Responsibilities
- Design, implement, operate, and support highly available and scalable infrastructure across AliCloud and AWS.
- Apply SRE principles to improve availability, latency, performance, capacity, scalability, and operational efficiency.
- Define, track, and continuously improve SLIs, SLOs, and reliability metrics for critical services.
- Build and maintain automated CI/CD pipelines for application and infrastructure deployments.
- Automate infrastructure provisioning, configuration, deployments, health checks, maintenance, and repetitive operational activities.
- Develop automation using Python, Bash/Shell, or equivalent scripting languages.
- Administer and troubleshoot Linux-based production systems, including:
- Operating systems
- Processes and services
- Storage
- Networking
- Permissions
- Patching
- System performance
- Security
- Support cloud services covering:
- Compute
- Storage
- Networking
- IAM/Security
- Load Balancing
- Monitoring
- Logging
- Backup
- Disaster Recovery
- Implement and maintain Infrastructure as Code using tools such as Terraform.
- Implement configuration automation using tools such as Ansible.
- Deploy, scale, troubleshoot, and support Docker and Kubernetes environments.
- Implement comprehensive monitoring, logging, alerting, and dashboards using:
- CloudWatch
- Prometheus
- Grafana
- ELK
- Splunk
- Comparable observability platforms
- Participate in production incident response, systematic troubleshooting, and root-cause analysis.
- Drive corrective and preventive actions through blameless post-incident reviews.
- Develop and maintain operational runbooks, automation playbooks, recovery procedures, and technical documentation.
- Identify recurring operational issues and eliminate manual toil through automation and engineering.
- Support capacity planning, performance tuning, high availability, failover, backup/recovery, and disaster recovery.
- Collaborate with development, infrastructure, security, network, database, and application support teams.
- Promote secure DevOps/SRE practices, including:
- Least-privilege access
- Secrets management
- Vulnerability remediation
- Patching
- Compliance controls
- Participate in change, release, incident, and problem management processes.
- Provide production support as required.
Required Skills
- 11+ years of DevOps, Cloud Engineering, SRE, or Infrastructure Engineering experience.
- Strong experience with AWS and AliCloud.
- Strong understanding of SRE principles, SLIs, SLOs, and reliability engineering.
- Hands-on experience with CI/CD pipelines.
- Strong scripting experience with Python and Bash/Shell.
- Strong Linux administration and production troubleshooting skills.
- Hands-on experience with Terraform and Infrastructure as Code.
- Experience with Ansible/configuration automation.
- Strong experience with Docker and Kubernetes.
- Experience with cloud services including compute, storage, networking, IAM, load balancing, monitoring, logging, backup, and disaster recovery.
- Strong observability experience with CloudWatch, Prometheus, Grafana, ELK/Splunk, or similar tools.
- Experience with production incident management, troubleshooting, RCA, and preventive actions.
- Strong understanding of cloud security, secrets management, vulnerability remediation, and compliance.
- Excellent communication and cross-functional collaboration skills.
Core Skills
AWS | AliCloud | DevOps | Cloud Engineering | SRE | SLIs | SLOs | Reliability Engineering | CI/CD | Terraform | Ansible | Python | Bash | Linux | Docker | Kubernetes | CloudWatch | Prometheus | Grafana | ELK | Splunk | Infrastructure as Code | Cloud Infrastructure | IAM | Networking | Load Balancing | Monitoring | Logging | Disaster Recovery | Backup & Recovery | Incident Management | Root Cause Analysis | Automation | Security | Secrets Management | Vulnerability Management | Capacity Planning | Performance Tuning
DevOps Cloud Engineer Β· Apolis