
Site Reliability Engineer β Oil & Gas
- Azure
- AKS
- Incident Response
- Terraform
- Ansible
- GitHub Actions
- Azure DevOps
- Disaster Recovery
- Devops
- Kubernetes
- Linux
- IaC
- AWS
- Python
- PowerShell
- SCADA
- IEC
Client: oil and gas company (upstream operator, Houston HQ) Β·Location: remote, LATAM, Houston business hours Β·Level: senior Β·Engagement: full-time contractor, on-call rotation
You will own the reliability of the platforms behind field operations: production monitoring, data ingestion and the applications engineers use around the clock.
What you will do
Set SLOs for operations-critical systems and report on them to IT and operations leadership
Run hybrid infrastructure: Azure (AKS) plus on-premise and edge sites with limited connectivity
Build monitoring and alerting that covers both cloud services and field data flows
Lead incident response and postmortems, coordinating with OT and network teams
Automate infrastructure with Terraform and Ansible and deployments with GitHub Actions or Azure DevOps
Plan capacity and disaster recovery for 24/7 operations
What you bring
5+ years in SRE, DevOps or infrastructure engineering
Strong Kubernetes, Linux and infrastructure as code
Experience with Azure or AWS in hybrid setups
Scripting in Python, Go or PowerShell
Track record running on-call for systems where downtime has operational cost
Advanced English
Nice to have
Exposure to OT/IT integration, SCADA networks or IEC 62443
Energy, utilities or industrial background
Experience with edge computing or low-bandwidth sites
Site Reliability Engineer β Oil & Gas Β· Itmates