
Site Reliability Engineer β Fintech
- AWS
- Incident Response
- Kubernetes
- EKS
- Terraform
- Datadog
- Prometheus
- Grafana
- OpenTelemetry
- PCI DSS
- SOC2
- Devops
- Python
- Load Testing
Client: US fintech (payments platform) Β·Location: remote, LATAM, US Eastern overlap Β·Level: senior Β·Engagement: full-time contractor, on-call rotation
You will keep a payments platform on AWS reliable and compliant, owning SLOs, incident response and the automation behind both.
What you will do
Define and track SLOs and error budgets for payment-critical services
Operate Kubernetes (EKS) workloads and the Terraform code that provisions them
Run incident response and blameless postmortems, and drive the follow-up work
Build observability with Datadog or Prometheus/Grafana and OpenTelemetry
Automate deployments, rollbacks and capacity management
Support PCI DSS and SOC 2 controls with evidence and tooling
What you bring
5+ years in SRE, DevOps or production engineering
Strong AWS, Kubernetes and Terraform
Coding ability in Go or Python
Experience with on-call for high-availability, transaction-processing systems
Advanced English
Nice to have
Fintech or regulated-industry background
Chaos engineering and load testing practice
AWS certifications
Site Reliability Engineer β Fintech Β· Itmates