Resiliency Test Lead
- Performance Testing
- JMeter
- AWS
- Azure
- Disaster Recovery
- CI/CD
- Jenkins
- GitHub Actions
- GitLab CI/CD
- Azure DevOps
- Dynatrace
- Datadog
- Grafana
- Kibana
- Splunk
- Devops
Resiliency Test Lead
Location: Remote
Duration: Long Term
Experience: 6β10 Years
Job Summary
We are seeking an experienced Resiliency Test Lead with strong expertise in Performance Testing, Resiliency Testing, and Chaos Engineering. The ideal candidate will have hands-on experience designing and executing performance and resiliency test strategies, troubleshooting complex performance issues, and validating system reliability, scalability, and availability.
Key Responsibilities
-
Design, develop, and execute Load, Scalability, Stress, Endurance, and Capacity tests using tools such as NeoLoad, LoadRunner, and JMeter.
-
Lead Resiliency and Chaos Engineering initiatives using tools such as Gremlin, Litmus Chaos, Chaos Mesh, AWS Fault Injection Simulator (FIS), or Azure Chaos Studio.
-
Troubleshoot performance issues, identify system bottlenecks, perform Root Cause Analysis (RCA), and provide optimization recommendations.
-
Conduct code-level investigations to identify performance and reliability issues.
-
Validate High Availability (HA), Failover, Disaster Recovery (DR), and System Reliability scenarios.
-
Apply Site Reliability Engineering (SRE) principles to performance and resiliency testing.
-
Integrate and execute performance tests through CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
-
Monitor and analyze application performance using AppDynamics, Dynatrace, Datadog, Grafana, Kibana, and Splunk.
-
Identify performance bottlenecks across multiple layers, including UI, Application, Database, Network, and Message Queues.
-
Collaborate with development, DevOps, infrastructure, architecture, and SRE teams to resolve performance and resiliency issues.
-
Prepare test strategies, test plans, execution reports, performance analysis, and resiliency assessment reports.
-
Support continuous improvement of application reliability, scalability, and operational resilience.
-
Healthcare domain experience is a plus.
Required Skills
-
6β10 years of experience in Performance and Resiliency Testing.
-
Strong hands-on experience with Chaos Testing / Chaos Engineering.
-
Strong experience with Gremlin.
-
Hands-on experience with Dynatrace and other APM/monitoring tools.
-
Experience with performance testing tools such as NeoLoad, LoadRunner, or JMeter.
-
Strong understanding of SRE, HA, Failover, DR, and System Reliability Testing.
-
Experience integrating performance testing into CI/CD pipelines.
-
Strong troubleshooting, bottleneck identification, and Root Cause Analysis skills.
-
Experience analyzing performance across application, database, network, UI, and messaging layers.
Top 3 Required Skills
-
Chaos Testing / Chaos Engineering
-
Gremlin
-
Dynatrace
Preferred / Value-Added Skills
-
Healthcare domain experience
-
AWS Fault Injection Simulator (FIS)
-
Azure Chaos Studio
-
Litmus Chaos
-
Chaos Mesh
-
Jenkins / GitHub Actions / GitLab CI/CD / Azure DevOps
-
AppDynamics / Datadog / Grafana / Kibana / Splunk
Resiliency Test Lead Β· Apolis