IL
Systems Architect, Disaster Recovery
ICONMA, LLC
🇺🇸 United States
On-site
Manager or above
1 month ago
$121.81 – $126.81 / hour
- Disaster Recovery
- Kubernetes
- NIST
- FedRAMP
- Threat Modeling
- MBSE
- AWS
- Azure
- GCP
- IAM
- VMware
- Hyper-V
- IaC
- Terraform
- CloudFormation
- CI/CD
- Jenkins
- GitLab
1 month ago
Responsibilities:
- Architectural Leadership
- Define end to end DR and high availability (HA) architectures for enterprise wide workloads, incorporating multi region cloud, hybrid, and on prem solutions.
- Develop architectural blueprints, reference designs, and pattern libraries that align with LM’s security, compliance, and cost optimization policies.
- Solution Design & Implementation
- Design and implement automated fail over, replication, and fail back mechanisms (e.g., Site Recovery Manager, Kubernetes based HA, database mirroring, storage level replication).
- Evaluate and integrate emerging technologies (e.g., Immutable Infrastructure, Chaos Engineering, Serverless DR) to improve resiliency and reduce mean time to recover (MTTR).
- Governance & Compliance
- Ensure all DR solutions meet corporate policies CRX 301, CRX 302, and relevant regulatory requirements (e.g., NIST?800 34, ISO?22301, FedRAMP).
- Create and maintain DR documentation, run books, and test plans; conduct periodic reviews and updates.
- Testing & Validation
- Lead full scale DR exercise planning, execution, and post mortem analysis for multi site, multi cloud environments.
- Define success criteria, metrics, and KPIs; report findings to senior leadership and stakeholders.
- Stakeholder Collaboration
- Partner with IT Infrastructure, Cloud Engineering, Application Development, Security, and Governance teams to embed DR/HA considerations early in the SDLC.
- Serve as the technical authority for DR during design reviews (SRR, PDR, CDR, TRR) and program risk assessments.
- Continuous Improvement
- Conduct risk assessments, threat modeling, and capacity planning to anticipate emerging resiliency challenges.
- Drive adoption of Model Based Systems Engineering (MBSE) and automated documentation tools to keep architecture artefacts current.
- DR Plan Modernization & Compliance
- Review existing DR plan architectures across the enterprise, assessing their alignment with current resilience standards, best practices, and organizational Recovery Objectives.
- Collaborate with internal teams (Application Owners, IT Service Managers, Engineering) to update and refine DR plans, ensuring that all applications and IT services meet the latest RTO/RPO targets.
- Develop and implement remediation plans to bring legacy systems and applications up to date with modern resilience standards, ensuring compliance with corporate policies (CRX 301, CRX 302) and regulatory requirements.
- Track progress and report status to senior leadership, providing insights into plan modernization efforts and risk mitigation strategies.
Requirements:
- 5+years of experience designing and implementing DR/HA solutions for enterprise scale workloads in cloud, hybrid, and on prem environments.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical discipline (or equivalent experience).
- Proven hands on experience with cloud platforms (AWS, Azure, GCP) and related services (e.g., Disaster Recovery, Site Recovery Manager, cross region replication, networking, IAM).
- Strong understanding of networking, storage, virtualization, container orchestration (Kubernetes), and database technologies as they relate to resiliency.
- Excellent written and verbal communication skills; ability to translate complex technical concepts for both technical and non technical audiences.
- Desired skills :
- Familiarity with automated disaster recovery (DR) solutions, including but not limited to:
- Amazon Web Services (AWS) Disaster Recovery Service (DRS): Experience with configuring and managing replication, fail over, and fail back processes for AWS workloads.
- Microsoft Azure Site Recovery (ASR): Knowledge of setting up and managing site recovery between on premises environments, Azure, and other clouds.
- Zerto: Hands on experience with continuous data protection (CDP) and near zero RPO replication across VMware, Hyper V, and cloud environments.
- Veeam Backup & Replication: Experience with agent less backup, replication, and automated fail over testing for virtual, physical, and cloud workloads.
- IBM Resiliency Services (formerly IBM Disaster Recovery as a Service): Familiarity with managed DR services for hybrid cloud environments, including integration with IBM Cloud and on premises infrastructure.
- Experience with DR automation, including:
- Scripting and integration with IaC tools (Terraform, CloudFormation) and CI/CD pipelines (Jenkins, GitLab) to automate DR workflows.
- DR exercise planning and execution, including defining success criteria, metrics, and KPIs for recovery processes.
- Strong analytical skills: ability to perform risk assessments, impact analysis, and cost benefit modeling for DR solutions.
- Hybrid/multi-cloud deployments, with ability to manage DR across multiple cloud providers and on premises environments.
- Advanced certifications (e.g., AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, VMware VCAP DCV).
Why Should You Apply?
- Health Benefits
- Referral Program
- Excellent growth and advancement opportunities
Systems Architect, Disaster Recovery · ICONMA, LLC