VS
Principal Azure Capacity Manager
Veterans Sourcing Group
🇺🇸 United States
On-site
Manager or above
7 months ago
- Azure
- IaC
- App Services
- Key Vault
- FedRAMP
- Azure SQL
- RBAC
- AKS
- Performance Testing
- k6
- JMeter
- Prometheus
- Grafana
- FIPS
- Terraform
- Bicep
- Ansible
- CI/CD
7 months ago
Duration: 12+ Months (Possible extension)
Location:New York, NY 10286
Onsite Role (4 days a week)
Responsibilities:
- Principal Azure Capacity Manager (Consultant) to lead capacity planning and optimization for an Azure public cloud project operating to High requirements.
- This role ensures adequate, resilient capacity and buffer across compute, storage, network, and platform services; supports Site Reliability Engineers (SREs) with performance and reliability goals; and drives evidence-based compliance with program High control expectations.
- Own the end-to-end Capacity Management operating model for Azure services in scope of the High program—planning, modeling, forecasting, monitoring, tuning, and governance.
- Ensure sufficient capacity and engineered buffer to meet service-level objectives (SLOs), recovery objectives (RTO/RPO), and regulatory/contractual requirements, with particular focus on U.S.-only region restrictions and continuous monitoring.
- Partner closely with SREs to operationalize capacity practices through IaC, gated change control, performance baselines, autoscaling policies, and resilience patterns.
- Contribute to documentation and evidence (e.g., SSP updates, control narratives, POA&M items, continuous monitoring artifacts).
- Capacity Planning & Optimization: Build and maintain service-level capacity models, App Services, databases, storage, messaging, networking, Key Vault/HSM, and other Azure/PaaS components.
- Secure System/Service Acquisition & Region Restrictions: Ensure external services supporting capacity (e.g., third-party telemetry or scaling tools) conform to required requirements with documented oversight and continuous monitoring
- Resilience, DR, and Performance Engineering: Perform criticality analysis to prioritize capacity for high-critical components; align hardening, monitoring, backup/DR, and buffer policies to criticality tiers.
- Metrics & Reporting: Define and publish capacity KPIs:utilization, saturation, headroom %, runway weeks, scaling efficacy, quota consumption, DR readiness, cost-to-performance efficiency.
- Bachelor's degree in computer science or related discipline; advanced degree preferred.
- 10–12+ years in infrastructure capacity/performance engineering across compute, storage, network, and platform services; financial services experience is a plus.
- Demonstrated experience operating in regulated environments; familiarity with FedRAMP High concepts and evidence requirements.
- Strong data analysis skills; capable of translating telemetry and forecasts into clear decisions and stakeholder communications.
- Experience coordinating cross-functional engineering teams and aligning delivery across multiple platforms and tools.
- Familiarity with Azure services and concepts (e.g., Entra ID, managed identities, Azure SQL/MI, storage, networking, policies, RBAC) from a PM perspective.
- Azure capacity ecosystem: Monitor/Log Analytics/Metrics, Advisor, Cost Management, Reservations/Savings Plans, quotas/limits management.
- Compute/container scaling:AKS, VMSS, App Service; HPA/VPA, autoscaling policies; performance testing (k6/JMeter); observability (Prometheus/Grafana).
- Storage and database performance: tiering, IOPS/throughput planning, caching, indexing, and connection management.
- Networking and security capacity:Azure Firewall, NSGs, private endpoints, Bastion; throughput/latency planning and allow-listing discipline.
- Cryptography services: Key Vault, managed HSM; FIPS-validated modules; key lifecycle capacity considerations.
- IaC and config management: Terraform/Bicep/ARM; Ansible/Chef; integration with gated CI/CD.
- Governance: Azure Policy/Blueprints/Initiatives for configuration baselines and region restrictions; SSP and evidence artifact production.
- Financial services experience
- Building capacity models for multi-region architectures with strict U.S.-only constraints.
- DR planning and execution with validated failover capacity and documented evidence.
- POA&M management and continuous monitoring submissions in a FedRAMP context.
- Collaboration with SREs on SLI/SLOs, error budgets, and reliability patterns.
Principal Azure Capacity Manager · Veterans Sourcing Group