
Director of DevOps & SRE
InvestorFlow
🇬🇧 United Kingdom
Hybrid
Manager or above
3 weeks ago
- CRM
- Salesforce
- Devops
- CI/CD
- Secrets Management
- Azure
- AWS
- Snowflake
- GitHub Actions
- Vulnerability Management
- IaC
- Terraform
- Auth0
- Incident Response
- Disaster Recovery
- Change Management
- Grafana
- AI
- triage
- WAF
- Azure DevOps
- Docker
- Kubernetes
- SAML
- OAuth
- Python
- Bash
- PowerShell
- TCP/IP
- DNS
- Load Balancing
- Prometheus
- Penetration Testing
- Cloudflare
- Redis
- Equity
3 weeks ago
You Will:
- Strategy and roadmap. The DevOps/SRE strategy for the InvestorFlow platform, aligned to measurable outcomes across reliability, delivery throughput, and cost.
- The team. Leading, hiring, and developing technical leads, senior and mid-level engineers across DevOps and SRE, and setting priorities across infrastructure, CI/CD, security, and production support.
- Cloud platform ownership. Reliability, performance, and cost-efficiency across DEV, QA, STG, and PRD -compute, data, secrets management, caching, and networking/edge security. Azure is our primary platform, with a supporting AWS and Snowflake footprint.
- CI/CD and release engineering. GitHub Actions end to end, multi-region deployment automation, advanced deployment strategies (blue/green, canary, rolling) and rollback, plus security scanning, compliance checks, and vulnerability management embedded in the pipeline.
- Infrastructure as code. Terraform practice and standards that standardise provisioning and reduce manual, ticket-driven infrastructure work.
- Automation-first operating standard. Setting and holding the expectation that repeatable work is codified rather than performed, across provisioning, releases, access, remediation, and reporting, and driving developer enablement through self-service tooling that increases engineering capacity while maintaining controls.
- Identity and access. Enterprise directory services, role-based and privileged access controls, Auth0 machine-to-machine credentials, and least-privilege access across engineering and QA.
- Incident response and resilience. Escalation paths, root-cause analysis and postmortems, production readiness reviews, automated failover, and our disaster recovery and RTO/RPO commitments. You are the escalation point for high-severity incidents and major client environment changes, ensuring appropriate change management and CAB governance, with occasional off-hours support for the teams.
- Observability and service levels. SLO/SLA and monitoring strategy across Grafana and our wider observability tooling, building alert coverage that is meaningful rather than noisy.
- AI in engineering operations. The operational rollout of AI-assisted engineering and support tooling: automated triage, agent-based workflows, and developer-facing AI capability.
- Architectural review and technical governance. Reviewing significant infrastructure and platform change with a security- and compliance-first lens, ensuring designs meet our obligations by default rather than by exception, and owning the shared-responsibility operating model between DevOps/SRE and product engineering.
- Vendor and cost management. Relationships across cloud, security, CI/CD, and observability platforms, including spend and renewal negotiation.
You Have:
- 8+ years in DevOps, SRE, or infrastructure engineering, including 3+ years leading and developing engineering teams.
- Deep hands-on Microsoft Azure experience at production scale: compute, data, secrets, identity services, networking and WAF, and cost/capacity management. Azure depth is essential for this role.
- Strong CI/CD and release engineering background (GitHub Actions and/or Azure DevOps Pipelines) and infrastructure as code with Terraform, including module development and state management.
- Production support and incident response experience for a multi-tenant SaaS platform.
- Containerisation and orchestration in production (Docker, Kubernetes).
- Solid grounding in identity and access management (role-based and privileged access, SSO/SAML, OAuth/Auth0) and secrets management practices.
- A genuine automation-first instinct, backed by strong scripting (Python, Bash, PowerShell or similar) and a track record of replacing manual process with reliable, supportable tooling.
- Networking fundamentals (TCP/IP, DNS, load balancing) and the ability to partner effectively with network and security teams.
- Observability experience with Grafana, Prometheus, or equivalent, and a track record of building signal rather than noise.
- Comfort operating in a security and compliance-conscious environment: financial services, private markets, or another regulated industry, including audit, penetration testing, and data retention requirements.
- Excellent written and verbal communication, able to translate infrastructure risk and trade-offs for engineering peers, executives, and clients.
- The judgement to know when to go deep technically and when to lead through others.
Nice to have: AWS exposure alongside Azure, Snowflake or another cloud data platform, CDN and edge platforms (Cloudflare or similar), Redis, event-driven or message-oriented middleware (Kafka, Azure Service Bus), and prior experience in private equity, private credit, real assets, or capital-markets SaaS.
Director of DevOps & SRE · InvestorFlow