
Senior Release & Reliability Engineer
- Disaster Recovery
- Windows
- Configuration Management
- Devops
- Python
- Linux
- CI/CD
- Jenkins
- TeamCity
- GitLab CI
- Git
- Maven
- Gradle
- npm
- Nexus
- Artifactory
- YAML
- Elasticsearch
- Logstash
- Kibana
- Grafana
- iOS
- Android
- IaC
- Ansible
- SonarQube
- ServiceNow
- Jira
- Pension
We are looking for a Senior Release & Reliability Engineer to deliver and automate the release and deployment of our backend services: approval, deploy, verification and rollback.
More than sixty backend services and web applications are released and restarted across production regions in North America, Europe and Asia, behind our web portal, mobile apps and desktop trading platform.
The role also owns disaster recovery for that estate: standby regions that stay current, backups that restore and failover that has been rehearsed.
You will
The release and deployment team, two engineers today, working with every development team in the group as well as product managers, compliance and the infrastructure team.
Responsibilities
- Deliver releases and deployments for backend services and web applications, from test through production, to whole regions or single hosts
- Run the release request lifecycle with the requesting developer: approvals, rollback path, verification, change records
- Execute rollbacks under time pressure, and verify the rolled-back state
- Set up disaster recovery: standby regions, replication, backup and restore, failover and failback
- Run DR drills and regional failover tests, and keep recovery procedures proven
- Verify releases with telemetry rather than by eye: dashboards, log queries, error and latency checks, smoke tests
- Onboard new services into the release system, including clustered multi-instance topologies
- Automate manual release work: pipelines, deployment and configuration templates, validation, verification harnesses, restart jobs
- Bring infrastructure, application and release pipelines under one versioned, testable template language, and retire per-project scripts
- Help teams meet readiness standards before new traffic reaches production: monitoring coverage, load-test evidence, rate limiting, staggered restarts
- Own build and release infrastructure and access: build jobs, artifact promotion and versioning, release permissions
- Provide L2 and L3 support for the release path: off-hours windows, weekend coverage, emergency releases, incidents
- Mentor engineers on release practice, and keep release and deployment procedures current
Requirements and skills
- 2+ years in release engineering, configuration management, DevOps, SRE or platform engineering, including production releases of server-side services
- Rigor about change control: every release documented and reconstructable afterwards, never shipped without a tested rollback path
- Disaster recovery: standby environments, replication, backup and restore, and failover you have actually run
- Strong Python and shell, enough to own and extend release tooling and automate operations against tool APIs
- Deep Linux experience, ideally Red Hat, including diagnosing a failing service on an unfamiliar host
- CI/CD and artifact management: Jenkins, TeamCity or GitLab CI, Git, Maven, Gradle or npm, Nexus or Artifactory
- Configuration as code: YAML models, environment, region and host overrides, everything in version control
- Observability and log analysis: Elasticsearch, Logstash and Kibana, Grafana or equivalent, plus scripting and regular expressions for raw logs
- Clear written communication, and comfort working across teams in several countries and time zones
- BSc/BA in Computer Science, Engineering or a relevant field
Good to have
- 5+ years of the above, in banking, brokerage or another regulated production environment
- Multi-region disaster recovery at scale, including regional failover of stateful services
- Reusable CI/CD and deployment pipeline templates across a large application estate
- Automating build and release pipelines for customer facing apps (e.g. iOS, Android, Desktop): code signing, beta distribution and store submission
- Cloud migration of on-premises release and deployment processes, infrastructure as code, Ansible, containers or orchestration
- Internal developer platforms and self-service release tooling
- Code quality gates such as SonarQube, job scheduling and change tracking such as ServiceNow or Jira
- Exposure to market data, order routing or brokerage systems
Β
Benefits
- Competitive salary, annual performance-based bonus, and stock grant awards
- 401(k) retirement plan with competitive company match
- Excellent health and wellness benefits, including medical, dental, and vision benefits. 100% employer-paid medical premiums, with generous employer contributions to dental & vision plans as well.
- Wellness screening and assessments, health coaches, and counseling services through an Employee Assistance Program (EAP)
- Generous paid parental leave (up to 16 weeks paid parental leave for eligible employees)
- Company-paid basic life insurance, accidental death & dismemberment (AD&D), and short- and long-term disability coverage
- Flexible Spending Accounts (Healthcare, Dependent Care, and Commuter FSAs)
- Quarterly fitness stipend to offset costs associated with traditional gym and fitness memberships or fees
- Education reimbursement and professional development opportunities
- Legal services, telehealth access, and voluntary insurance options
- Backup child and adult care support through Care.com
- Daily lunch allowance and fully stocked kitchen with healthy breakfast and snack options has context menu
Senior Release & Reliability Engineer Β· Interactive Brokers