AL
OBSERVABILITY ENGINEER
Artech LLC
🇨🇦 Canada
On-site
Manager or above
3 days ago
- Kubernetes
- AI/ML
- Prometheus
- Grafana
- Thanos
- Loki
- GitOps
- IaC
- Splunk
- Fluentd
- Fluent Bit
3 days ago
Location: Halifax, NS - Hybrid (4 Days WFO)
Duration: 6-12 months
Pay Range: C$53 INCÂ
Role Descriptions: OBSERVABILITY ENGINEEREnterprise Kubernetes Platform | Financial Services========================================================================ABOUT THE ROLEWe are seeking an experienced Observability Engineer to join our EnterpriseKubernetes Platform team at a leading financial services organization. Youllown the complete observability stack across 50+ production Kubernetes clusters|providing metrics| logging| tracing| and alerting capabilities that ensureexceptional reliability and performance for mission-critical applications.This role combines deep technical expertise in modern observability tools withemerging AI/ML capabilities to build intelligent monitoring solutions|predictive alerting| and self-healing infrastructure.========================================================================WHAT YOULL DO========================================================================OBSERVABILITY STACK OWNERSHIP-------------------------------------------------------------------------------- Design| deploy| and maintain enterprise-scale observability infrastructureincluding Prometheus| Grafana| Thanos| Loki| and modern collection agents Manage observability deployments using GitOps principles and infrastructureas code Implement long-term metrics storage solutions with cloud object storage Maintain and upgrade observability components across development| QA| UAT|production| and DR environments Configure distributed observability architecture spanning multiple datacenters and cloud providersMETRICS & MONITORING-------------------------------------------------------------------------------- Design and implement Prometheus monitoring strategies for Kubernetesinfrastructure and containerized applications Create ServiceMonitors| PodMonitors for automated metrics collection Develop rules for intelligent alerting with minimal false positives Configure multi-cluster metrics federation and aggregation Optimize metrics cardinality| storage e_iciency| and query performance Implement recording rules for pre-aggregated metrics and SLI calculationsDASHBOARDS & VISUALIZATION-------------------------------------------------------------------------------- Build comprehensive Grafana dashboards for infrastructure health| applicationperformance| and business metrics Create reusable dashboard templates and libraries for development teams Implement dashboard-as-code Configure multiple datasources Design executive dashboards with SLO/SLI tracking and business KPIs Implement role-based access control and multi-tenancy in GrafanaLOGGING INFRASTRUCTURE-------------------------------------------------------------------------------- Deploy and manage centralized logging solutions (Loki| ELK| Splunk| orsimilar) Configure log collection agents (Promtail| Fluentd| FluentBit| Vector| etc.) Design log retention policies balancing cost| compliance| and operationalneeds Create LogQL/Lucene queries and log-base
Custom Fields:
Name: Assets
Value: None
Name: Parent
Value: IS-BFSI-Canada-2-Parent
Name: Type of Asset
Value: None
Name: Child
Value: IS-BFSI-Canada-2-Group1
Name: Work Assignment Location
Value: Hybrid
Name: Requisition ID
Value: 10969790
Name: Business Vertical
Value: CL-BFSI Americas 2
Name: Work Location CDF
Value: ~HALIFAX, NOVA SCOTIA~
OBSERVABILITY ENGINEER · Artech LLC