Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
AS

AI Platform Engineering Specialist

Axelon Services Corporation
๐Ÿ‡จ๐Ÿ‡ฆ Canada
On-site
Senior
6 days ago
  • AI
  • Azure
  • AWS
  • Python
  • FastAPI
  • Flask
  • Azure AI
  • Foundry
  • Azure OpenAI
  • AWS Bedrock
  • AWS IAM
  • PostgreSQL
  • Kubernetes
  • AKS
  • EKS
  • Helm
  • GitOps
  • Terraform
  • CI/CD
  • Jenkins
  • GitHub Actions
  • Prometheus
  • Grafana
  • Loki
  • Snowflake
  • OIDC
  • Key Vault
  • Redis
  • Azure Monitor
  • IAM
  • Bedrock
  • VPC
  • CloudWatch
  • IaC
  • Bicep
  • SQL
  • Data Modeling
  • API Gateway
  • Server-Sent Events
  • OpenTelemetry
  • Valkey
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Experience Level: Level 3 (senior): minimum 5 years
Duration: 12 Months contract position
Location: Montreal
Work Mode: Onsite (Day 1 onboarding onsite/in office presence 3x/week)

Responsibilities:

  • Design, build, and operate the AI Gateway's Azure and AWS deployments, transitioning from proof of concept to production.
  • Develop and extend Python services (FastAPI / Flask) for the Gateway's inference, onboarding, and administrative APIs.
  • Integrate new model providers and model families, including Azure AI Foundry / Azure OpenAI and AWS Bedrock, covering request signing, streaming responses, failover, and quota handling.
  • Implement cloud-native authentication and secrets handling โ€” Entra ID with Managed Identity and workload federation, AWS IAM roles, and STS โ€” aiming to eliminate stored credentials.
  • Build and evolve the entitlement and authorization data layer across SQL Server and PostgreSQL, including schema changes, migrations, and data-correctness controls.
  • Manage platform controls for governance: rate limiting, token accounting, content guardrails, audit logging, and chargeback reporting.
  • Deploy and run services on Kubernetes (on-premises, AKS, and EKS) using Helm, GitOps, and Terraform, and maintain CI/CD pipelines (Jenkins, GitHub Actions).
  • Develop observability tools to track requests post-factum โ€” metrics, logs, and dashboards across Prometheus, Grafana, Loki, and Snowflake.
  • Collaborate with cloud platform, network, and security teams on connectivity, egress policy, network controls, and architecture review, and provide necessary evidence for reviews.
  • Support production: participate in on-call, investigate incidents, and implement fixes and hardening into the code.
  • Write tests and documentation as part of delivery and review peers' changes.

Requirements:

  • Strong, production-grade Python, including a web framework โ€” FastAPI or Flask โ€” and a real testing discipline.
  • Hands-on experience with Kubernetes: deploying, configuring, and troubleshooting workloads.
  • Practical understanding of OIDC / OAuth 2.0: token validation, JWKS, client-credentials flows, claim, and audience handling.
  • Microsoft Azure experience in at least three of the following: AKS, Entra ID, Azure OpenAI or Azure AI Foundry, Key Vault, Azure Database for PostgreSQL, Azure Cache for Redis, Azure Monitor.
  • Amazon Web Services experience in at least three of the following: IAM and STS / assume-role, SigV4 request signing, Bedrock, EKS, VPC endpoints and private networking, Secrets Manager, CloudWatch.
  • Experience with Infrastructure as code โ€” Terraform, Bicep, or CDK โ€” and CI/CD with Jenkins or GitHub Actions.
  • Proficiency in SQL and relational data modeling, including schema migrations.
  • Clear written and verbal communication skills, and the ability to work directly with security, network, and platform teams.

Preferred Skills:

  • Experience building or operating an API gateway, reverse proxy, or multi-tenant platform.
  • LLM platform engineering specifics: streaming and server-sent events, token accounting, prompt and response guardrails, model evaluation.
  • Experience with Kafka and Snowflake for audit and consumption data pipelines.
  • Observability depth: Prometheus and PromQL, Grafana, Loki, OpenTelemetry.
  • Advanced Redis or Valkey use beyond basic caching โ€” counters, TTLs, distributed rate-limiter semantics.
  • Experience delivering in a regulated enterprise environment with corporate proxies, private networking, and strict change control.

This role is for an existing vacancy.

AI Platform Engineering Specialist ยท Axelon Services Corporation

Auto apply with Likeremote