Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
N

Machine Learning Engineer

NimrodCareers
๐Ÿ‡ฐ๐Ÿ‡ช Kenya
On-site
4 weeks ago
$1,100 โ€“ $1,545 / month
  • Machine Learning
  • Devops
  • MLOps
  • CI/CD
  • IaC
  • Terraform
  • CloudFormation
  • Microservices
  • Kubernetes
  • Docker
  • Helm
  • MLflow
  • Kubeflow
  • Vertex AI
  • Azure ML
  • GitHub Actions
  • GitLab CI/CD
  • Jenkins
  • Azure DevOps
  • AWS
  • Azure
  • GCP
  • IAM
  • Secrets Management
  • Risk Management
  • Change Management
  • AI
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Purpose

To design, build, automate and operate the technology infrastructure, platforms and deployment pipelines required to develop, train, deploy, monitor and continuously improve machine learning models at scale.

The role will bridge Data Science, Machine Learning Engineering, Software Engineering, DevOps and Cloud Infrastructure, ensuring that machine learning solutions move reliably from experimentation into secure, scalable and production-ready environments.

The role will establish and maintain MLOps capabilities covering model development, experimentation, data and model versioning, automated testing, continuous integration and delivery, continuous training, model serving, observability, performance monitoring and infrastructure automation.

Key Responsibilities

MLOps Platform & Infrastructure

  • Design, build and maintain scalable infrastructure for machine learning workloads across cloud, on-premise or hybrid environments.

  • Develop and maintain enterprise MLOps platforms supporting the full machine learning lifecycle.

  • Provision and manage compute, storage, networking and GPU/accelerator resources required for ML workloads.

  • Design reusable infrastructure and deployment patterns for machine learning teams.

  • Implement Infrastructure as Code using tools such as Terraform, CloudFormation or equivalent technologies.

  • Establish development, testing, staging and production environments for ML workloads.

  • Optimise infrastructure utilisation, scalability, reliability and cost.

ML Pipeline Engineering

  • Design and automate end-to-end machine learning pipelines covering data preparation, feature engineering, training, evaluation, validation and deployment.

  • Implement Continuous Integration, Continuous Delivery and Continuous Training (CI/CD/CT) pipelines for ML systems.

  • Automate model retraining based on scheduled events, new data, performance degradation or defined business triggers.

  • Build reusable pipeline components and templates for Data Scientists and ML Engineers.

  • Implement automated data, model, code and pipeline validation.

  • Support reproducibility and traceability of ML experiments and production models.

Model Deployment & Serving

  • Develop scalable and reliable model-serving infrastructure for real-time, batch and event-driven inference.

  • Deploy models through APIs, microservices, containers or other appropriate serving architectures.

  • Implement model versioning, model promotion and rollback mechanisms.

  • Support blue-green, canary and A/B deployment strategies where appropriate.

  • Optimise inference latency, throughput and resource utilisation.

  • Ensure production models meet defined availability, performance and scalability requirements.

Containerisation & Kubernetes

  • Containerise ML workloads using Docker or equivalent technologies.

  • Deploy and manage ML workloads on Kubernetes and/or managed container platforms.

  • Develop Helm charts, Kubernetes manifests and deployment templates.

  • Configure autoscaling, resource allocation, service discovery and workload scheduling.

  • Support GPU-enabled workloads where required.

  • Implement secure and reliable container deployment practices.

Model & Experiment Management

  • Implement and manage model registries and model lifecycle processes.

  • Support experiment tracking, model versioning and ML metadata management.

  • Establish processes for tracking datasets, features, models, parameters, metrics and deployment history.

  • Support tools such as MLflow, Kubeflow, Vertex AI, SageMaker, Azure ML or equivalent platforms.

  • Ensure models can be reproduced, audited and traced from development through production.

  • Modern MLOps platforms typically incorporate source control, model registries, feature stores, metadata management and pipeline orchestration.

Monitoring & Observability

  • Develop comprehensive monitoring for ML infrastructure, pipelines and deployed models.

  • Monitor model accuracy, latency, throughput, resource utilisation and availability.

  • Implement monitoring for data drift, concept drift, model degradation and anomalous behaviour.

  • Establish alerting mechanisms for failed pipelines, infrastructure issues and model performance degradation.

  • Integrate ML monitoring with enterprise logging and observability platforms.

  • Develop dashboards and operational metrics for ML services.

  • Continuous monitoring is critical because ML models can degrade as production data and business conditions change.

DevOps & Automation

  • Integrate ML development into enterprise DevOps practices.

  • Develop CI/CD pipelines using platforms such as GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps or equivalent.

  • Automate build, test, packaging and deployment processes.

  • Implement automated unit, integration, regression, performance and model validation testing.

  • Establish version control standards for code, configuration, pipelines and infrastructure.

  • Automate infrastructure provisioning and environment configuration.

Cloud & Platform Engineering

  • Design and manage ML infrastructure across AWS, Microsoft Azure, Google Cloud Platform or other cloud environments.

  • Implement secure cloud architectures for machine learning workloads.

  • Configure cloud compute, storage, networking, IAM, secrets management and monitoring.

  • Optimise cloud resource consumption and ML infrastructure costs.

  • Support hybrid-cloud and multi-cloud ML architectures where required.

Security, Risk & Governance

  • Implement security controls across ML infrastructure, pipelines, APIs and model-serving environments.

  • Apply identity and access management principles to ML platforms and workloads.

  • Secure model artefacts, datasets, credentials, APIs and infrastructure.

  • Implement secrets management, encryption and appropriate access controls.

  • Support model governance, auditability and regulatory requirements.

  • Ensure ML infrastructure complies with organisational technology, cybersecurity, data protection and risk-management standards.

  • Maintain appropriate audit trails for model development, approval, deployment and change management.

Collaboration with Data & Technology Teams

  • Partner with Data Scientists to productionise machine learning models.

  • Work closely with Data Engineers to ensure reliable and scalable data pipelines.

  • Collaborate with Software Engineers to integrate ML services into enterprise applications.

  • Partner with DevOps, Cloud, Cybersecurity and Infrastructure teams.

  • Provide technical guidance and MLOps standards to ML development teams.

  • Support the establishment of engineering standards and best practices across the organisation.

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Technology, Data Science, Artificial Intelligence, Engineering or a related discipline.

  • A Master's degree in a relevant technical field is an added advantage.

  • Relevant professional certifications in cloud, DevOps, Kubernetes, machine learning or data engineering are desirable.

Professional Experience

  • 5 years of experience in software engineering, DevOps, cloud engineering, data engineering, ML engineering or related technical disciplines.

  • At least 3 years' practical experience supporting machine learning systems in production or implementing MLOps capabilities.

  • Demonstrated experience designing and operating CI/CD pipelines.

  • Experience deploying and operating machine learning models at scale.

  • Experience working with cloud infrastructure and containerised workloads.

  • Experience with Kubernetes and Infrastructure as Code is highly desirable.

  • Experience working in regulated environments such as banking, financial services, insurance or telecommunications is an advantage.

Machine Learning Engineer ยท NimrodCareers

Auto apply with Likeremote