Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
N

Distributed Systems & Platform Engineer

NimrodCareers
🇰🇪 Kenya
On-site
1 month ago
$1,100 – $1,545 / month
  • AI
  • Machine Learning
  • Kubernetes
  • AI/ML
  • IaC
  • Canary Releases
  • CI/CD
  • CUDA
  • Slurm
  • Ray
  • InfiniBand
  • Terraform
  • Pulumi
  • GitOps
  • Node.js
  • AWS
  • Azure
  • GCP
  • Load Balancing
  • Incident Response
  • RBAC
  • Secrets Management
  • Devops
  • Incident Management
  • C++
  • Rust
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Purpose

To design, build, and operate the distributed computing and platform infrastructure that powers the organization’s AI, machine learning, generative AI, and data-intensive workloads.

The role will be responsible for developing highly scalable, reliable, secure and high-performance platforms for distributed model training, inference, experimentation and production AI services. The position will work across Kubernetes, cloud and/or bare-metal infrastructure, GPU compute, distributed systems, storage, networking, orchestration and platform automation.

The role combines deep software engineering with infrastructure and platform engineering, with a strong focus on creating self-service platforms that enable AI/ML engineering teams to deploy and operate workloads efficiently at scale. Current AI infrastructure roles increasingly emphasize Kubernetes control-plane expertise, GPU scheduling, distributed training/inference, multi-tenancy, observability and infrastructure-as-code.

Job Responsibilities

Distributed Systems Engineering

  • Design and develop highly available, fault-tolerant and horizontally scalable distributed systems.

  • Build services capable of operating reliably across multiple nodes, clusters, availability zones and cloud environments.

  • Design mechanisms for distributed coordination, state management, caching, queuing, scheduling and workload orchestration.

  • Develop systems capable of graceful degradation, automated recovery and failure isolation.

  • Identify and resolve performance bottlenecks across compute, storage, networking and application layers.

  • Design APIs, services and control planes that abstract complex infrastructure capabilities for internal users.

  • Apply appropriate distributed-systems patterns including replication, sharding, partitioning, consensus, idempotency and eventual consistency.

AI/ML Infrastructure Platform

  • Build and maintain the foundational platform supporting AI/ML model training, fine-tuning, experimentation and inference.

  • Develop reusable infrastructure and services for AI workloads across development, testing and production environments.

  • Enable distributed training and inference across GPU clusters.

  • Develop mechanisms for job submission, scheduling, prioritisation, retries, checkpointing, logging and artifact management.

  • Support model-serving platforms with scalable routing, autoscaling, deployment, versioning, canary releases and rollback capabilities.

  • Build self-service capabilities that allow AI/ML engineers to provision and manage compute resources without extensive infrastructure intervention.

  • Integrate AI workloads with model registries, experiment tracking, data platforms and CI/CD pipelines.

Kubernetes & Container Platforms

  • Design, deploy and operate production-grade Kubernetes platforms for AI workloads.

  • Develop Kubernetes operators, controllers, Custom Resource Definitions (CRDs) and automation.

  • Understand and optimise Kubernetes control-plane components including API server, scheduler, controller manager, kubelet and etcd.

  • Design multi-cluster and multi-tenant Kubernetes architectures.

  • Implement workload scheduling, autoscaling, resource quotas and workload isolation.

  • Manage cluster lifecycle, upgrades, configuration, capacity and resilience.

  • Optimise Kubernetes for GPU-intensive and distributed workloads.

  • Implement containerization standards and platform engineering best practices.

GPU & High-Performance Computing Infrastructure

  • Design and operate GPU-enabled infrastructure supporting AI training and inference.

  • Implement efficient GPU allocation, scheduling, sharing and utilisation strategies.

  • Work with GPU technologies such as NVIDIA CUDA, GPU Operator, MIG, MPS and NCCL where applicable.

  • Support GPU-aware scheduling, topology-aware workload placement and distributed GPU workloads.

  • Optimise GPU utilisation, throughput, latency and cost.

  • Work with cluster schedulers such as Kubernetes, Slurm, Ray, Kueue or Volcano where appropriate.

  • Diagnose GPU, compute, memory, communication and networking bottlenecks.

  • Support high-performance interconnects and technologies such as RDMA, InfiniBand and RoCE where applicable.

Platform Engineering & Developer Experience

  • Build internal developer platforms, APIs, CLIs, SDKs and automation tools.

  • Establish "golden paths" for AI/ML engineers to deploy, test and operate workloads.

  • Provide reusable infrastructure components and platform services.

  • Reduce infrastructure complexity and manual operational effort through automation.

  • Develop self-service provisioning and deployment capabilities.

  • Create platform documentation, technical standards and reusable engineering patterns.

  • Treat internal engineering teams as platform customers and continuously improve platform usability and adoption.

Infrastructure as Code & Automation

  • Develop and maintain Infrastructure as Code using Terraform, Pulumi or equivalent technologies.

  • Automate infrastructure provisioning, configuration, deployment and lifecycle management.

  • Implement GitOps practices for infrastructure and application deployment.

  • Develop automated remediation and self-healing capabilities.

  • Automate cluster provisioning, node management, GPU configuration and workload deployment.

  • Build repeatable infrastructure patterns across cloud, on-premise and hybrid environments.

Cloud & Infrastructure Architecture

  • Design scalable infrastructure across AWS, Microsoft Azure, Google Cloud and/or private infrastructure.

  • Develop hybrid and multi-cloud infrastructure patterns where required.

  • Design compute, storage and networking architectures optimised for AI workloads.

  • Implement capacity planning and resource allocation strategies.

  • Support cloud bursting and dynamic scaling for high-demand AI workloads.

  • Optimise infrastructure cost without compromising performance, security or reliability.

Storage & Data Infrastructure

  • Design high-throughput storage solutions for AI training datasets, model artefacts, checkpoints and logs.

  • Work with object, block and distributed file storage systems.

  • Implement caching and data-locality strategies to reduce training and inference bottlenecks.

  • Support high-performance distributed storage environments.

  • Implement appropriate data lifecycle, retention and quota-management mechanisms.

  • Ensure storage systems meet required performance, availability and resilience objectives.

Networking

  • Design and optimise networking for distributed AI workloads.

  • Work with Kubernetes networking, CNI, service discovery, load balancing and ingress.

  • Support high-throughput, low-latency communication between distributed compute nodes.

  • Troubleshoot network performance and reliability issues.

  • Support RDMA-enabled networking where required for distributed GPU workloads.

  • Implement appropriate network segmentation, security policies and multi-tenant isolation.

Reliability, Observability & Performance

  • Define and implement SLIs, SLOs and operational metrics for platform services.

  • Build comprehensive monitoring, logging and distributed tracing.

  • Monitor CPU, memory, GPU, storage, network and application-level performance.

  • Develop dashboards, alerting and automated incident detection.

  • Conduct performance benchmarking, capacity testing and scalability assessments.

  • Participate in incident response and root-cause analysis.

  • Build systems for fault detection, automated recovery and resilience testing.

Security & Multi-Tenancy

  • Design secure, multi-tenant AI infrastructure platforms.

  • Implement RBAC, namespaces, quotas, secrets management and workload isolation.

  • Apply security controls across Kubernetes, cloud infrastructure, containers and APIs.

  • Support identity and access management across platform services.

  • Implement audit logging and appropriate security monitoring.

  • Ensure infrastructure complies with organisational security and data-protection requirements.

Engineering Leadership & Collaboration

  • Work closely with AI/ML engineers, data scientists, software engineers, DevOps/SRE, cybersecurity and enterprise architects.

  • Translate AI/ML requirements into scalable infrastructure and platform solutions.

  • Lead technical design discussions and architecture reviews.

  • Conduct code reviews and promote engineering quality standards.

  • Mentor engineers and contribute to technical capability development.

  • Evaluate emerging infrastructure and AI technologies and recommend appropriate adoption.

  • Contribute to engineering standards, technical roadmaps and platform strategy.

Key Deliverables / Expected Outcomes

The role holder will be expected to:

  • Deliver a scalable and reliable AI infrastructure platform.

  • Improve GPU and compute-resource utilisation.

  • Reduce deployment and provisioning times for AI workloads.

  • Enable self-service infrastructure capabilities for AI/ML engineering teams.

  • Improve reliability and availability of distributed AI workloads.

  • Reduce infrastructure-related operational toil through automation.

  • Improve performance and efficiency of distributed training and inference.

  • Establish effective monitoring, observability and incident-management capabilities.

  • Maintain secure and appropriately isolated multi-tenant environments.

  • Support rapid and reliable deployment of AI models into production.

Job Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Technology or a related discipline.

  • A Master's degree in Computer Science, Distributed Systems, AI/ML, Cloud Computing or a related field would be advantageous.

  • Relevant professional certifications in cloud, Kubernetes or infrastructure engineering are desirable.

Professional Experience

  • 3-5 years of software engineering, distributed systems, infrastructure, cloud or platform engineering experience, depending on seniority.

  • Demonstrable experience designing and operating production-scale distributed systems.

  • Strong production experience with Kubernetes and containerised environments.

  • Experience building infrastructure platforms for AI/ML workloads is highly desirable.

  • Experience with GPU infrastructure, distributed training or model serving is strongly preferred.

  • Experience with cloud and/or large-scale bare-metal infrastructure.

  • Experience developing production software in C++, Rust or equivalent systems-oriented languages.

  • Experience with Infrastructure as Code and CI/CD automation.

  • Experience troubleshooting complex production infrastructure across compute, storage and networking.

Distributed Systems & Platform Engineer · NimrodCareers

Auto apply with Likeremote