Distributed Systems & Platform Engineer
- AI
- Machine Learning
- Kubernetes
- AI/ML
- IaC
- Canary Releases
- CI/CD
- CUDA
- Slurm
- Ray
- InfiniBand
- Terraform
- Pulumi
- GitOps
- Node.js
- AWS
- Azure
- GCP
- Load Balancing
- Incident Response
- RBAC
- Secrets Management
- Devops
- Incident Management
- C++
- Rust
Job Purpose
To design, build, and operate the distributed computing and platform infrastructure that powers the organization’s AI, machine learning, generative AI, and data-intensive workloads.
The role will be responsible for developing highly scalable, reliable, secure and high-performance platforms for distributed model training, inference, experimentation and production AI services. The position will work across Kubernetes, cloud and/or bare-metal infrastructure, GPU compute, distributed systems, storage, networking, orchestration and platform automation.
The role combines deep software engineering with infrastructure and platform engineering, with a strong focus on creating self-service platforms that enable AI/ML engineering teams to deploy and operate workloads efficiently at scale. Current AI infrastructure roles increasingly emphasize Kubernetes control-plane expertise, GPU scheduling, distributed training/inference, multi-tenancy, observability and infrastructure-as-code.
Job Responsibilities
Distributed Systems Engineering
Design and develop highly available, fault-tolerant and horizontally scalable distributed systems.
Build services capable of operating reliably across multiple nodes, clusters, availability zones and cloud environments.
Design mechanisms for distributed coordination, state management, caching, queuing, scheduling and workload orchestration.
Develop systems capable of graceful degradation, automated recovery and failure isolation.
Identify and resolve performance bottlenecks across compute, storage, networking and application layers.
Design APIs, services and control planes that abstract complex infrastructure capabilities for internal users.
Apply appropriate distributed-systems patterns including replication, sharding, partitioning, consensus, idempotency and eventual consistency.
AI/ML Infrastructure Platform
Build and maintain the foundational platform supporting AI/ML model training, fine-tuning, experimentation and inference.
Develop reusable infrastructure and services for AI workloads across development, testing and production environments.
Enable distributed training and inference across GPU clusters.
Develop mechanisms for job submission, scheduling, prioritisation, retries, checkpointing, logging and artifact management.
Support model-serving platforms with scalable routing, autoscaling, deployment, versioning, canary releases and rollback capabilities.
Build self-service capabilities that allow AI/ML engineers to provision and manage compute resources without extensive infrastructure intervention.
Integrate AI workloads with model registries, experiment tracking, data platforms and CI/CD pipelines.
Kubernetes & Container Platforms
Design, deploy and operate production-grade Kubernetes platforms for AI workloads.
Develop Kubernetes operators, controllers, Custom Resource Definitions (CRDs) and automation.
Understand and optimise Kubernetes control-plane components including API server, scheduler, controller manager, kubelet and etcd.
Design multi-cluster and multi-tenant Kubernetes architectures.
Implement workload scheduling, autoscaling, resource quotas and workload isolation.
Manage cluster lifecycle, upgrades, configuration, capacity and resilience.
Optimise Kubernetes for GPU-intensive and distributed workloads.
Implement containerization standards and platform engineering best practices.
GPU & High-Performance Computing Infrastructure
Design and operate GPU-enabled infrastructure supporting AI training and inference.
Implement efficient GPU allocation, scheduling, sharing and utilisation strategies.
Work with GPU technologies such as NVIDIA CUDA, GPU Operator, MIG, MPS and NCCL where applicable.
Support GPU-aware scheduling, topology-aware workload placement and distributed GPU workloads.
Optimise GPU utilisation, throughput, latency and cost.
Work with cluster schedulers such as Kubernetes, Slurm, Ray, Kueue or Volcano where appropriate.
Diagnose GPU, compute, memory, communication and networking bottlenecks.
Support high-performance interconnects and technologies such as RDMA, InfiniBand and RoCE where applicable.
Platform Engineering & Developer Experience
Build internal developer platforms, APIs, CLIs, SDKs and automation tools.
Establish "golden paths" for AI/ML engineers to deploy, test and operate workloads.
Provide reusable infrastructure components and platform services.
Reduce infrastructure complexity and manual operational effort through automation.
Develop self-service provisioning and deployment capabilities.
Create platform documentation, technical standards and reusable engineering patterns.
Treat internal engineering teams as platform customers and continuously improve platform usability and adoption.
Infrastructure as Code & Automation
Develop and maintain Infrastructure as Code using Terraform, Pulumi or equivalent technologies.
Automate infrastructure provisioning, configuration, deployment and lifecycle management.
Implement GitOps practices for infrastructure and application deployment.
Develop automated remediation and self-healing capabilities.
Automate cluster provisioning, node management, GPU configuration and workload deployment.
Build repeatable infrastructure patterns across cloud, on-premise and hybrid environments.
Cloud & Infrastructure Architecture
Design scalable infrastructure across AWS, Microsoft Azure, Google Cloud and/or private infrastructure.
Develop hybrid and multi-cloud infrastructure patterns where required.
Design compute, storage and networking architectures optimised for AI workloads.
Implement capacity planning and resource allocation strategies.
Support cloud bursting and dynamic scaling for high-demand AI workloads.
Optimise infrastructure cost without compromising performance, security or reliability.
Storage & Data Infrastructure
Design high-throughput storage solutions for AI training datasets, model artefacts, checkpoints and logs.
Work with object, block and distributed file storage systems.
Implement caching and data-locality strategies to reduce training and inference bottlenecks.
Support high-performance distributed storage environments.
Implement appropriate data lifecycle, retention and quota-management mechanisms.
Ensure storage systems meet required performance, availability and resilience objectives.
Networking
Design and optimise networking for distributed AI workloads.
Work with Kubernetes networking, CNI, service discovery, load balancing and ingress.
Support high-throughput, low-latency communication between distributed compute nodes.
Troubleshoot network performance and reliability issues.
Support RDMA-enabled networking where required for distributed GPU workloads.
Implement appropriate network segmentation, security policies and multi-tenant isolation.
Reliability, Observability & Performance
Define and implement SLIs, SLOs and operational metrics for platform services.
Build comprehensive monitoring, logging and distributed tracing.
Monitor CPU, memory, GPU, storage, network and application-level performance.
Develop dashboards, alerting and automated incident detection.
Conduct performance benchmarking, capacity testing and scalability assessments.
Participate in incident response and root-cause analysis.
Build systems for fault detection, automated recovery and resilience testing.
Security & Multi-Tenancy
Design secure, multi-tenant AI infrastructure platforms.
Implement RBAC, namespaces, quotas, secrets management and workload isolation.
Apply security controls across Kubernetes, cloud infrastructure, containers and APIs.
Support identity and access management across platform services.
Implement audit logging and appropriate security monitoring.
Ensure infrastructure complies with organisational security and data-protection requirements.
Engineering Leadership & Collaboration
Work closely with AI/ML engineers, data scientists, software engineers, DevOps/SRE, cybersecurity and enterprise architects.
Translate AI/ML requirements into scalable infrastructure and platform solutions.
Lead technical design discussions and architecture reviews.
Conduct code reviews and promote engineering quality standards.
Mentor engineers and contribute to technical capability development.
Evaluate emerging infrastructure and AI technologies and recommend appropriate adoption.
Contribute to engineering standards, technical roadmaps and platform strategy.
Key Deliverables / Expected Outcomes
The role holder will be expected to:
Deliver a scalable and reliable AI infrastructure platform.
Improve GPU and compute-resource utilisation.
Reduce deployment and provisioning times for AI workloads.
Enable self-service infrastructure capabilities for AI/ML engineering teams.
Improve reliability and availability of distributed AI workloads.
Reduce infrastructure-related operational toil through automation.
Improve performance and efficiency of distributed training and inference.
Establish effective monitoring, observability and incident-management capabilities.
Maintain secure and appropriately isolated multi-tenant environments.
Support rapid and reliable deployment of AI models into production.
Job Qualifications
Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Technology or a related discipline.
A Master's degree in Computer Science, Distributed Systems, AI/ML, Cloud Computing or a related field would be advantageous.
Relevant professional certifications in cloud, Kubernetes or infrastructure engineering are desirable.
Professional Experience
3-5 years of software engineering, distributed systems, infrastructure, cloud or platform engineering experience, depending on seniority.
Demonstrable experience designing and operating production-scale distributed systems.
Strong production experience with Kubernetes and containerised environments.
Experience building infrastructure platforms for AI/ML workloads is highly desirable.
Experience with GPU infrastructure, distributed training or model serving is strongly preferred.
Experience with cloud and/or large-scale bare-metal infrastructure.
Experience developing production software in C++, Rust or equivalent systems-oriented languages.
Experience with Infrastructure as Code and CI/CD automation.
Experience troubleshooting complex production infrastructure across compute, storage and networking.
Distributed Systems & Platform Engineer · NimrodCareers