Senior Site Reliability Engineer (MAAS)
- Kubernetes
- Linux
- Debian
- Ubuntu
- Node.js
- VLANs
- DNS
- Ansible
- Bash
- Python
- OpenTofu
- Terraform
- Git
- Prometheus
- Grafana
- VictoriaMetrics
- Incident Response
- KVM
- OpenStack
- VMware
- RBAC
- Secrets Management
- Ceph
- Vault
- Cloudflare
- GDPR
Location: Fully remote, EU timezone (CET ยฑ2h)
Start date: ASAP
Languages: Fluent English required
Industry: Cloud Computing / GPU Infrastructure
About the Opportunity
Pragmatike is hiring aSenior SRE / Infrastructure Engineer to help operate and scale a distributed infrastructure platform spanningbare-metal GPU nodes, Kubernetes, virtualization, networking, and multi-site environments.
Youโll work close to the infrastructure itself, fromBMCs, hardware and MAAS provisioning through Kubernetes, networking, observability, automation, and site operations.
This is a hands-on role for someone who enjoys owning infrastructure end-to-end, building reliable systems, and automating everything that can be automated. Startup or hyper-growth experience is a strong plus: autonomy, ownership, and speed matter here.
What Youโll Do
Operate and maintain large-scaleLinux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments.
OwnMAAS-based bare-metal provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
Operate and maintainproduction Kubernetes clusters, including upgrades, node pools, networking, storage, security hardening, and troubleshooting.
Design and maintain multi-site networking acrossVLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS.
Automate infrastructure provisioning and operations usingAnsible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
Build and maintain automated deployment workflows includingPXE, Preseed, and cloud-init.
Operate observability platforms usingPrometheus, Grafana, Alertmanager, VictoriaMetrics/VictoriaLogs, or comparable tooling.
Define and improveSLIs, SLOs, alerting, and reliability practices across infrastructure and platform services.
Lead infrastructureincident response, troubleshooting, escalation, and post-incident improvements.
Maintain on-call processes and operational coverage across distributed environments.
Work close to the hardware layer, includingIPMI/Redfish, BMCs, RAID, storage, hardware diagnostics, and GPU infrastructure.
Manage virtualization platforms includingProxmox, KVM/libvirt, OpenStack, or VMware, including GPU passthrough where required.
Build and maintain internal infrastructure tooling forhost discovery, configuration, IPAM, hardware health, and operational automation.
Own infrastructure lifecycle activities includingsite onboarding, maintenance, decommissioning, drift detection, and operational runbooks.
Work closely with engineering and cross-functional teams to improve reliability, resource utilization, and operational efficiency.
What Weโre Looking For
5+ years of hands-on SRE, Infrastructure, Systems, or Platform Engineering experience.
Expert-levelLinux administration, particularly Debian/Ubuntu.
Strong production experience withMAAS and bare-metal provisioning.
Expert-level, hands-on experience operatingKubernetes in production, including cluster lifecycle, networking, storage, upgrades, and troubleshooting.
Strong network engineering skills acrossVLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
Strong automation skills withAnsible, Bash and/or Python.
Experience withTerraform/OpenTofu and Git-based infrastructure workflows.
Production experience withPrometheus/Grafana or comparable observability platforms.
Experience withincident response, on-call operations, monitoring, alerting, and reliability practices.
Experience withProxmox, KVM/libvirt, OpenStack, VMware, or comparable virtualization technologies.
Experience withbare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
Strong understanding of distributed systems, container orchestration, and infrastructure reliability.
Experience with infrastructure security includingRBAC, firewalls, network policies, secrets management, and security hardening.
Ability to createSOPs, runbooks, and operational processes from scratch.
Comfortable working autonomously in a fast-paced, engineering-driven environment.
Nice to Have
Experience operatingGPU infrastructure or GPU-heavy Kubernetes/bare-metal environments.
Proxmox VE withZFS/Ceph and GPU passthrough.
VictoriaMetrics / VictoriaLogs or similar large-scale observability platforms.
NetBox or other IPAM / infrastructure inventory platforms.
Vault, SOPS, Atlantis, or similar infrastructure security and automation tooling.
Experience withCloudflare APIs, DNS automation, or tunnels.
Experience withUniFi or comparable site networking platforms.
Experience withservice mesh or advanced CNI implementations.
Go experience, particularly for infrastructure tooling or custom exporters.
Experience withCeph or distributed databases running on bare metal.
Experience operating infrastructure across multiple sites and timezones.
Experience establishingSRE frameworks and reliability practices in growing organizations.
Why Join Us
100% remote with flexible working hours
High-impact role with significanttechnical ownership and autonomy
Work directly withbare-metal, Kubernetes, networking, and GPU infrastructure
International, engineering-driven team
Strong focus onautomation, reliability, and infrastructure at scale
Opportunity to shape the architecture and operational foundations of a growing cloud platform
Pragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation.
In accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration.
Senior Site Reliability Engineer (MAAS) ยท Pragmatike