
Site Reliability Engineer
Loft Orbital Solutions
🇦🇪 United Arab Emirates
On-site
Senior
8 months ago
- Devops
- CI/CD
- IaC
- Grafana
- GCP
- Kubernetes
- Prometheus
- Loki
- Python
- Rust
- C++
- Java
- TCP/IP
- DNS
- GitOps
- ArgoCD
- FinOps
- Terraform
- Ansible
- Square
8 months ago
Responsibilities
- Collaborate with developers, test engineers and satellite operators to foster a strongSatDevOps culture.
- Design and roll-out cloud solutions for ourtesting and operations infrastructure. Find the best trade-offs between existing and additional cloud resources toscale and help Orbitworks achieve its mission.
- Design, implement, and maintainscalable, reliable, and secure infrastructure in a hybrid cloud environment.
- Improve ourdeveloper and test engineers experience by building better tools, workflows, and environment to streamline
- Lead efforts to automate and optimize systems, includingCI/CD pipelines,infrastructure provisioning (IaC), and deployment workflows for test on the ground and operations in space.
- Own and evolve ourobservability stack (metrics, tracing, logs) to improve usability and performance. Grafana-centric ecosystems are a plus.
- Implement and advocate forbest practices in software reliability, fault tolerance, and performance tuning.
- Proactivelyidentify, investigate, and resolve system reliability issues, performing root cause analyses and implementing long-term fixes.
- Partner with teams to design and operateSoftware Defined Network (SDN) solutions.
- Contribute to acollaborative and inclusive team culture where respectful debate and continuous learning are celebrated.
- Initially,handle and manage the link between cloud and network/software/hardware infrastructure. Assume I&T (Information technology) responsibilities as much as necessary to start with.
Must Haves:
- Strong experience withpublic cloud infrastructure, ideally GCP.
- Deep expertise inKubernetes, architecture, deployment, ops, and resource optimization.
- Demonstrated ability todesign and build scalable, highly available systems.
- Familiarity withSoftware Defined Networking (SDN) concepts and tools.
- Experience implementing and maintainingobservability stacks (Grafana, Prometheus, Loki, etc.).
- Proficiency in at least one backend language:Go, Python, Rust, C/C++, or Java.
- Deep understanding and hands-on experience withDevOps practices: CI/CD, infrastructure as code (IaC), and automation.
- Proven track record of working infast-paced, high-growth technical environments.
- Strongnetworking knowledge (TCP/IP, DNS, routing, switching, firewalls, VPNs, secure networks).
- Deep experience inSystems Administration.
- Excellent problem-solving skills and ability to operate independently with aproactive, results-driven mindset.
- Strongcommunication skills; thrives in a multicultural, cross-functional team.
Nice to Have:
- Hands-on experience withGitOps frameworks (ArgoCD, FluxCD).
- Interest or experience inFinOps andcost-optimized architectures.
- Understanding oforchestration in resource-constrained environments, like space systems.
- Knowledge ofinfrastructure as code frameworks (Terraform, Ansible or similar)
- Knowledge ofsystems engineering tools and SDLC governance.
- Cybersecurity Awareness.
- Familiarity withsecurity practices, vulnerability scanning, threat detection, risk mitigation.
Site Reliability Engineer · Loft Orbital Solutions