
Site Reliability Engineer (SRE)
Unlimit
🇧🇦 Bosnia and Herzegovina
On-site
7 months ago
- AI
- Incident Response
- Linux
- Kubernetes
- Incident Management
- AWS
- IaC
- Terraform
- Ansible
- Configuration Management
- CI/CD
- Prometheus
- Grafana
- Zabbix
- Splunk
- PagerDuty
- GitLab
- Agile
- Scrum
- GCP
- Python
- Golang
- Bash
- PostgreSQL
- MongoDB
- RabbitMQ
- Apache
- NGINX
7 months ago
As aSite Reliability Engineer (SRE) at Unlimit, you will help ensure the reliability, scalability, and performance of our core platform and services. You’ll work closely with Engineering and other stakeholders to design, build, and operate cloud-based infrastructure and distributed systems—while continuously improving automation, observability, and incident response.This role blendsLinux systems engineering,cloud infrastructure,Kubernetes, andautomation with a strong focus onservice availability,operational excellence, andcontinuous improvement.
Key responsibilities
- Platform reliability & operations:
- Ensure the availability, resilience, and performance of the platform and supporting services.
- Own and improve incident management, including troubleshooting, escalation handling, and follow-ups aligned to SLAs.
- Participate in an on-call rotation, supporting production systems and driving reliability improvements from real incidents.
- Infrastructure engineering (Linux / Cloud / Kubernetes):
- Design, deploy, configure, and manageLinux-basedsystem architecture across environments.
- Build and support platform implementations usingAWSand other cloud technologies (compute-centric services and related infrastructure).
- Design and implement large and complex technology projects, from initial design through production rollout and operational handover.
- Support and maintain Kubernetes-based workloads and platform components.
- Automation & Infrastructure as Code:
- Build tooling and solutions to automate recurring operational tasks.
- UseInfrastructure as Code (IaC) to standardize and scale:Terraform for provisioning ,Ansible for configuration management and automation
- Improve reliability by reducing manual steps and enabling repeatable deployments.
- CI/CD & developer enablement:
- Manage and maintainCI/CD pipelines across 20+ repositoriesspanning multiple technology stacks.
- Partner with Engineering teams to improve build/release consistency, pipeline reliability, and deployment safety.
- Observability & operational readiness:
- Implement and enhancemonitoring, logging, and alerting, using tools such as:Prometheus,Grafana, Zabbix,Splunk, PagerDuty (or equivalent incident alerting/response tooling).
- Use metrics and incident learnings to reduce noise, improve signal, and shorten time-to-detect/time-to-recover.
- Documentation & standards:
- Produce clear, formal documentation including: Configuration standards, Troubleshooting runbooks, Infrastructure and architecture design documentation.
- Contribute to internal standards that improve consistency, security, and operational maturity.
Required skills & experience
- 5+ years of hands-on experience inLinux systems administration / engineering in production environments.
- Strong working knowledge of the following (or equivalents):Linux,Kubernetes,GitLab, Terraform,Ansible.
- Experience working inAgile (Scrum) teams.
- Experience withAWS (compute-focused services) and/orGoogle Cloud Platform.
- Proven experience withdistributed systems design, maintenance, and troubleshooting.
- Strong scripting/coding ability in at least one of:Python,Golang,bash.
- Experience with observability and incident response tooling such as:Zabbix,Splunk,Prometheus,Grafana,PagerDuty.
- Strong communication skills inEnglish, with the ability to work effectively with customers, vendors, partners, and internal teams across levels.
- Working knowledge (expected familiarity) with datastores and messaging systems such as:PostgreSQL,MongoDB,RabbitMQ.Also Web/application infrastructure components such as:Apache,Nginx
- Demonstrated ability to learn quickly, work independently, make good decisions, and collaborate as a team player in fast-changing environments.
- StrongAI-driven mindsetand curiosity about emerging AI technologies.
- Hands-on experience usingAI tools(e.g., LLMs, automation frameworks, AI-assisted development tools) to enhance productivity or system performance.
Nice to have
- Experience operatinghighly available, high-volume web services.
- Strong initiative and self-starter attitude with minimal supervision.
- Demonstrated success reducing operational toil through automation and better tooling.
- Experience improving SLOs/SLIs, error budgets, or formal reliability practices (if applicable to your background).
Site Reliability Engineer (SRE) · Unlimit