IW
JD Grade 10 - Senior Site Reliability Engineer (Private Cloud) - FY26Q4- Engineering
Info Way Solutions LLC
Location not stated
Remote
Senior
2 months ago
- IaC
- VMware
- AWS
- GCP
- Kubernetes
- Helm
- ArgoCD
- Python
- Bash
- Ruby
- Scala
- Configuration Management
- Ansible
- Terraform
- Unix
- Linux
2 months ago
Location: Remote, prefer PST hours
Visa Requirements: UC Citizenship
Job Summary
Our SRE Private Cloud team is responsible for building and scaling the platform that
supports millions of devices across the world. Our customer base has grown by a factor
of 2-3 every year, serving more than 4 billion HTTP requests per day across eight data
centers. Our customers depend on the our cloud to monitor their critical infrastructure
of network switches, security appliances, wireless Client, and security cameras. We
embrace the *nix way, automate away tedious tasks and build infrastructure as code.
As a Site Reliability Engineer within the Private Cloud team, you will architect, build, and
support highly available, scalable, and resilient private cloud infrastructure. You will
focus on automating operations, improving platform reliability, and enabling seamless
service deployment for internal customers. This role requires collaboration with cross-
functional teams to ensure platform stability, security, and continuous improvement
aligned with the mission to provide reliable compute and storage services.
Responsibilities
• Design, implement, and manage scalable, highly available private cloud
infrastructure and orchestration platforms supporting compute and storage
services.
• Automate infrastructure provisioning, configuration, and deployment using
Infrastructure as Code (IaC) tools.
• Monitor, troubleshoot, and optimize platform performance, reliability, and
security.
• Develop and maintain standardized OS images and deployment frameworks for
bare metal, virtual machines, and cloud instances.
• Collaborate with other engineering teams to ensure platform uptime, cultivate
sound engineering principles and represent our engineering values.
• Build end-to-end documentation, instrumentation, and automation to enable
self-healing and resiliency.
• Participate in on-call rotations to support production environments.
• Lead and contribute to large-scale projects, promoting engineering standards
and team values such as inclusivity, customer focus, and continuous
improvement.
You are an ideal candidate if you have:
• 7+ years of experience designing, deploying, and operating mid to large-scale
enterprise or cloud environments.
• Familiarity with private cloud technologies including VMware, bare metal
environments, file and block storage, as well as cloud infrastructure providers
like AWS, GCP.
• Operational experience with container orchestration platforms (Kubernetes) at
production scale, workload deployment via Helm, ArgoCD
• Familiarity with automation frameworks for OS image creation and deployment.
• Strong scripting and coding skills in languages such as Python, Bash, Ruby, or
Scala.
• Expertise with infrastructure/configuration management tools such as Ansible,
or Terraform.
• Deep understanding of systems and application design with operational trade-
offs.
• Proficiency with modern Unix/Linux operating systems and distributions.
• Strong collaboration, mentoring, documentation and communication skills.
• Ability to prioritize tasks, work independently, and call out issues effectively.
• BS or MS degree in Computer Science, Engineering, or equivalent experience.
JD Grade 10 - Senior Site Reliability Engineer (Private Cloud) - FY26Q4- Engineering · Info Way Solutions LLC