Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
NC

Site Reliability Engineering

NR Consulting - India
🇮🇳 India
On-site
3 days ago
  • System Design
  • CI/CD
  • Devops
  • Incident Response
  • Ansible
  • NGINX
  • Active Directory
  • SAML
  • OAuth
  • Jenkins
  • Groovy
  • Bitbucket
  • Git
  • TCP
  • Configuration Management
  • mTLS
  • SSL
  • SSH
  • AWS
  • Incident Management
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV
Title: Site Reliability Engineering
Location: Pune
Exp: 5+ Yrs

Job Description:

• Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operations, and refinement.
• Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
• Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
• Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
• Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
• Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Client in DevOps automation and best practices.
• Practice sustainable incident response and blameless postmortems.
• Take a holistic approach to problem solving, by connecting the dots during a production event thru the various technology stack that makes up the platform, to optimize mean time to recover
• Collaborate with a global team spread across tech hubs in multiple geographies and time zones
• Share knowledge and mentor junior resources.
• Develop and maintain automation pipelines for certificate renewal, traffic routing, alerting, and compliance reporting using tools like Ansible, Venafi.
• Drive improvements in ITSM and DQ SLOs, ensuring timely CRQ status updates and incident closure.
• Lead initiatives for Safety & Soundness and Operational Excellence across quarterly EPICs, covering areas such as PCI compliance, threat/toil management, self-healing, and ITSM defect resolution

All About You
• Background in operational resiliency and self-healing systems.
• Understanding of two factor authentication.
• Strong documentation and communication skills.
• Strong understanding and experience in implementing NGINX configuration.
• Intermediate understanding of Active Directory (Users / Groups), SAML, LTPA, SSO, Oauth.
• Understanding of DEVOPS technologies like Chef, Jenkins, Groovy, shell scripting, bitbucket, GIT.
• Experience in working with or implementing automation workflows and/or scripting development.
• Understanding of:
o Client-server relationships
o Network concepts (Layer 1 to Layer 3)
o Stack trace analysis (TCP dumps, heap dumps, CPU/memory analysis, thread dumps).
o Load balancers and application firewalls.
o Operating System navigation.
o Logging and monitoring methods, standards, and tools.
o High availability and business continuity planning
o Caching concepts
o Configuration management
• Awareness of security implementations, certificate management lifecycle, mutual TLS, SSL handshake, SSH keys, symmetric and asymmetric encryptions.
o Experience with AWS infrastructure and secure access practices.
o Familiarity with ITSM processes, compliance frameworks, and incident management.
o Excellent communication and collaboration skills across cross-functional teams

Site Reliability Engineering · NR Consulting - India

Auto apply with Likeremote