E
Applications Support Specialist
Ensono
- 🇮🇳 India
- Hybrid
- 14 hours ago
- Java
- JVM
- Python
- PowerShell
- .NET
- Unix
- Linux
- SQL
- Splunk
- Dynatrace
- Grafana
- Prometheus
- ITIL
14 hours ago
Key Responsibilities
Incident & Problem Management
- Leadmajor incident (MI) bridges and restore service with minimum business impact.
- Handle allL3 escalations, perform deep diagnostics across Java, JVM, middleware, OS, and infra.
- Owntechnical RCAs, drive long‑term and systemic remediation.
- Identify recurring failure patterns and risks.
Reliability Engineering
- ApplySRE principles: SLIs/SLOs, error budgets, resilience patterns.
- TuneJVM parameters, analyze thread/heap dumps, and improve performance.
- Influence application architecture forfault tolerance, scalability, and recoverability.
- ValidateDR readiness, failover behavior, and resilience testing outcomes.
Change, Release & Risk
- Providetechnical approval and risk assessment for high-risk changes.
- Enforceoperational readiness for new apps and major releases.
- Ensure changes meetaudit, compliance, and regulatory expectations.
Automation, Monitoring & Observability
- Build advanced automation usingShell/Python/PowerShell.
- Develop frameworks forhealth validation, automated recovery, and compliance checks.
- Define observability standards; optimize alerts and improveMTTR.
Leadership & Mentorship
- Mentor L1/L2 teams; review and approve runbooks, SOPs, and KB articles.
- Act as a trusted technical advisor to stakeholders and leadership.
Skills & Qualifications
Technical (Mandatory)
- Strong knowledge ofapplication architecture, distributed systems, and middleware.
- Java expertise: JVM internals, GC, memory management, thread/heap dump analysis, performance tuning.
- .Net --CLR internals, garbage collection, memory management, thread/dump analysis, and application performance tuning.
- StrongUnix/Linux, networking basics, and advanced scripting (Shell/Python/PowerShell/VBS).
- AdvancedSQL and understanding of databases; Autosys (or equivalent scheduler).
- Handson withobservability tools: Splunk, AppDynamics/Dynatrace, ELK, Grafana, Prometheus.
Reliability & Operations
- Major incident leadership, deep RCA, change/release readiness, DR & resilience engineering.
- Experience inregulated production environments.
Soft Skills
- Strong technical leadership and decision‑making.
- Clear communication during high‑pressure incidents.
- Ownership mindset and business awareness.
Experience & Education
- 7–12+ years in Application Reliability, Production Support, SRE, or platform operations.
- Bachelor’s degree inComputer Science/Engineering or equivalent.
- ITIL, cloud, or industry certifications (preferred).
- Banking/financial domain experience (preferred).
Working Conditions
- On‑call and after‑hours support as required.
- Fast‑paced environment with multiple priorities.
- Hybrid working model
Applications Support Specialist · Ensono