Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
A

Lead SRE Engineer

Apolis
πŸ‡ΊπŸ‡Έ United States
Hybrid
Staff / Principal
3 weeks ago
$60 – $65 / hour
  • Java
  • CI/CD
  • Incident Management
  • Devops
  • System Design
  • Python
  • JVM
  • Disaster Recovery
  • Spring Boot
  • Linux
  • Unix
  • REST API
  • Microservices
  • Jenkins
  • GitHub Actions
  • GitLab CI
  • Kubernetes
  • Docker
  • AWS
  • Azure
  • GCP
  • Splunk
  • Prometheus
  • Grafana
  • SQL
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Job Title:Lead SRE Engineer – Strong Java background

Location: Jersey City, NJ – 3 Days onsite role

Long term Project

Job Summary:We are seeking a highly experiencedLead Site Reliability Engineer (SRE) with strong Java development experience to join the technology organization supporting highly available, scalable, resilient, and business-critical applications.

The ideal candidate will have12+ years of overall technology experience, with strong hands-on expertise inJava, application reliability, production engineering, observability, automation, cloud technologies, CI/CD, incident management, and performance engineering.

The Lead SRE will apply asoftware engineering mindset to production operations, developing automation and reliability solutions rather than relying solely on traditional infrastructure support. The role will work closely with application developers, architects, DevOps engineers, platform teams, and technology stakeholders to improve system availability, stability, scalability, and operational efficiency.

Banking's technology organization emphasizes application/infrastructure availability, operational stability, system design, Java/Python, resiliency, automation, and strong engineering practices, making this a particularly strong match for aJava-heavy SRE profile.

Key Responsibilities

  • LeadSite Reliability Engineering initiatives for critical enterprise applications and platforms.
  • Develop and maintainJava-based automation, reliability, monitoring, and operational tools.
  • Apply software engineering principles to improve applicationavailability, scalability, resiliency, and performance.
  • Own production stability and participate inincident management, problem management, and root-cause analysis.
  • Design and implement proactive monitoring, alerting, health checks, and automated remediation.
  • Analyze production issues, identify systemic problems, and implement permanent corrective actions.
  • Establish and improveSRE practices, SLOs, SLIs, SLAs, error budgets, and reliability metrics.
  • Build automation to eliminate repetitive manual operational activities.
  • Work closely with Java development teams to improve application reliability and production readiness.
  • Troubleshoot complexJava/JVM, application, API, database, network, and infrastructure-related issues.
  • Perform application performance analysis, includingJVM, memory, CPU, thread, garbage collection, latency, and throughput analysis.
  • Design and improveCI/CD pipelines and automated deployment processes.
  • Support highly available applications acrosscloud and distributed environments.
  • Implement resiliency patterns including fault tolerance, failover, disaster recovery, and capacity planning.
  • Develop dashboards and observability solutions for application and infrastructure health.
  • Participate in production deployments, release management, and post-production validation.
  • Drive automation and continuous improvement across the application lifecycle.
  • Provide technical leadership and mentoring to other engineers.
  • Collaborate with architecture, development, infrastructure, security, and platform engineering teams.

Required Technical Skills

Must Have:

  • 12+ years of IT/software engineering experience
  • Strong hands-onJava development
  • StrongSite Reliability Engineering / Production Engineering experience
  • Java/Spring Boot or enterprise Java application experience
  • Strong understanding ofJVM internals and Java application performance
  • Production support and troubleshooting of large-scale applications
  • Linux/Unix
  • REST APIs / Microservices
  • CI/CD
  • Jenkins / GitHub Actions / GitLab CI or similar
  • Kubernetes / Docker
  • Cloud experience β€”AWS / Azure / GCP
  • Monitoring and observability tools
  • Splunk / ELK or equivalent logging platforms
  • Prometheus / Grafana or equivalent monitoring tools
  • Strong scripting/automation usingPython, Shell, or similar
  • Incident management andRoot Cause Analysis (RCA)
  • Application performance and capacity management
  • High availability, resiliency, scalability, and disaster recovery concepts
  • Strong SQL/database troubleshooting skills
  • Strong communication and stakeholder-management skills

Lead SRE Engineer Β· Apolis

Auto apply with Likeremote