Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
D

Stellar - Director of SRE

deCircle
  • 🇺🇸 United States
  • Hybrid
  • Manager or above
  • 1 day ago
  • Kubernetes
  • CI/CD
  • GitHub
  • IaC
  • Disaster Recovery
  • Incident Response
  • AI
  • AWS
  • GCP
  • Lean
  • Secrets Management
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About Stellar Development Foundation

TheStellar Development Foundation (SDF) is a mission-driven organization supporting the development and growth of theStellar blockchain network, an open-source platform designed to expand access to the global financial system.

Since 2014, Stellar has grown into a global blockchain ecosystem used by developers and companies building financial applications and infrastructure around the world.

SDF is now looking for aDirector of Site Reliability Engineering to lead its SRE function and shape how engineering teams own, operate and improve production services.

The Role

This is a senior engineering leadership positionreporting directly to the CTO.

You’ll lead a small, high-leverage SRE team while defining the broaderSRE vision, operating model and reliability culture across engineering.

Rather than SRE acting as the operational owner of every production system, engineering teams at SDF own the services they build. Your role will be to create theinfrastructure, frameworks, tooling, standards and observability practices that allow those teams to operate their services reliably and independently.

You’ll combine hands-on technical judgment with organizational leadership, helping SDF improve reliability, infrastructure maturity and developer productivity without introducing unnecessary process or complexity.

What You’ll Work On

You’ll lead, coach and develop a distributed SRE team while establishing its charter, priorities, operating model and measures of success.

A major focus will be defining and rolling out aService Ownership & Maturity Framework, establishing appropriate reliability and operational standards based on the criticality of individual services.

You’ll own and evolve core engineering infrastructure across:

  • Cloud infrastructure and foundations

  • Kubernetes and containerized compute

  • CI/CD and deployment infrastructure

  • Observability, monitoring and alerting

  • Secrets and access management

  • GitHub workflows

  • Infrastructure-as-code and automation

You’ll help engineering teams become stronger owners of their production services through better dashboards, runbooks, alerting, escalation paths, deployment practices and operational readiness.

You’ll also improvedeployment automation, resilience, self-healing systems, disaster recovery and service reliability, prioritizing improvements based on real operational risk and impact.

Another important part of the role will be evolvingincident response, postmortems, escalation and on-call practices across a geographically distributed engineering organization.

You’ll build paved paths and self-service infrastructure that reduce engineering toil and cognitive load while allowing teams to ship faster without compromising reliability.

The role also works closely withSecurity, Compliance, Legal, Finance, Procurement and Corporate IT wherever cloud infrastructure, access management, vendors or operational controls intersect with engineering.

SDF is also interested in pragmatically exploringAI-assisted and agentic workflows where they can improve infrastructure operations, observability, developer productivity and service ownership.

What We’re Looking For

You bring10+ years of experience across Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, cloud infrastructure, production operations or closely related areas.

You also have5+ years of leadership experience, managing or formally developing SRE, infrastructure, platform or reliability engineers.

You have strong experience defining:

  • Team charters and operating models

  • Infrastructure and reliability roadmaps

  • Engineering standards and practices

  • Success metrics and operational maturity frameworks

You bring deep technical judgment acrossdistributed systems, cloud infrastructure, production operations, automation, reliability engineering and operational risk.

You should also have strong practical experience with:

  • AWS, GCP or comparable cloud platforms

  • Kubernetes and container orchestration

  • Infrastructure-as-code and declarative infrastructure

  • CI/CD and deployment safety

  • Observability, logging and monitoring

  • SLOs and SLIs

  • Incident response and postmortems

  • On-call systems and operational readiness

You’ve helped application or product engineering teams take greater ownership of production systems and understand how to balancedeveloper velocity, reliability and operational responsibility.

You’re pragmatic about tooling and comfortable deciding when tobuild, buy, adapt, simplify or retire infrastructure based on the underlying engineering problem.

Finally, you’re comfortable operating in a lean engineering organization where influence comes fromtechnical credibility, judgment and execution, and can communicate effectively with the CTO and other senior engineering leaders.

Particularly Relevant Experience

Experience in any of the following would be especially valuable:

  • Leading SRE, Platform or Infrastructure teams inlean, high-agency organizations

  • Supportingglobally distributed engineering teams and 24/7 production environments

  • Buildingself-service infrastructure and paved paths

  • Improving developer productivity through automation and toil reduction

  • Infrastructure security, secrets management and cloud access controls

  • Financial services or regulated environments

  • Blockchain, crypto or Web3 infrastructure

  • Vendor and infrastructure platform evaluation

  • ApplyingAI-assisted or agentic systems to infrastructure, operations, observability or developer workflows

Stellar - Director of SRE · deCircle

Auto apply with Likeremote