Stellar - Director of SRE
- 🇺🇸 United States
- Hybrid
- Manager or above
- 1 day ago
- Kubernetes
- CI/CD
- GitHub
- IaC
- Disaster Recovery
- Incident Response
- AI
- AWS
- GCP
- Lean
- Secrets Management
About Stellar Development Foundation
TheStellar Development Foundation (SDF) is a mission-driven organization supporting the development and growth of theStellar blockchain network, an open-source platform designed to expand access to the global financial system.
Since 2014, Stellar has grown into a global blockchain ecosystem used by developers and companies building financial applications and infrastructure around the world.
SDF is now looking for aDirector of Site Reliability Engineering to lead its SRE function and shape how engineering teams own, operate and improve production services.
The Role
This is a senior engineering leadership positionreporting directly to the CTO.
You’ll lead a small, high-leverage SRE team while defining the broaderSRE vision, operating model and reliability culture across engineering.
Rather than SRE acting as the operational owner of every production system, engineering teams at SDF own the services they build. Your role will be to create theinfrastructure, frameworks, tooling, standards and observability practices that allow those teams to operate their services reliably and independently.
You’ll combine hands-on technical judgment with organizational leadership, helping SDF improve reliability, infrastructure maturity and developer productivity without introducing unnecessary process or complexity.
What You’ll Work On
You’ll lead, coach and develop a distributed SRE team while establishing its charter, priorities, operating model and measures of success.
A major focus will be defining and rolling out aService Ownership & Maturity Framework, establishing appropriate reliability and operational standards based on the criticality of individual services.
You’ll own and evolve core engineering infrastructure across:
Cloud infrastructure and foundations
Kubernetes and containerized compute
CI/CD and deployment infrastructure
Observability, monitoring and alerting
Secrets and access management
GitHub workflows
Infrastructure-as-code and automation
You’ll help engineering teams become stronger owners of their production services through better dashboards, runbooks, alerting, escalation paths, deployment practices and operational readiness.
You’ll also improvedeployment automation, resilience, self-healing systems, disaster recovery and service reliability, prioritizing improvements based on real operational risk and impact.
Another important part of the role will be evolvingincident response, postmortems, escalation and on-call practices across a geographically distributed engineering organization.
You’ll build paved paths and self-service infrastructure that reduce engineering toil and cognitive load while allowing teams to ship faster without compromising reliability.
The role also works closely withSecurity, Compliance, Legal, Finance, Procurement and Corporate IT wherever cloud infrastructure, access management, vendors or operational controls intersect with engineering.
SDF is also interested in pragmatically exploringAI-assisted and agentic workflows where they can improve infrastructure operations, observability, developer productivity and service ownership.
What We’re Looking For
You bring10+ years of experience across Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, cloud infrastructure, production operations or closely related areas.
You also have5+ years of leadership experience, managing or formally developing SRE, infrastructure, platform or reliability engineers.
You have strong experience defining:
Team charters and operating models
Infrastructure and reliability roadmaps
Engineering standards and practices
Success metrics and operational maturity frameworks
You bring deep technical judgment acrossdistributed systems, cloud infrastructure, production operations, automation, reliability engineering and operational risk.
You should also have strong practical experience with:
AWS, GCP or comparable cloud platforms
Kubernetes and container orchestration
Infrastructure-as-code and declarative infrastructure
CI/CD and deployment safety
Observability, logging and monitoring
SLOs and SLIs
Incident response and postmortems
On-call systems and operational readiness
You’ve helped application or product engineering teams take greater ownership of production systems and understand how to balancedeveloper velocity, reliability and operational responsibility.
You’re pragmatic about tooling and comfortable deciding when tobuild, buy, adapt, simplify or retire infrastructure based on the underlying engineering problem.
Finally, you’re comfortable operating in a lean engineering organization where influence comes fromtechnical credibility, judgment and execution, and can communicate effectively with the CTO and other senior engineering leaders.
Particularly Relevant Experience
Experience in any of the following would be especially valuable:
Leading SRE, Platform or Infrastructure teams inlean, high-agency organizations
Supportingglobally distributed engineering teams and 24/7 production environments
Buildingself-service infrastructure and paved paths
Improving developer productivity through automation and toil reduction
Infrastructure security, secrets management and cloud access controls
Financial services or regulated environments
Blockchain, crypto or Web3 infrastructure
Vendor and infrastructure platform evaluation
ApplyingAI-assisted or agentic systems to infrastructure, operations, observability or developer workflows
Stellar - Director of SRE · deCircle