Senior Production Engineer
- CI/CD
- IaC
- GitOps
- AI
- Python
- Kubernetes
- Datadog
- Terraform
- ArgoCD
- GitHub Actions
- Java
- Redis
- Snowflake
- AWS
- PostgreSQL
- gRPC
- Protocol Buffers
About Clear Street:
We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.
For more information, visit https://clearstreet.io.
The Role
As a Production Engineer, you sit at the intersection of software reliability and operationalexcellence. You own the health, resilience, and recovery of our production systems—whilespending equal energy innovating solutions that eliminate human toil, reduce incident blastradius, and raise the reliability bar across the entire platform. You will partner closely with engineering, operations, and business teams to understand dailypain points and translate them into lasting automated solutions. Half your time is spent in thetrenches—supporting production, responding to incidents, and deeply understanding how our
systems behave under real conditions. The other half is yours to build: automation, tooling, andobservability platforms that make tomorrow's on-call shift meaningfully easier than today's.
You will work on challenges like:
● Design and build comprehensive monitoring and observability platforms that surface theright signal at the right time—eliminating alert fatigue and accelerating root-causeanalysis.
● Develop intelligent automation and self-healing capabilities that diagnose issues, triggerrecovery workflows, and reduce mean time to recovery (MTTR) without manual intervention.
● Analyze incidents, identify systemic trends, and engineer solutions that prevent entireclasses of failures from recurring.
● Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribalknowledge into scalable platform capabilities.
● Create golden-path operational workflows—making the safest, most reliable path alsothe easiest one for engineering teams to follow.
● Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, andinfrastructure resilience from a production reliability perspective.
● Champion Infrastructure as Code, GitOps, and SRE best practices while helping teamsadopt modern engineering workflows.
● Continuously measure production health through SLIs/SLOs/SLAs, and driveengineering priorities based on reliability data.
● Explore emerging technologies—including AI-assisted diagnostics and developertooling—that transform how we operate production systems.
The Team
We believe resilient systems are built by engineers who understand them end to end. Our Production Engineering team is the first and last line of defense for our production platform. We treat reliability as a product, with uptime and engineer experience as our north stars. Wecombine the discipline of SRE with a builder's mindset: when we see a recurring problem, webuild a solution—not a workaround.
You will work across every engineering and operations team to understand failure modes, quantifyreliability gaps, and build platform capabilities that scale with the organization. Whether it'sreducing MTTR from hours to minutes, building self-service diagnostic tools, or designingproactive alerting that catches issues before customers notice, your work will have immediate,measurable impact.
If you're passionate about making production systems invisible to end users—and you getenergy from both firefighting and building the systems that make fires less likely—you'll thrivehere.
What We're Looking For
We're looking for engineers who combine operational instinct with a builder's discipline.
You should have:
● Strong hands-on Python skills—this is your primary language for automation and tooling.
● Experience in SRE, Production Engineering, Platform Engineering, or a related disciplinewith direct production ownership.
● Proven track record of building automation and diagnostic tooling that improved recoverytimes or reduced operational toil.
● Deep familiarity with cloud-native technologies—Kubernetes, containers, distributedsystems—and how they fail in production.
● Experience with observability platforms such as Datadog, and a strong intuition for what "good" monitoring looks like.
● Exposure to Infrastructure as Code (Terraform) and GitOps-based deploymentworkflows (ArgoCD, GitHub Actions, or similar).
● Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, andPostgres.
● Strong analytical and problem-solving skills—you thrive on ambiguous, high-stakesproduction problems.
● A product mindset applied to operational tooling: you think about usability, adoption, anddocumentation when building internal solutions.
● Excellent communication skills and the ability to work fluidly across engineering,operations, and business stakeholders.
● Self-starter mentality—you identify opportunities, take initiative, and deliver with minimalsupervision.
● Curiosity and a continuous learning mindset; fintech or financial industry background is aplus.
The Technology You'll Work With
You'll operate and build on a modern cloud-native platform that includes:
● Kubernetes & AWS
● Terraform & ArgoCD
● GitHub Actions
● Kafka, Redis
● PostgreSQL & Snowflake
● Datadog
● Python, Go, Java
● gRPC & Protobuf
● Internal Platform APIs and Developer Tooling
What Success Looks Like
Within your first year, you'll have made a measurable impact on production reliability. Success looks like:
● Reducing mean time to detection (MTTD) and mean time to recovery (MTTR) across keyproduction systems.
● Building automation that handles a meaningful percentage of incident scenarios withouthuman intervention.
● Becoming a trusted subject matter expert for core platform components and their failuremodes.
● Delivering observability and diagnostic tools that other engineers actually use anddepend on.
● Establishing SLO baselines and driving engineering investment based on reliability data.
● Spending less of your time—and your teammates time—on repetitive manual toil.
Your impact won't be measured by the number of incidents you respond to—it will be measuredby how reliably our systems run and how quickly we recover when they don't.
Senior Production Engineer · Clear Street