
Software Engineer, Infrastructure
- AI
- Gusto
- Stripe
- Plaid
- Docker
- Kubernetes
- SOC2
- Datadog
- Incident Response
- IaC
- VPC
- IAM
- Devops
- Helm
- AWS
- GCP
- Azure
- Terraform
- Kustomize
- GitOps
- TypeScript
- AI/ML
- Equity
- Pension
About Bretton AI
AI-native operations for the banks
The role
# **About Bretton AI**Bretton AI is the leading AI agent platform for financial services. Companies like Robinhood, Mercury and Gusto trust us to automate mission critical work, starting with anti-money laundering and counter-terrorism investigations.We’ve raised over $95M from Greylock, Y Combinator, Thomson Reuters Ventures and other top tier investors. We’re based in downtown San Francisco and our team comes from world-class organizations like SpaceX, Google, Netflix, Stripe, Plaid and more.# **The Role**As a **Senior Infrastructure Engineer** , you will own the foundation that enables us to deploy secure, compliant AI systems at major financial institutions fighting financial crime at a massive scale. Our infrastructure is built on a modern, container-native architecture, leveraging Docker and Kubernetes to deliver consistent, auditable deployments across diverse customer environments.You will work directly with our largest customers—institutions serving over a billion people—to architect, automate, and harden our on-premises and cloud environments to meet the strictest regulatory and performance requirements, including SOC 2 compliance. Your work will be informed by real customer needs and will ship to everyone, so you must build enterprise-grade systems, work effectively with engineering and customer teams, understand financial services compliance, and adapt quickly.# **What You’ll Do**- Own and evolve our Kubernetes infrastructure, including cluster management, service mesh configuration, and container security policies.- Design and implement progressive delivery pipelines with canary deployments, automated rollbacks, and deployment health validation.- Build and maintain our observability infrastructure in Datadog, including dashboards, monitors, SLOs, and distributed tracing.- Drive incident response for high-severity outages and proactively model capacity needs for low-latency AI inference.- Architect and automate secure infrastructure using Infrastructure-as-Code for VPCs, IAM policies, Kubernetes manifests, and private cloud deployments.- Maintain and improve the infrastructure controls that support our SOC 2 compliance posture.- Lead customer engagements for enterprise rollouts and mentor mid-level engineers on infrastructure best practices.# **What We’re Looking For**## **Must-Haves:**- 8+ years in infrastructure engineering or DevOps at high-growth or hyperscale companies.- Experience with Docker and Kubernetes, including production cluster management, Helm, and service mesh technologies.- A proven track record of architecting and operating AWS (preferred), GCP, or Azure at an enterprise scale.- Experience with observability platforms, preferably Datadog (metrics, logs, APM, distributed tracing).- A strong background in Infrastructure-as-Code (Terraform, Helm, Kustomize) and safe deployment practices (progressive delivery, canary deployments, GitOps, automated rollbacks).- "Battle scars" from leading outages, capacity events, and large-scale incident reviews.- Strong programming skills in Python.## **Bonus Points:**- Familiarity with TypeScript.- Direct involvement in SOC 2 or other compliance audit preparation or remediation.- Direct experience with private-cloud or on-premises deployments for regulated customers.- Previous experience at startups scaling infrastructure from the early stages to the enterprise level.- A background in fintech or building systems for highly regulated industries.- Experience with AI/ML infrastructure and model deployment at scale.# **Why You’ll Love Working Here**- **Build for Scale:** You thrive at the intersection of technical leadership and customer impact, building systems that enable rapid development while maintaining the highest standards of security, compliance, and reliability.- **Infrastructure as a Product:** You see infrastructure as a product for your engineering peers and understand the value of platform automation in enabling developer velocity.- **High-Impact Work:** Your contributions will have a direct, measurable impact on how financial institutions adopt AI to fight crime.- **Mentorship and Leadership:** You are comfortable balancing technical excellence with mentoring others and leading customer engagements.# **Compensation & Benefits**- $168k - $213k + equity- Comprehensive healthcare, 401k matching, commuter benefits- 15 days PTO + holidays, unlimited sick days- Flexible leave options- Working late? We’ve got you covered with DoorDash and an Uber home_Join us in building AI that protects the global financial system from financial crimes that fund terrorism, human trafficking, and other serious threats._
Software Engineer, Infrastructure · Bretton AI