
Engineering Lead - Bildbot App Health
- AI
- Claude
- Devops
- Next.js
- React.js
- Node.js
- Unit Testing
- End-to-End Testing
- Load Testing
- SQL Injection
- Ruby on Rails
- Django
- SvelteKit
- Claude Code
- OWASP
- Secrets Management
- XSS
- CSRF
- CI/CD
- IaC
- Copilot
- Gemini
- Codex
- Square
- Cursor
- Equity
Engineering Lead - App Health Lead
Build Something Beautiful.
Our Beliefs
Everyone builds.
Engineering is free. Distribution is the bottleneck.
The constraint is knowing what to ask.
AI-first.
Launch 40 times a day, not 4 times a month.
Velocity over perfection.
Don't do the thing 10 times. Build an AI-powered system that can do it 1,000 times, 100 times faster.
Help Builders Make Better Apps
AI lets anyone build an app in a weekend. But "it works" and "it's ready" are very different things.
Bildbot is building App Health β automated scanners that tell vibe coders exactly what's wrong with their app and how to fix it. Security vulnerabilities, performance bottlenecks, missing tests, broken accessibility, outdated dependencies, legal compliance gaps. We scan the codebase, surface findings by severity, track them over time, and give actionable fix plans β so a solo founder shipping with Claude doesn't have to become a security expert or a DevOps engineer to have a trustworthy product.
We've built the infrastructure: 10 scanner categories, a findings database, severity tracking, trend charts, fix-plan workflows, and a dashboard that makes it all legible.What we need now is someone who knows the domain cold β someone who can make every scanner accurate, reduce false positives to near-zero, catch what matters across different tech stacks, and tell us what we're missing.
This isn't a traditional DevOps role. You're not managing infrastructure. You'rebuilding a product that replaces the security consultant, the performance auditor, and the reliability engineer β all wrapped in scanners that run in seconds.
What You'll Do
Own scanner accuracy. Our scanners are static analysis scripts that audit codebases and output structured findings. You'll make them smarter β reducing false positives, catching real issues, tuning severity levels, and ensuring findings are actionable across Next.js, React, Node, and beyond.
Build new scanner categories. We have security, performance, legal, code cleanup, dependencies, unit testing, e2e testing, load testing, accessibility, and mobile. What's missing? Rate limiting? Infrastructure hardening? Database optimization? API security? You'll identify the gaps and build the scanners to fill them.
Define what "healthy" means. When should a finding be urgent vs. low? When does a security issue matter at 100 users vs. 100,000? Our scanners use stage-aware severity (pre-launch, live, scaling) β you'll calibrate these thresholds based on real-world experience, not guesswork.
Validate against real attacks and real failures. You've seen what actually breaks in production. SQL injection that a scanner missed. A performance cliff that only shows up under load. A dependency vulnerability that's technically present but not exploitable. You bring the scar tissue that makes our tooling credible.
Make scanners work across tech stacks. Our scanners currently target Next.js/React. You'll generalize them β or build stack-specific variants β so they work for Rails, Django, Express, SvelteKit, whatever vibe coders are shipping with.
Build the fix-plan intelligence. A finding without a fix is just noise. You'll write the remediation guidance β the specific code changes, config tweaks, and architectural patterns that resolve each finding. Eventually, you'll help build AI-powered auto-fix that does it for them.
Use AI to build faster. Claude Code is your daily driver. You'll use it to write scanners, test edge cases, and prototype detection rules. When you find something Claude gets wrong about security or performance, that's signal β it means our users are getting it wrong too, and our scanner needs to catch it.
Who You Are
You've built scanners or static analysis tools. Maybe at a security company, maybe as internal tooling, maybe open-source. You know the architecture: AST parsing, pattern matching, config file analysis, output formatting. You've dealt with false positive rates and know how to tune them.
Deep in security. OWASP Top 10 isn't a checklist you Google β it's muscle memory. You understand CSP headers, auth patterns, input validation, secrets management, XSS vectors, CSRF, and injection attacks. You can look at a codebase and spot what's wrong without running a tool.
You get our audience. Our customers are solo founders, indie hackers, and first-time builders using AI to ship products they couldn't have built alone. They've told Claude "make my app secure" and gotten a result that looks right but might not be. You understand that gap β between "Claude said it's fine" and "it's actually fine" β because you've seen what AI-assisted code gets wrong. You respect what these builders are doing and you want to give them a safety net, not a lecture.
Performance and reliability instincts. You know why N+1 queries kill you at scale, when to add caching, what makes a database choke, and which performance issues are noise vs. ticking time bombs. You've been on-call and you know what breaks.
DevOps fluency. CI/CD, monitoring, alerting, infrastructure-as-code. You've set up pipelines, configured scanners in CI, and understood why a scan that works locally fails in production. Not your whole job β but part of your toolkit.
AI-native. You use Claude Code, Copilot, Gemini or Codex constantly as a primary tool. You've used AI for security audits and know where it's strong (finding patterns) and where it's weak (understanding context). You're not threatened by AI β you're building tools that make AI-generated code trustworthy.
Multi-stack. You've worked across at least 2-3 web frameworks. You know Next.js extremely well. You know that security patterns differ between Next.js and Rails, that performance bottlenecks look different in serverless vs. traditional hosting, and that a one-size-fits-all scanner will miss things.
Pragmatic. You know the difference between a theoretical vulnerability and an exploitable one. You don't flag everything as urgent. You prioritize based on real risk, real impact, and the stage of the product.
Systems thinker. When you find a false positive, you don't just suppress it β you fix the detection logic so that class of false positive never happens again. You think in patterns, not instances.
This Isn't For You If...
You've only used scanners, never built them
You think security is someone else's problem until there's a breach
You need enterprise-scale infrastructure to do good work
You dismiss AI-generated code without understanding what it gets right
You over-classify β everything is "critical" and nothing is actionable
You can't explain a vulnerability to a non-technical team member in plain English
You've never been surprised by a false positive or false negative in production
Details
Location: Hybrid, 3-4 days/week in office (Cambridge/Kendall Square area). Scanner design benefits from rapid iteration in person.
Equity: Everyone at Bildbot will be an owner.
Compensation: Competitive salary + equity. We're flexible on the cash/equity mix.
Tools: Claude Code, Cursor, whatever scanner frameworks and security tools you bring. We'll pay for what you need.
To Apply
Show us something you've built that detects problems in code. Include:
A scanner, linter rule, or static analysis tool you've written
An example of a false positive you caught and how you fixed the detection logic
Your experience with AI-assisted development β what it gets right and wrong about security/performance
Bonus: a finding that surprised you in production β something a scanner should have caught but didn't
bildbot.com
Engineering Lead - Bildbot App Health Β· Bildbot