Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

DevOps Engineer (Founding Team)

Gradientflo Labs
🇮🇳 India | 🇺🇸 United States
On-site
12 months ago
$24,000 – $32,000 / year
  • AI
  • Devops
  • CI/CD
  • GCP
  • Notion
  • GitHub Actions
  • Cloud Run
  • IaC
  • Terraform
  • Firestore
  • Pub/Sub
  • VPC
  • OpenTelemetry
  • Grafana
  • Gemini
  • IAM
  • Disaster Recovery
  • Neo4j
  • Redis
  • Prometheus
  • Vertex AI
  • Next.js
  • FastAPI
  • Secrets Management
  • GitHub
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

Role Overview

You will be theguardian of infrastructure and velocity. Every AI artifact, every mini-app, every TutorAgent interaction depends on the reliability of the pipelines and cloud environment you build. At Vibecoderz, DevOps isn’t a support role — it’s thebackbone of product velocity and reliability.

As the founding DevOps engineer, you will design ourCI/CD pipelines, observability stack, infra-as-code, and security posture from day zero. You’ll ensure that when a developer ships code, it’s live in production safely, quickly, and traceably. You’ll architect theserverless-first cloud strategy on GCP, balancing performance, cost, and scale as we grow from MVP to 100K+ MAUs.

You’ll useLinear (execution), Notion (runbooks/docs), and GitHub (actions & infra code) to make every workflow reproducible, transparent, and automated.

Key Responsibilities

  1. CI/CD Pipeline Ownership

    • Design and maintain GitHub Actions workflows for FE, BE, AI, and agent services.

    • Automate build, test, and deploy to Cloud Run with zero-downtime releases.

  2. Infrastructure as Code (IaC)

    • Implement Terraform scripts for GCP (Cloud Run, Firestore, Pub/Sub, VPCs).

    • Maintain environment parity (dev, staging, prod).

  3. Observability & Monitoring

    • Set up OpenTelemetry tracing for multi-agent workflows.

    • Configure dashboards (Cloud Trace, Grafana) for latency, errors, and throughput.

  4. Cost & Resource Optimization

    • Track infra costs, optimize workloads, and enforce scaling policies.

    • Benchmark agent workloads across Gemini Flash vs. Pro vs. custom models.

  5. Cloud Security & Compliance

    • Enforce IAM best practices, firewall rules, and secret management.

    • Build guardrails for prompt injection and unsafe agent actions at the infra level.

  6. Release Management

    • Define release pipelines with feature flags, rollbacks, and canary deploys.

    • Ensure smooth collaboration between PM, engineers, and QA.

  7. Disaster Recovery & Resilience

    • Build automated backup and recovery strategies for Firestore + Neo4j + Redis.

    • Design failover strategies for critical agent services.

  8. Agent Infrastructure Support

    • Support Browser-Use scaling for the Vibe Browser.

    • Manage GPU/TPU allocations for Gemini/Vertex pipelines if required.

  9. Collaboration & Enablement

    • Write runbooks and incident playbooks in Notion.

    • Train engineering team to self-serve common workflows.

  10. Problem Solving

    • Debug infra bottlenecks, trace latency across services, and enforce SLAs.

Success Metrics

90 Days (Probation):

  • CI/CD pipeline live for FE + BE services.

  • Terraform-based infra deployed and reproducible.

  • OpenTelemetry traces visible for at least 2 core user flows.

12 Months:

  • 99.9% uptime across production workloads.

  • <200ms latency for API responses across multi-agent workflows.

  • Fully automated deployments with rollback & feature flag system.

  • Disaster recovery tested with<5 min RTO (Recovery Time Objective).

Must-Haves

  • 10+ years in DevOps/SRE roles for high-scale products.

  • Mastery of CI/CD, Terraform, GCP services (Cloud Run, Pub/Sub, Firestore).

  • Proven experience with observability stacks (OpenTelemetry, Prometheus, Grafana).

  • Deep knowledge of cloud security, IAM, and infra cost management.

  • Background in scaling infra for developer or AI products.

Nice-to-Haves

  • Experience withVertex AI/ML infra and GPU/TPU scaling.

  • Prior work onreal-time, multi-agent systems.

  • Contributions to open-source DevOps tooling.

  • Startup/founding engineer experience.

Tech Stack Visibility

  • Infra: Terraform, GCP (Cloud Run, Pub/Sub, Firestore, VPC)

  • CI/CD: GitHub Actions

  • Observability: OpenTelemetry, Cloud Trace, Grafana

  • Security: IAM, Secrets Manager, GCP Firewall

  • Other: Redis, Neo4j, Browser-Use scaling infra

Assessment

Objective: Validate ability to design and operate production-grade infra for Vibecoderz.

Challenge (Candidate PoC):

  1. CI/CD Setup

    • Create a GitHub Actions workflow to:

      • Run unit tests for FE (Next.js) + BE (FastAPI).

      • Deploy BE service to Cloud Run on merge tomain.

  2. IaC

    • Write Terraform scripts to provision:

      • Cloud Run service

      • Firestore DB

      • Pub/Sub topic for agent comms

  3. Observability

    • Add OpenTelemetry traces for one workflow (Text → Course).

    • Export traces to Cloud Trace and provide a screenshot of latency breakdown.

  4. Security

    • Configure IAM policy with least-privilege roles.

    • Add secrets management (e.g., API keys) to the workflow.

Deliverables:

  • GitHub repo with workflows + Terraform configs.

  • Cloud Run URL for deployed BE service.

  • Tracing screenshot with latency insights.

  • Short README explaining infra choices + tradeoffs.

Evaluation Criteria:

  • CI/CD Workflow Robustness (25%)

  • IaC Quality & Reproducibility (25%)

  • Observability & Monitoring Depth (20%)

  • Security & IAM Best Practices (15%)

  • Documentation & Clarity (15%)

DevOps Engineer (Founding Team) · Gradientflo Labs

Auto apply with Likeremote