
Prompt Engineer (Founding Team)
- AI
- Notion
- GitHub
- JSON
- LangSmith
- Gemini
- Claude
- RAG
- CI/CD
- Natural Language Processing
- JSON Schema
- LangChain
- CrewAI
- ADK
- Firestore
- Redis
- Neo4j
- Cloud Run
- Pub/Sub
- GitHub Actions
- React.js
- YAML
Role Overview
You will be thelanguage architect of Vibecoderz. The TutorAgent and its family of sub-agents only think as clearly as the prompts that guide them — and you will design those “thought patterns.” As the founding Prompt Engineer, you’ll build aprompt library and evaluation system that ensures Vibecoderz produces consistent, safe, and high-quality artifacts: slides, quizzes, code snippets, and runnable mini-apps.
This is a role for abuilder of mental scaffolding. You’ll be deeply involved in orchestratingsystem prompts, tool schemas, chaining strategies, and eval frameworks that make our AI both reliable and delightful. You’ll collaborate with AI Engineers on orchestration, Backend Engineers on schema integration, and PM on defining product outcomes.
You’ll useLinear for execution, Notion for prompt specs and experiments, and GitHub for version control, making every iteration observable, testable, and reproducible.
Key Responsibilities
Prompt Library Ownership -Design, version, and maintain system prompts for TutorAgent, PlannerAgent, CodeAgent, QuizAgent, and ContentAgent.
Prompt Chaining & Orchestration -Implement advanced strategies: reflection, planning, few-shot, self-consistency, and tool-use flows.
Schema Definition -Collaborate with BE and AI engineers to define JSON schemas and tool contracts that LLMs must respect.
Evaluation Pipelines -Use LangSmith or Langfuse to build eval frameworks that test accuracy, latency, drift, and hallucination rates.
Performance Optimization -Continuously refine prompts to minimize cost and latency while maximizing reliability and user trust.
Hybrid Model Router Support -Adapt prompt strategies for Gemini Flash vs. Pro vs. Claude depending on use case (speed vs. depth).
Guardrails & Safety -Embed safety layers in prompts to prevent unsafe, biased, or irrelevant outputs.
RAG Integration -Build prompt strategies for retrieval-augmented generation (GitHub docs, MDN, StackOverflow).
Documentation & Experimentation -Maintain prompt playbooks in Notion, documenting rationale, variations, and outcomes for reproducibility.
Cross-Team Collaboration -Work with PM to translate learning flows into prompt chains, AI engineers for orchestration, and QA for regression testing.
Problem Solving -Debug prompt failures, identify root causes (model vs. schema vs. orchestration), and propose iterative solutions.
Success Metrics
90 Days (Probation):
Deliver prompt library v1 for TutorAgent, PlannerAgent, and QuizAgent.
Build baseline LangSmith eval pipeline with at least 20 golden prompts.
Achieve<15% hallucination rate in core flows.
12 Months:
Maintain<5% hallucination rate across all core workflows.
Prompt system powering 10+ specialized agents in production.
Automated regression evals integrated into CI/CD pipeline.
AI outputs powering >50% of artifacts with consistent schema compliance.
Must-Haves
10+ years in NLP, prompt engineering, or applied AI.
Expertise in LLM prompting strategies and evaluation frameworks.
Strong grasp of JSON schema enforcement, tool-use, and prompt chaining.
Proven ability to debug and refine AI outputs in production products.
Product company background with evidence of shipping AI-first systems.
Nice-to-Haves
Contributions to LangChain, CrewAI, or LangSmith open source.
Research background in prompt optimization, LLM evals, or safety.
Prior work on developer-facing AI tutor or EdTech products.
Startup/founding engineer experience.
Tech Stack Visibility
Prompting & Orchestration: LangSmith, Langfuse, Google ADK
Models: Gemini Pro, Gemini Flash, Gemini Vision, Gemini Live API, Claude 3 Opus
Data: Firestore, Redis, Neo4j (schema enforcement targets)
Infra: Cloud Run, Pub/Sub for agent comms
CI/CD: GitHub Actions with prompt regression evals
Tools: Linear (execution), Notion (prompt specs), GitHub (prompt versions)
Assessment
Objective: Validate ability to design, evaluate, and optimize a prompt system powering multi-agent tutoring.
Challenge (Candidate PoC):
Designsystem prompts for:
TutorAgent: teaching “React Hooks.”
QuizAgent: generating 5 MCQs with answers + explanations.
CodeAgent: generating runnable JS snippets.
Implementprompt chaining strategy for:
Outline → Lesson → Quiz → Mini-App.
Build aLangSmith eval pipeline:
Test against 20 golden prompts.
Measure accuracy, latency, hallucination rate, schema compliance.
Optimize:
Compare Gemini Flash vs. Pro routing for cost/latency.
Document tradeoffs and improvements.
Deliverables:
Prompt library (YAML/JSON).
Eval report with metrics table (baseline vs. optimized).
GitHub repo with eval scripts + LangSmith integration.
Notion doc summarizing prompt strategies.
5-min Loom walkthrough of the workflow.
Evaluation Criteria:
Prompt Design & Schema Compliance (30%)
Prompt Chaining & Orchestration Strategy (20%)
Evaluation Pipeline & Metrics (20%)
Performance Optimization (15%)
Documentation & Reproducibility (15%)
Prompt Engineer (Founding Team) · Gradientflo Labs