Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
DATACLAP DIGITAL logo

AI Evaluation Lead (Part Time Contract, Engagement Based)

DATACLAP DIGITAL
  • ๐Ÿ‡ฎ๐Ÿ‡ณ India
  • On-site
  • Staff / Principal
  • 17 hours ago
  • AI
  • RLHF
  • MLOps
  • Computer Vision
  • Machine Learning
  • RAG
  • Python
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

AI Evaluation Lead (Part Time Contract, Engagement Based)

Dataclap Digital Ventures | Bangalore, Chennai or Coimbatore (remote first, occasional visits to Coimbatore)

About us

Dataclap is an AI data services company based in Coimbatore, working with AI companies across North America and Europe on data annotation, human in the loop workflows, RLHF and MLOps. We're building a dedicated AI evaluation practice covering LLMs, AI agents, computer vision and audio models, and we're looking for experienced practitioners to lead client engagements and mentor our engineering team.

The role

This is a part time, engagement based contract role, not full time employment. You'll lead AI evaluation projects for our clients, typically 8 to 15 hours a week per engagement. You'll design the evaluation approach, guide a team of junior engineers who do the hands on execution, review their work, and own the quality of what we deliver. You'll also help us turn every engagement into reusable methods, so the practice gets stronger with each project.

What you'll do

Scope evaluation projects with clients: understand the model, the use case, and what "good" looks like for them

Design evaluation frameworks, including metrics, test sets, rubrics, and pass/fail criteria

Build or guide the building of golden datasets and evaluation harnesses

Design and calibrate LLM as a judge setups, and know when human evaluation is needed instead

Plan red teaming, safety and robustness testing where relevant

Lead and mentor 2 to 5 junior engineers per engagement through regular reviews and working sessions

Review results, catch flawed methodology, and present findings to client stakeholders

Document each engagement as a playbook, rubric set or template the team can reuse

What we're looking for

6+ years in machine learning, data science or applied AI, with at least 2 years focused on model evaluation, benchmarking or quality assurance of ML systems

Hands on experience evaluating LLM applications: RAG systems, chatbots, agents or fine tuned models

Solid grounding in evaluation fundamentals: metric selection, test set design, statistical significance, inter annotator agreement, and data leakage

Experience with eval tooling such as DeepEval, Ragas, Promptfoo, Inspect, Langfuse, Braintrust or EleutherAI's evaluation harness (or custom built equivalents)

Experience leading or mentoring engineers, and comfort reviewing others' work critically

Clear written and spoken communication; you'll present findings to international clients

Python proficiency

Strong plus

Computer vision evaluation experience (detection, segmentation, tracking; metrics such as mAP and IoU; tools such as FiftyOne or CVAT)

Audio or speech model evaluation (word error rate, speaker diarization, MOS style human ratings)

Experience with AI agent evaluation: task completion, tool use, multi step trajectories

Familiarity with AI governance frameworks such as the EU AI Act or NIST AI RMF

Prior consulting or client facing delivery experience

Engagement model

Contract basis, paid per engagement, with a monthly retainer during active projects

Flexible hours, mostly remote, with occasional in person sessions in Coimbatore

Start with a short paid pilot project so we can both see if it's a fit

If you're currently employed, please make sure your employer permits outside consulting work

How to apply

Send us a short note (not just a CV) describing one model evaluation you designed or led: what you were evaluating, how you decided what to measure, and one thing that surprised you in the results.

AI Evaluation Lead (Part Time Contract, Engagement Based) ยท DATACLAP DIGITAL

Auto apply with Likeremote