
Computer Vision Developer
- AI
- RLHF
- MLOps
- Computer Vision
- OCR
- ONNX
- TensorRT
- Devops
- Python
- PyTorch
- TensorFlow
- OpenCV
- Docker
- Weights & Biases
- MLflow
- CI/CD
- AWS
- GCP
- Azure
About the role
Dataclap builds AI data pipelines ā annotation, RLHF, human-in-the-loop, and MLOps ā for clients across North America and Europe. We're looking for a Computer Vision Developer who can go beyond running a single off-the-shelf model: someone who has stitched multiple models into working pipelines, and who has trained or fine-tuned models rather than only consuming APIs.
You'll design and build vision systems that power model-assisted labeling, automated QA of annotations, and delivery pipelines for our clients' datasets. This is a hands-on engineering role ā you'll own problems end to end, from data and model selection through to deployment and monitoring.
What you'll do
⢠Design and build multi-model computer vision pipelines (e.g. detection ā tracking ā segmentation ā classification / OCR) that run reliably at scale..
⢠Fine-tune and train custom models from open-source checkpoints to hit client-specific accuracy and edge-case requirements..
⢠Evaluate and benchmark candidate models, select the right architecture for each task, and document trade-offs (accuracy, latency, cost)..
⢠Build model-assisted labeling and auto-QA tooling to accelerate our annotation and HITL workflows..
⢠Handle the full lifecycle: data preparation, augmentation, training, validation, error analysis, and iteration..
⢠Optimize models for inference ā quantization, ONNX/TensorRT export, batching ā and package them for deployment..
⢠Set up experiment tracking, versioning, and reproducible training runs..
⢠Collaborate with annotation, delivery, and DevOps teams, and communicate results clearly to non-ML stakeholders and clients..
Required qualifications
⢠3+ years of total software/ML engineering experience, with at least 1ā2 years working specifically in computer vision..
⢠Hands-on experience with multiple computer vision models across different task families ā not just one. For example: object detection (YOLO family, Faster R-CNN, DETR), segmentation (SAM, Mask R-CNN, U-Net), classification (ResNet, EfficientNet, ViT), plus any of OCR, pose estimation, or object tracking (ByteTrack, DeepSORT)..
⢠Demonstrated experience building pipelines that chain multiple models together ā feeding the output of one model into another, with proper pre/post-processing between stages..
⢠Proven experience training custom models or fine-tuning from open-source models, including preparing datasets, running training, and doing error analysis (please be ready to walk us through a specific example)..
⢠Strong Python and solid experience with PyTorch and/or TensorFlow and OpenCV..
⢠Comfort with the data side: dataset curation, augmentation, handling class imbalance, and evaluating with the right metrics (mAP, IoU, precision/recall, F1)..
⢠Ability to read a recent CV paper or model repo and get it running..
Nice to have
⢠Experience with vision-language / multimodal models (VLMs) or vision components for VLA / robotics training data..
⢠Inference optimization and edge deployment experience (ONNX, TensorRT, quantization, distillation)..
⢠MLOps exposure: Docker, experiment tracking (Weights & Biases, MLflow), model versioning, CI/CD for models..
⢠Familiarity with annotation platforms and data-labeling workflows (CVAT, Label Studio, or similar)..
⢠Cloud experience (AWS / GCP / Azure) for training and serving..
⢠Experience working with or delivering to international clients..
Who this role suits
You'll do well here if you're genuinely curious about models ā the kind of person who benchmarks three architectures before picking one, who reads the eval numbers critically, and who has actually broken and fixed a training run. If your CV experience is limited to calling a single pre-trained model through an API, this role will likely stretch you beyond what it's asking for.
Computer Vision Developer Ā· DATACLAP DIGITAL