
Senior Data Engineer
- ๐ Worldwide
- Remote
- Senior
- 13 hours ago
- โฌ35 โ โฌ41 / hour
- AI
- AML
- KYC
- OCR
- RAG
- Excel
- PowerPoint
- CI/CD
- AWS
- Git
- Python
- SQL
- AWS S3
- Step Functions
- CloudWatch
- Azure Databricks
We are looking for aSenior Data / Ingestion Engineer for our client, a consultancy that helpsbanks and insurers transform their operations through AI, automation, and advanced analytics, with a strong focus onAnti-Financial Crime (fraud prevention, AML, KYC) andenterprise AI platforms. You will join their delivery for aDutch insurance client.
We're looking for an engineer with deep experience buildingrobust ingestion pipelines for unstructured documents and integratingOCR and document extraction technologies. You will turn high volumes of insurance documents (PDFs, scans, emails, Office files) into structured, high-quality, AI-ready data that feeds downstream RAG and AI models.
This is along-term remote-first contract position for candidates based in Europe.
Responsibilities
Design and implement scalable pipelines for ingesting high-volume unstructured insurance documents (PDFs, scans, emails, Word, Excel, PowerPoint).
Build connectors to document sources such as SharePoint and email.
Integrate, configure, and optimise OCR and document parsing technologies to extract high-accuracy text and layouts.
Build automated workflows for text cleaning, normalisation, semantic chunking, and metadata tagging.
Design vector storage schemas and robust retrieval mechanisms (RAG) to feed downstream AI models.
Ensure document processing pipelines meet enterprise security and low-latency SLA requirements.
Build automated error monitoring and extraction validation loops that flag low-confidence OCR outputs.
Apply engineering best practices across the pipeline lifecycle: version control, CI/CD, and testing.
Work Conditions
Start Date: ASAP
Location: Remote within Europe (CEE preferred)
Long-term contract-based role: until July 2027 with possible extension
Contract with EU LCC
5โ10 years' experience in data engineering.
Proven experience building data processing and document ingestion pipelines on public cloud platforms.
Hands-on experience processing unstructured documents (PDF, Word, Excel, PowerPoint, scans, emails).
Experience building connectors to enterprise sources such as SharePoint and email.
Practical experience with document extraction / OCR tools, e.g. AWS Textract or equivalent.
Strong engineering practices: Git, CI/CD, automated testing.
Experience in banking or insurance is an advantage.
Technical Requirements
Mandatory
Python, SQL
AWS: S3, Step Functions, CloudWatch
OCR / document extraction: AWS Textract or equivalent
Unstructured document processing and ingestion pipelines
Git, CI/CD, testing
Nice to have
Vector databases and RAG architectures
Azure, Databricks
Financial services / insurance domain experience
Senior Data Engineer ยท Jimmy Technologies