Senior Data & Document Ingestion Engineer (OCR / RAG)
from
About Gramian
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
About the Role
Our client is a Big4 Consultancy group that works with leading financial institutions onAI-driven transformation, automation, advanced analytics, and financial crime prevention. Their work spans intelligent fraud detection, AML/KYC modernization, autonomous workflows, enterprise AI platforms, and the secure industrialization of AI in highly regulated environments.
We are looking for aSenior Data & Document Ingestion Engineer to build robust pipelines for processing high-volume unstructured insurance content. The role focuses onOCR, document parsing, ingestion pipelines, text normalization, semantic chunking, metadata extraction, and retrieval-ready data preparation for downstream AI systems.
CONTRACT: Contractor assignment, expected October 2026 – July 2027, with potential extension
COMMITMENT: Full-time
LOCATIONS: Europe-based, preferably CEE; remote, with potential future hybrid work in Prague
PROCESS: Initial qualification followed by technical and client interviews
NOTES: Fluent English is required.
Responsibilities
- Design and build scalabledocument ingestion pipelines for PDFs, scans, emails, and office documents.
- Integrate and optimizeOCR and document extraction technologies for high-accuracy text and layout extraction.
- Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging.
- Process unstructured formats including PDF, Word, Excel, and PowerPoint.
- Develop connectors for enterprise sources such asSharePoint and email systems.
- Design data schemas and retrieval mechanisms for downstream AI andRAG use cases.
- Build validation and monitoring loops to detect low-confidence OCR or extraction results.
- Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements.
- Implement logging, testing, and operational monitoring across data-processing workflows.
- Apply Git, CI/CD, and software-engineering best practices to pipeline development.
- Approximately5–10 years of professional data engineering or backend/data-platform experience.
- Strong hands-on experience withPython and SQL.
- Proven experience buildingdata ingestion and document-processing pipelines.
- Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails.
- Experience withOCR/document extraction tools such as AWS Textract or equivalent.
- Professional experience building data-processing pipelines onpublic cloud platforms.
- Experience with AWS services such asS3, Step Functions, and CloudWatch, or comparable cloud services.
- Strong development practices includingGit, CI/CD, and automated testing.
Preferred Qualifications
- Experience withAzure, AWS, or Databricks in enterprise data environments.
- Experience withvector databases, embeddings, or RAG architectures.
- Experience designing connectors to SharePoint, email, or other enterprise content systems.
- Background in insurance, financial services, or regulated-data environments.