Data & Knowledge Engineer
- AI
- SQL
- Python
- Data Modeling
- Vector Search
- Microsoft Fabric
- Azure Data Factory
- Databricks
- Snowflake
- OCR
- Azure AI
- PostgreSQL
- pgvector
- Elasticsearch
- Pinecone
- Weaviate
- Milvus
- Neo4j
- Risk Management
Job Description & Summary
The opportunity
Provide trusted, contextual and well-governed enterprise data and knowledge services that ground agentic workflows and improve their reliability.
What you will be doing
·       Design and build ingestion, transformation and serving pipelines for structured and unstructured data.
·       Create retrieval indexes, metadata models, semantic layers, knowledge graphs or data products as appropriate.
·       Implement chunking, enrichment, lineage, quality and access-control patterns.
·       Optimize retrieval quality, freshness, latency and cost with the AI engineering team.
·       Integrate cloud and on-premises data sources for hybrid solutions.
·       Support evaluation datasets, monitoring data and traceability requirements.
What we need from you
·       4+ years in data engineering, analytics engineering, information retrieval or knowledge platforms.
·       Strong SQL and Python skills and experience with data pipelines, APIs and data modeling.
·       Practical knowledge of vector search, embeddings, metadata, document processing and retrieval evaluation.
·       Experience with enterprise security, data quality and hybrid data integration.
Relevant AI technologies and tooling
·       Strong SQL and Python capability with practical experience in Spark and data engineering platforms such as Microsoft Fabric, Azure Data Factory, Databricks, Snowflake or equivalent.
·       Hands-on experience processing structured and unstructured content, including parsing, OCR, chunking, enrichment, metadata extraction, lineage and incremental indexing.
·       Experience with vector and hybrid search technologies such as Azure AI Search, PostgreSQL with pgvector, Elasticsearch, Pinecone, Weaviate, Milvus or equivalent.
·       Understanding of embedding selection, semantic and lexical retrieval, metadata filtering, reranking, query transformation, evaluation datasets and retrieval quality metrics.
·       Experience with graph and knowledge technologies such as Neo4j, RDF or property graphs, ontologies, entity resolution and GraphRAG patterns is desirable.
·       Ability to implement secure hybrid data access, row or document-level permissions, data masking and traceable ingestion from cloud and on-premises repositories.
Measures of success
·       Data freshness, quality and availability
·       Retrieval relevance and traceability
·       Speed of onboarding new knowledge sources
·       Pipeline reliability and performance
·       Compliance with data-access requirements
Key interfaces
·       Other members of the AI Transformation & Agentic Systems Practice
·       PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists
·       Client business owners, product owners, technology teams and operational users
·       Technology alliance and implementation partners where relevant
Contribution to the practice
·       Support proposals, client workshops and market development appropriate to seniority.
·       Contribute reusable methods, patterns, code, assets and lessons learned.
·       Coach colleagues and participate in the capability’s continuous learning agenda.
·       Uphold PwC quality, independence, confidentiality and risk-management requirements.
#LI-BS1 #LI-HybridÂ
Data & Knowledge Engineer · wd3:pwc:Global_Experienced_Careers