
Associate Information Data Architect (HYBRID)
- Azure
- Azure SQL
- Python
- Informatica
- SQL
- XML
- JSON
- Azure Data Factory
- Fabric
- Delta Lake
- Parquet
- ETL
- Data Modeling
- Excel
- Power BI
- JSON Schema
- Azure DevOps
- CI/CD
- PySpark
- Microsoft Fabric
- Microsoft Purview
- YAML
- Git
- SOAP
We are currently seeking an Associate Information Data Architect to join our organization in our Worcester, MA / Windsor, CT / or Michigan office, in a hybrid work arrangement.
POSITION SUMMARY:
The Associate Information Data Architect plays a key role in delivering trusted, high‑quality analytics, data products, and insights that support strategic and operational decision‑making across the organization. Working within a modern Azure-based data ecosystem (ADF, Synapse, ADLS, Azure SQL, Python), this role transforms complex data into curated datasets, automated reporting, and actionable intelligence for teams across a wide range of business domains.
The Associate Information Data Architect collaborates with business stakeholders, architects, and engineering teams to translate requirements into scalable analytical solutions; design and govern KPI logic; validate data quality; and support both legacy (Informatica) and cloud native data platforms. Success in this role requires the ability to conceptualize solutions end-to-end, innovate within evolving cloud and data environments, drive results under tight timelines, build consensus across diverse stakeholders, and communicate complex concepts clearly to both technical and business audiences.
This a full time, exempt position.
IN THIS ROLE, YOU WILL:
Analytical & Architectural Responsibilities
- Translate reporting and analytical requirements into analytical datasets, metric definitions, and reporting solutions.
- Design and document source‑to‑target mappings, transformation logic, and data lineage for cross-system integrations.
- Conduct deep‑dive data analysis to validate logic, resolve discrepancies, and support modernization or migration initiatives.
- Respond to ad hoc and urgent requests by delivering accurate, timely information in a fast-paced environment.
- Communicate complex data concepts clearly through written and verbal channels for both technical and business audiences.
- Collaborate with product owners to define and document KPIs, metric logic, and shared business rules.
Technical & Data Responsibilities
- Create and maintain curated analytics datasets (SQL views, tables, semantic layers) for dashboards and self-service analytics.
- Perform data profiling, quality checks, and reconciliation using SQL and Python to ensure accuracy and completeness.
- Parse, transform, and validate XML/JSON using Python, XPath, and schema rules; ensure consistency across systems.
- Maintain source‑to‑target mappings supporting Azure Data Factory, Synapse, and Informatica workflows.
- Improve analytics performance through query optimization, semantic modeling, and scalable dataset design.
- Support data governance through documentation, lineage tracking, and adherence to data management standards.
Cloud & Platform Responsibilities
- Collaborate with Data Engineering to translate requirements into cloud-based data pipelines (ADF, Synapse, Informatica).
- Work with architecture and engineering teams to align datasets with enterprise data models and Azure standards.
- Support Azure and Fabric monitoring (pipeline health, data freshness, validation checks) to ensure SLA adherence.
- Validate and analyze data stored in ADLS containers, ensuring correct structure and schema.
- Work with Delta Lake and Parquet files to perform integrity checks and ensure consistency across storage layers.
- Conduct reconciliation and quality verification as data moves through cloud storage zones.
Collaboration
- Contribute to tool enhancements, process improvements, and innovation initiatives.
WHAT YOU NEED TO APPLY:
- Bachelor’s degree preferred.
- Authorization to work in the United States without current or future sponsorship (U.S. citizen or lawful permanent resident).
- 2+ years of experience in a related field such as data engineering, ETL, data modeling, analysis, validation, reporting, or analytical solution delivery.
- Strong proficiency with SQL for data analysis, data profiling and quality with large and complex datasets.
- P&C insurance experience is strongly preferred, including understanding of operational and financial metrics.
- Strong analytical and problem‑solving skills.
- Advanced Excel, Power BI and other visualization tools; ability to translate data into meaningful insights.
- Excellent written and verbal communication skills.
- Strong time management, prioritization, and organizational skills.
- Knowledge of technical mapping, including database schema, XML schema, and JSON schema.
- Hands-on experience with Azure Data Factory, Synapse (SQL & Spark), ADLS, Azure SQL, and Azure DevOps (CI/CD).
- Python experience for data wrangling, XML/JSON parsing, XPath, automation, and validation.
- Experience creating semantic models and analytical datasets for BI platforms.
- Understanding of KPI design, data governance, lineage, and business glossary concepts.
- Ability to work across legacy (Informatica) and cloud platforms and validate logic during migrations.
- Ability to plan for future capability needs and continuous improvements.
PREFERRED SKILLS
- Data quality frameworks and automated testing with SQL/Python.
- Synapse Spark / PySpark; Delta Lake and notebook development.
- Experience with Microsoft Fabric (datasets, pipelines, governance).
- Exposure to Microsoft Purview or similar metadata/lineage tools.
- Experience with Parquet, partitioning, and performance optimization in lakehouse architectures.
- Exposure to Azure DevOps CI/CD (YAML, Git branching).
- Exposure to Event streaming (Event Hubs, Service Bus) and APIs/Web Services (REST/SOAP).
Associate Information Data Architect (HYBRID) · The Hanover Insurance Group