PL
Data Architect β OCP (OpenShift) | IAM Data Modernization
PeopleNTech LLC
πΊπΈ United States
On-site
4 months ago
- OpenShift
- IAM
- SQL
- Machine Learning
- PySpark
- CI/CD
- GCP
- Data Architecture
- Airflow
- Devops
- ETL
- ELT
- Kubernetes
- Parquet
- Apache Avro
- BigQuery
- Dataflow
- Agile
4 months ago
Role :Data Architect β OCP (OpenShift) | IAM Data Modernization
Location : Dallas, TX / Charlotte, NC (Hybrid β 3 days office)
Rate: $85/hr to $90/hr
Project/Program
Identity & Access Management (IAM) Data Modernization
Migration of an on-premises SQL data warehouse to a modern enterprise Data Lake platform, enabling analytics and GenAI use cases. The platform leveragesPySpark-based processing, CI/CD pipelines, and containerized deployments on OpenShift (OCP), withGCP as a preferred cloud platform, to deliver scalable, secure, and high-performance data solutions
About Program/Project
The IAM Data Modernization program focuses on transforming legacy data platforms into a scalable and cloud-compatible architecture.
Key Highlights:
- Integration Scope: 30+ source systems with multiple downstream integrations[
- Capabilities: Metrics, reporting, advanced analytics, and GenAI use cases (NL querying, summarisation, cross-domain insights)
- Benefits:
- Scalable and resilient data platform
- High-performance semantic and analytics layer
- Single source of truth for enterprise-wide reporting and analytics
We are looking for aData Architect with strong expertise in OpenShift (OCP), PySpark, and CI/CD pipelines to design and govern scalable data platforms.
The role requires definingend-to-end data architecture, containerised deployment patterns, orchestration strategies (Airflow/Autosys), and platform standards, along with hands-on involvement in implementation.
Key Responsibilities
Data Architecture & Platform Design
- Defineenterprise data architecture for IAM data lake and analytics platform
- Designscalable, modular, and containerised data pipeline architectures on OCP
- Establishdata models, schema governance, and data lifecycle strategies
- Define best practices fordata partitioning, performance optimisation, and cost efficiency
- Architect and governcontainerised data workloads on OpenShift (OCP)
- Define standards fordeployment, scaling, and workload isolation
- Collaborate with DevOps teams forplatform engineering and infrastructure alignment
- Define architecture forPySpark-based batch and near real-time processing pipelines
- Provide guidance ondistributed processing design, optimisation, and performance tuning
- Establish reusable frameworks forETL/ELT processing
- Architectdata ingestion frameworks (batch, streaming, CDC)
- Define orchestration strategies usingAirflow / Autosys
- Implement standards forretry, backfills, dependency management, and error handling
- Define and overseeCI/CD strategy for data and platform deployments
- Enableautomation of build, test, and deployment processes
- Ensure integration of CI/CD pipelines withOCP-based environments
- Provide architecture guidance forGCP-based data platforms (preferred, not mandatory)
- Define integration patterns forcloud-native and on-premise hybrid environments
- Guide teams oncloud migration strategies and modern data platform adoption
- Define frameworks for:
- Data quality, validation, and lineage
- Metadata management and cataloguing
- Establishmonitoring, logging, alerting, and SLOs for platform reliability
- Ensure compliance withdata security and audit requirements
- Work closely withclient architects, IAM teams, and business stakeholders
- Translate business requirements intoscalable technical architecture
- Provide architectural guidance and mentorship to engineering teams
Core Skills (Must Have)
- Strong experience in:
- OpenShift (OCP) / Kubernetes-based platforms
- PySpark / Spark ecosystem
- CI/CD implementation for data platforms
- Airflow / Autosys orchestration tools
- Solid understanding of:
- Data lake architectures (layered models)
- ETL/ELT design patterns
- Distributed data processing concepts
- Expertise in:
- Data formats: Parquet, ORC, Avro
- Partitioning and performance tuning
- Large-scale data modelling for analytics
- Experience withGoogle Cloud Platform (GCP) (preferred)
- Exposure to services like BigQuery, Dataproc, Dataflow, GCS is a plus
- Experience defining:
- Monitoring, logging, alerting frameworks
- Dashboards, SLOs, and operational runbooks
- Experience withIAM domain / cybersecurity data
- Understanding ofdata security and access control frameworks
- Exposure toGenAI-enabled data platforms
- Experience inAgile delivery and team leadership
- Experience:
- 10β14+ years in Data Architecture / Data Engineering
- Strong experience in OCP, PySpark, CI/CD, and orchestration frameworks
- Prior experience indata modernisation / migration programs
- Education:
Bachelor's/Master's in Computer Science, Information Systems, or equivalent - Certifications (Preferred):
- OpenShift / Kubernetes certifications
- GCP certifications (preferred, not mandatory)
==========
This e-mail may contain privileged and confidential information which is the property of Persistent Systems Ltd. It is intended only for the use of the individual or entity to which it is addressed. If you are not the intended recipient, you are not authorized to read, retain, copy, print, distribute or use this message. If you have received this communication in error, please notify the sender and delete all copies of this message. Persistent Systems Ltd. does not accept any liability for virus infected mails.
Data Architect β OCP (OpenShift) | IAM Data Modernization Β· PeopleNTech LLC