PL
Google Cloud Data Architect – IAM Data Modernization
PeopleNTech LLC
🇺🇸 United States
On-site
4 months ago
- GCP
- IAM
- SQL
- Machine Learning
- PySpark
- Devops
- CI/CD
- OpenShift
- AI
- Git
- ETL
- ELT
- Hadoop
- Parquet
- Apache Avro
- Pub/Sub
- Composer
- Airflow
- Dataflow
- BigQuery
- Hive
- Python
- Data Architecture
4 months ago
Role : Google Cloud Data Architect – IAM Data Modernization
Location : Dallas, TX / Charlotte, NC (Hybrid – 4 days office)
Rate: $80/hr to $85/hr
Highly Preferred OCP exp
Project/Program
Identity & Access Management (IAM) Data Modernization – migration of an on‐premises SQL data warehouse to a target‐stateData Lake on Google Cloud (GCP), enabling metrics & reporting, advanced analytics, andGenAI use cases (natural language querying, accelerated summarization, cross‐domain trend analysis)leveragingPySpark‐based processing, cloud‐native DevOps CI/CD pipelines, and containerized deployments on OpenShift (OCP) to deliver scalable, secure, and high‐performance data solutions.
About Program/Project
The IAM Data Modernization project involves migrating an on-premises SQL data warehouse to a target state Data Lake in GCP cloud environment. Key highlights include:
- Integration Scope: 30+ source system data ingestions and multiple downstream integrations
- Capabilities: Metrics, reporting, and Gen AI use cases with natural language querying, advanced pattern/trend analysis, faster summarizations, and cross-domain metric monitoring
- Benefits:
- Scalability and access to advanced cloud functionality
- Highly available and performant semantic layer with historical data support
- Unified data strategy for executive reporting, analytics, and Gen AI across cyber domains
Required Skills
DevOps / CI‐CD
- Experienceimplementing CI/CD pipelines for data and analytics workloads
- Familiarity withGit‐based source control, build automation, and deployment strategies
- Experience withOpenShift Container Platform (OCP) for deploying data workloads and services
- Understanding of containerized architecture, scaling, and environment management
- Proven ability to buildCI/CD pipelines for data and infrastructure workloads
- Experience managingsecrets securely using GCP Secret Manager
- Ownership ofobservability, SLOs, dashboards, alerts, and runbooks
- Proficiency inlogging, monitoring, and alerting for data pipelines and platform reliability
- Hands‐on experience with PySpark for ETL/ELT, data transformation, and performance optimization
- Solid understanding of distributed data processing concepts
- Strong experience designing data platforms on Google Cloud Platform (GCP)
- Experience with Data Lakes, data warehousing, and large‐scale migration programs
Data Lake Architecture & Storage
- Proven experience designing and implementingdata lake architectures (e.g., Bronze/Silver/Gold or layered models).
- Strong knowledge ofCloud Storage (GCS) design, including bucket layout, naming conventions, lifecycle policies, and access controls
- Hands-on experience withcolumnar data formats (Parquet, Avro, ORC) and compression techniques
- Expertise inpartitioning strategies, backfills, and large-scale data organization
- Ability to designdata models optimized for analytics and BI consumption
Data Ingestion & Orchestration
· Experience buildingbatch and streaming ingestion pipelines using GCP-native services
· Knowledge ofPub/Sub-based streaming architectures, event schema design, and versioning
· Strong understanding ofincremental ingestion and CDC patterns, including idempotency and deduplication
· Hands-on experience withworkflow orchestration tools (Cloud Composer / Airflow)
· Ability to design robusterror handling, replay, and backfill mechanisms
Data Processing & Transformation
· Experience developing scalablebatch and streaming pipelines using Dataflow (Apache Beam) and/or Spark (Dataproc)
· Strong proficiency inBigQuery SQL, including query optimization, partitioning, clustering, and cost control.
· Hands-on experience with HadoopMapReduce and ecosystem tools (Hive, Pig, Sqoop)
· AdvancedPython programming skills for data engineering, including testing and maintainable code design
· Experience managingschema evolution while minimizing downstream impact
Analytics & Data Serving
· Expertise inBigQuery performance optimization and data serving patterns
· Experience buildingsemantic layers and governed metrics for consistent analytics
· Familiarity withBI integration, access controls, and dashboard standards
· Understanding of data exposure patterns viaviews, APIs, or curated datasets
Data Governance, Quality & Metadata
· Experience implementingdata catalogs, metadata management, and ownership models
· Understanding ofdata lineage for auditability and troubleshooting
· Strong focus ondata quality frameworks, including validation, freshness checks, and alerting
· Experience defining and enforcingdata contracts, schemas, and SLAs
Good to have
Security, Privacy & Compliance
· Hands-on experience implementingfine-grained access controls for BigQuery and GCS
· Experience withSprint planningand helping team technically.
·Strong stakeholder communication and solution‐architecture skills
Qualifications
- Experience: [10–14]+ years in DevOps and Data Architecture, 5+ years designing on Pyspark/GCP/OCP at scale; prior on‐prem → cloud migration a must.
- Education: Bachelor's/Master's in Computer Science, Information Systems, or equivalent experience.
Google Cloud Data Architect – IAM Data Modernization · PeopleNTech LLC