
Principal Software Engineer
- Large Language Models
- Microservices
- Apache Spark
- PySpark
- Databricks
- GCP
- MongoDB
- ETL
- Apache Airflow
- Composer
- Data Modeling
- SQL
- NoSQL
- Azure
- AWS
- Python
- TDD
- LangChain
- LangGraph
JOB TITLE: Senior Software Engineer
LOCATION: Southlake, TX (hybrid role, may work from home)
DUTIES: Develop automation tooling using LLMs (Large Language Models) and agent frameworks to assist with legacy code analysis, documentation generation, and migration/pipeline scaffolding; integrate outputs with engineering review, validation, and testing; Analyze and reverse-engineer legacy database and mainframe workloads to identify dependencies, data flows, and functional logic required for modernization and migration; Develop parsers and analysis utilities to extract metadata, lineage, and relationships from legacy codebases to support migration planning, documentation, and implementation; Design and build cloud-native migration services to convert legacy procedures and batch logic into modern microservices and standardized data processing jobs; Design and implement scalable data platforms and pipelines to support enterprise eligibility and operational data processing using Apache Spark (PySpark) and Databricks/Google Cloud Dataproc; Architect data models and storage patterns; redesign relational schemas into denormalized nested document models suitable for MongoDB to improve downstream query efficiency and application performance; Build configuration-driven Extract, Transform, Load (ETL) frameworks that generate Spark jobs from declarative specifications; Develop data validation, auditing, and threshold-based control frameworks to detect data discrepancies and enforce quality gates across pipeline stages; Orchestrate and monitor data workflows using Apache Airflow (Google Cloud Composer), including dependency management, retries, alerts, and operational controls; Implement Change Data Capture and reverse ETL mechanisms to synchronize changes from MongoDB to data lake storage in near real time for downstream analytics and reporting.
REQUIREMENTS: Bachelor’s or foreign equivalent degree in Computer Science, Computer or Electronic Engineering, or a related field, and 5 years of progressive, post-baccalaureate experience in the job offered or as a Software Engineer/Developer, Data Engineer, Programmer Analyst, or in a related/similar position. Experience therein to include 5 years in data engineering using Apache Spark (PySpark) and Databricks or Google Cloud Dataproc, databases and data modeling using SQL, NoSQL or MongoDB, schema design and optimization, cloud services such as Microsoft Azure, GCP or AWS for data/compute, and Python software development with Agentic AI, data processing, backend services, and TDD; and 2 years with Large Language Models (LLMs), LangChain and LangGraph agent frameworks. Hybrid role, ability to work from home.
CONTACT: To apply, email resume to anna.zhang@peerislands.io. Ref. Job code KAMN-W.
Principal Software Engineer · PeerIslands