
Java Spark developer
- Java
- Apache Spark
- ETL
- ELT
- SQL
- Scala
- Microservices
- Apache Kafka
- NoSQL
- Hive
- Cassandra
- HBase
- MongoDB
- Delta Lake
- Snowflake
- Kubernetes
- Apache
- Yarn
- Data Modeling
- Parquet
- Apache Avro
- CI/CD
- Agile
- Scrum
- JUnit
- Mockito
- Jenkins
- GitLab CI
- GitHub Actions
- OOP
- JVM
- RabbitMQ
- PostgreSQL
- Oracle
- MySQL
- Maven
- Gradle
- Git
- Test-Driven Development
- TDD
- AWS
- EMR
- Athena
- Azure Databricks
- GCP
- BigQuery
- Apache Iceberg
- Apache Hudi
- Docker
- Apache Airflow
- Python
- PySpark
We are seeking an experienced and motivatedJava Spark Developer to design, develop, and maintain high-performance, large-scale data processing pipelines and distributed applications. In this role, you will leverageCore Java andApache Spark to build resilient batch and real-time streaming data architectures, optimize distributed data workloads, and collaborate with cross-functional teams including Data Scientists, Cloud Engineers, and Solution Architects.
1. Data Pipeline & Application Development
- Design, implement, and maintain robust, scalable data ingestion and ETL/ELT pipelines usingApache Spark (Core, SQL, Streaming) written inJava (or Scala interoperability).
- Develop performant, low-latency microservices and distributed processing modules integrated with messaging platforms (e.g., Apache Kafka).
- Build and maintain interfaces to relational databases, distributed data lakes, and NoSQL stores (e.g., Hive, Cassandra, HBase, MongoDB, Delta Lake, Snowflake).
2. Performance Tuning & Optimization
- Profile, debug, and optimize Spark jobs by managing partitioning strategies, caching, broadcast variables, memory allocation (driver/executor memory), and data serialization (Kryo).
- Analyze query execution plans, DAGs, and Spark UI metrics to eliminate data skew, reduce shuffle overhead, and minimize bottleneck latencies.
- Monitor resource utilization on cluster managers such as Kubernetes, Apache YARN, or cloud-native orchestration engines.
3. Architecture & Data Modeling
- Design structured, semi-structured, and unstructured data storage schemas using columnar file formats (e.g., Parquet, ORC, Avro).
- Implement robust data validation, cleansing, data governance, and error-handling mechanisms across the ingestion lifecycle.
- Ensure data privacy and enterprise compliance by applying encryption at rest/transit and access-control policies.
4. Collaboration, CI/CD & Best Practices
- Participate in Agile/Scrum ceremonies, sprint planning, and code reviews to ensure adherence to high code quality standards.
- Write comprehensive unit, integration, and automated regression tests using frameworks such as JUnit, Mockito, and Spark Testing Base.
- Configure and maintain continuous integration and continuous deployment (CI/CD) pipelines using tools like Jenkins, GitLab CI, or GitHub Actions.
Required Qualifications & Skills
Technical Competencies
- Core Java: Deep proficiency in Java (Java 8/11/17+), including multithreading, concurrency, OOP principles, memory management, and JVM internals.
- Apache Spark: Hands-on experience developing distributed applications with Apache Spark (RDDs, DataFrames, Datasets, Spark SQL, Spark Structured Streaming).
- Distributed Ecosystem: Strong working knowledge of distributed architecture (HDFS, YARN), Hive, and distributed storage systems.
- Messaging & Streaming: Practical experience with event streaming platforms such asApache Kafka or RabbitMQ.
- Database & Query Languages: Advanced SQL capabilities, experience with relational databases (PostgreSQL, Oracle, MySQL) and NoSQL datastores.
- Build & Version Control: Proficiency with build tools (Maven, Gradle) and Git version control workflows.
- Testing: Solid track record in Test-Driven Development (TDD) using JUnit, Mockito, and distributed testing patterns.
Professional Experience & Education
- Education: Bachelor’s or Master’s degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
- Experience: 3-6 years of professional software engineering experience, with at least 2–4 years dedicated to building scalable distributed data processing applications using Java and Apache Spark.
Preferred / Desired Qualifications
- Cloud Platforms: Experience building and deploying data architectures on AWS (EMR, S3, Glue, Athena), Azure (Databricks, HDInsight, ADLS), or Google Cloud (Dataproc, BigQuery).
- Modern Lakehouse Technologies: Hands-on exposure to Apache Iceberg, Delta Lake, or Apache Hudi.
- Containerization & Orchestration: Familiarity with Docker, Kubernetes, and workflow schedulers like Apache Airflow or Luigi.
- Polyglot Exposure: Familiarity with Scala or Python (PySpark) is an added advantage.
------------------------------------------------------
Job Family Group:
Technology------------------------------------------------------
Job Family:
Applications Development------------------------------------------------------
Time Type:
Full time------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
Â
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity reviewAccessibility at Citi.
View Citi’sEEO Policy Statement and theKnow Your Rights poster.
Java Spark developer · Citibank, N.A. United Kingdom