Java Spark developer

CitiChennai, Tamil NaduOn-siteFull-timeJunior, 1–2 yearsListed 7 hours ago

Apply now

About this role

We are seeking an experienced and motivated Java Spark Developer to design, develop, and maintain high-performance, large-scale data processing pipelines and distributed applications. In this role, you will leverage Core Java and Apache Spark to build resilient batch and real-time streaming data architectures, optimize distributed data workloads, and collaborate with cross-functional teams including Data Scientists, Cloud Engineers, and Solution Architects.

1. Data Pipeline & Application Development

- Design, implement, and maintain robust, scalable data ingestion and ETL/ELT pipelines using Apache Spark (Core, SQL, Streaming) written in Java (or Scala interoperability).
- Develop performant, low-latency microservices and distributed processing modules integrated with messaging platforms (e.g., Apache Kafka).
- Build and maintain interfaces to relational databases, distributed data lakes, and NoSQL stores (e.g., Hive, Cassandra, HBase, MongoDB, Delta Lake, Snowflake).

2. Performance Tuning & Optimization

- Profile, debug, and optimize Spark jobs by managing partitioning strategies, caching, broadcast variables, memory allocation (driver/executor memory), and data serialization (Kryo).
- Analyze query execution plans, DAGs, and Spark UI metrics to eliminate data skew, reduce shuffle overhead, and minimize bottleneck latencies.
- Monitor resource utilization on cluster managers such as Kubernetes, Apache YARN, or cloud-native orchestration engines.

3. Architecture & Data Modeling

- Design structured, semi-structured, and unstructured data storage schemas using columnar file formats (e.g., Parquet, ORC, Avro).
- Implement robust data validation, cleansing, data governance, and error-handling mechanisms across the ingestion lifecycle.
- Ensure data privacy and enterprise compliance by applying encryption at rest/transit and access-control policies.

4. Collaboration, CI/CD & Best Practices

- Participate in Agile/Scrum ceremonies, sprint planning, and code reviews to ensure adherence to high code quality standards.
- Write comprehensive unit, integration, and automated regression tests using frameworks such as JUnit, Mockito, and Spark Testing Base.
- Configure and maintain continuous integration and continuous deployment (CI/CD) pipelines using tools like Jenkins, GitLab CI, or GitHub Actions.

Required Qualifications & Skills

Technical Competencies

- Core Java: Deep proficiency in Java (Java 8/11/17+), including multithreading, concurrency, OOP principles, memory management, and JVM internals.
- Apache Spark: Hands-on experience developing distributed applications with Apache Spark (RDDs, DataFrames, Datasets, Spark SQL, Spark Structured Streaming).
- Distributed Ecosystem: Strong working knowledge of distributed architecture (HDFS, YARN), Hive, and distributed storage systems.
- Messaging & Streaming: Practical experience with event streaming platforms such as Apache Kafka or RabbitMQ.
- Database & Query Languages: Advanced SQL capabilities, experience with relational databases (PostgreSQL, Oracle, MySQL) and NoSQL datastores.
- Build & Version Control: Proficiency with build tools (Maven, Gradle) and Git version control workflows.
- Testing: Solid track record in Test-Driven Development (TDD) using JUnit, Mockito, and distributed testing patterns.

Professional Experience & Education

- Education: Bachelor’s or Master’s degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
- Experience: 3-6 years of professional software engineering experience, with at least 2–4 years dedicated to building scalable distributed data processing applications using Java and Apache Spark.

Preferred / Desired Qualifications

- Cloud Platforms: Experience building and deploying data architectures on AWS (EMR, S3, Glue, Athena), Azure (Databricks, HDInsight, ADLS), or Google Cloud (Dataproc, BigQuery).
- Modern Lakehouse Technologies: Hands-on exposure to Apache Iceberg, Delta Lake, or Apache Hudi.
- Containerization & Orchestration: Familiarity with Docker, Kubernetes, and workflow schedulers like Apache Airflow or Luigi.
- Polyglot Exposure: Familiarity with Scala or Python (PySpark) is an added advantage.

------------------------------------------------------

## Job Family Group:
Technology
------------------------------------------------------

## Job Family:
Applications Development
------------------------------------------------------

## Time Type:
Full time
------------------------------------------------------

## Most Relevant Skills
Please see the requirements listed above.
------------------------------------------------------

## Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.
------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi .

View Citi’s EEO Policy Statement and the Know Your Rights poster.