Scala, Spark/pyspark
Job Description
Technical Requirements: • Primary skills:Domain->Finacle-Core-Functional->Finacle-Core-WMS->Grand Master,Technology->Big Data - Data Processing->Spark,Technology->Java->Apache Responsibilities: Big Data & Spark Development Design and implement scalable data pipelines using Apache Spark (Scala and/or PySpark) Work extensively with Spark Core, Spark SQL, DataFrames, and Datasets Develop batch and real-time data processing solutions using Spark Streaming / Structured Streaming Optimize Spark jobs for performance, memory management, and parallel processing Scala & Python Development Develop robust and efficient applications using Scala and Python Write reusable, modular, and maintainable code Implement business logic and transformations on large datasets Data Engineering & ETL Build and maintain ETL/ELT pipelines for large-scale data ingestion and transformation Process structured and unstructured data from multiple sources Ensure data validation, quality, and consistency Work with file formats like Parquet, ORC, Avro, JSON, CSV Big Data Ecosystem Work with Hadoop ecosystem (HDFS, Hive, YARN) Integrate Spark jobs with data lakes and warehouses Handle large datasets with distributed computing techniques Cloud & Integration (Optional but Preferred) Work with cloud platforms (AWS/Azure/GCP) for big data solutions Utilize services such as AWS EMR, Glue, S3 / Azure Databricks / Synapse Integrate pipelines with APIs and external systems Collaboration & Leadership Collaborate with data engineers, architects, and business teams Lead technical discussions and provide guidance to junior developers Participate in code reviews and best practice implementation Work in Agile/Scrum environments Preferred Skills: Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Data Processing->Spark->SparkSQL,Technology->Java->Apache->Scala