PySpark Developer
Job Description
We are looking for a skilled PySpark Developer with up to 3 years of experience in building and maintaining large-scale data processing pipelines. The ideal candidate should have hands-on experience in PySpark, Python, SQL, and Big Data technologies and be capable of developing efficient ETL workflows for enterprise data platforms. Technical Requirements: Strong experience in PySpark. Good programming knowledge of Python. Hands-on experience with SQL. Understanding of ETL/ELT concepts and data warehousing. Experience working with large datasets and distributed processing. Knowledge of Spark SQL, DataFrames, and Spark transformations. Familiarity with Linux/Unix environment. Responsibilities: Design, develop, and maintain data pipelines using PySpark. Develop ETL/ELT processes for ingesting, transforming, and loading large volumes of data. Write optimized PySpark code for data processing and transformation. Work with structured and semi-structured data from multiple sources. Develop and optimize SQL queries for data extraction and validation. Troubleshoot data quality and performance issues. Collaborate with Data Engineers, Analysts, and Business teams to understand requirements. Participate in code reviews and follow data engineering best practices. Monitor and support production data pipelines. Preferred Skills: Technology->Big Data - Data Processing->PySpark