Hadoop / PySpark
Job Description
Join a data-driven team where your work helps turn large-scale information into clear, actionable insights. In this role, you’ll collaborate with engineers, analysts, and stakeholders to build and optimize reliable big-data solutions that power reporting, analytics, and downstream applications. You’ll get hands-on exposure to modern distributed processing, contribute to production-grade pipelines, and learn best practices for performance, quality, and governance. If you enjoy solving complex data challenges, working in an environment that values curiosity and teamwork, and delivering measurable impact through scalable engineering, this opportunity will help you grow your technical depth while making a real difference across projects and business outcomes. Technical Requirements: echnology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop Administration->Hadoop Responsibilities: • Design, develop, and support scalable data processing workflows using Hadoop and PySpark for batch and large-volume processing. • Build and maintain data pipelines that ingest, transform, and validate data from multiple sources into curated datasets. • Optimize Spark jobs for performance (partitioning, caching, shuffle tuning) and improve overall pipeline efficiency and reliability. • Perform data quality checks, reconciliation, and root-cause analysis for pipeline failures or data anomalies. • Collaborate with cross-functional teams to understand requirements, translate them into technical solutions, and deliver within timelines. • Create clear technical documentation for workflows, data mappings, and operational runbooks. • Participate in code reviews, follow engineering best practices, and contribute to continuous improvement of standards and tooling. Preferred Skills: Technology->Big Data - Hadoop->Hadoop Administration->Hadoop,Technology->Big Data - Data Processing->PySpark