Hadoop / PySpark
Job Description
Step into a high-impact consulting role where you’ll lead the design and delivery of scalable big data solutions that turn complex datasets into actionable insights. As a Lead Consultant, you’ll collaborate closely with data engineers, architects, analysts, and business stakeholders to shape modern data platforms built on Hadoop ecosystems and PySpark-based processing. You’ll guide teams through best practices, performance tuning, and reliable delivery—balancing hands-on problem solving with technical leadership. This role is ideal for someone who enjoys mentoring, driving technical decisions, and building data pipelines that are resilient, efficient, and production-ready. If you’re motivated by solving large-scale data challenges and enabling teams to deliver measurable outcomes, you’ll find a collaborative environment here that values ownership, clarity, and continuous improvement. Technical Requirements: • Primary skills:Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop Responsibilities: • Lead end-to-end delivery of big data solutions using Hadoop and PySpark, from requirements to production rollout. • Design and implement scalable batch/ETL pipelines for large datasets with strong focus on reliability and performance. • Drive technical architecture discussions, define standards, and ensure best practices for distributed data processing. • Optimize Spark jobs through partitioning strategies, caching, shuffle tuning, and efficient file formats where applicable. • Collaborate with stakeholders to translate business needs into technical specifications, delivery plans, and milestones. • Establish data quality checks, validation frameworks, and operational monitoring for production pipelines. • Perform root-cause analysis for pipeline failures and performance bottlenecks; implement preventive fixes. • Mentor engineers through code reviews, design reviews, and knowledge sharing to uplift team capability. • Ensure documentation, runbooks, and handover artifacts are created and maintained for support readiness. Preferred Skills: Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Hadoop->Hadoop