CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

Pyspark

Infosys • Bangalore, India

Job Description

Join a fast-paced, collaborative team where data powers smarter decisions and better customer experiences. In this role, you’ll work hands-on with large-scale datasets to build reliable, high-performing data processing solutions using PySpark and Spark. You’ll partner closely with engineers, analysts, and stakeholders to understand business needs, translate them into scalable pipelines, and continuously improve data quality and performance. If you enjoy solving complex data challenges, optimizing distributed workloads, and taking ownership from design to delivery, this is a great opportunity to grow your impact. You’ll be encouraged to share ideas, learn from peers, and contribute to a culture that values clarity, craftsmanship, and continuous improvement. Technical Requirements: • Primary skills:Technology->Big Data - Data Processing->PySpark Responsibilities: Key Responsibilities • Design, develop, and maintain scalable batch data pipelines using PySpark and Apache Spark for large datasets. • Perform data ingestion, transformation, and enrichment while ensuring accuracy, completeness, and consistency of outputs. • Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and improve runtime efficiency and resource utilization. • Implement robust error handling, logging, and monitoring to ensure reliable pipeline execution and faster issue resolution. • Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver well-documented solutions. • Conduct code reviews, follow engineering best practices, and contribute to reusable components and standards. • Troubleshoot production issues, perform root-cause analysis, and drive corrective and preventive actions. Preferred Skills: Technology->Big Data - Data Processing->PySpark

Job Reference ID: CF-128355 • Posted on CloudFrame Job Scanner