- Pyspark Data Engineer:
- Hands-on expertise in designing, building, and maintaining Apache Spark pipelines in production environments.
- Proven experience building and scaling data ingestion frameworks that integrate data from multiple source systems , with a focus on reliability, reusability, and scalability .
- Deep understanding of Spark architecture (driver/executors, DAG, partitioning, shuffles, caching, cluster resource management) and experience operating pipelines at scale, including data transformations on datasets ~500 GB+ .
- Strong understanding of Oracle SQL and HDFS , including handling file formats and applying appropriate data cleansing, normalization, and formatting to produce curated output datasets.
- Ability to write Python , Pyspark , and shell scripts to process, transform, and automate data workflows. The Candidate should be good in writing application programs and automation manual data processing steps using python.
PySpark Developer / Senior Data Engineer
Skills
- Strong hands-on experience in PySpark, Python, and SQL.
- Experience designing and optimizing Spark-based ETL/ELT pipelines and data processing jobs.
- Strong understanding toa BigQuery.
- Strong understanding of data quality, governance, observability, and performance tuning.
- Good collaboration, debugging, and Agile delivery skills.
Experience
Bachelor’s or Master’s degree plus 6+ years of data engineering experience with strong PySpark expertise.
Pay: ₹2,200,000.00 per year
Application Question(s):
- current annual ctc
- expected annual ctc
- notice period
- linkedin profile link
- full current address with pincode
- pan card no (mandatory to upload into the portal to schedule interview, if not willing to share kindly do not apply
- Please share the candidate availability for interview (Date & Time) between 10 am to 5pm
Work Location: Hybrid remote in Chennai, Tamil Nadu