Job Summary
We are seeking a Senior Data Engineer with strong hands-on experience in Databricks, Spark SQL, Delta Lake, and AWS-based data lake architectures.
The successful candidate will support data warehouse migration and consolidation initiatives, modernize legacy SQL and PL/SQL workloads, and build scalable data pipelines and data lake solutions using Databricks, Amazon S3, and Delta Lake.
Mandate Skills
- Databricks , including Spark SQL and Delta Lake implementations.
- Data lake architectures on Databricks
- data Modeling and warehouse design principles
Key Responsibilities
- Analyze existing data warehouse solutions to support migration, modernization, and consolidation initiatives.
- Reverse-engineer legacy SQL and PL/SQL stored procedures and convert business logic into scalable Spark SQL code.
- Design and develop AWS data lake solutions using Amazon S3, Databricks, and Delta Lake.
- Build and maintain reliable ETL and data-ingestion pipelines for processing and transformation in Databricks.
- Collaborate with data architects to implement standardized ingestion and transformation frameworks.
- Design and optimize star, snowflake, and flattened data models for performance and scalability.
- Troubleshoot data lake, pipeline, job-scheduling, and performance issues.
- Perform foundational data administration activities, including job monitoring, error handling, tuning, and backup coordination.
- Document data flows, ETL processes, transformation rules, and technical designs.
- Maintain source code and deployment artifacts using Git-based repositories.
- Work with cross-functional teams to integrate data sources into a unified data platform.
- Participate in Agile ceremonies, including sprint planning, backlog grooming, reviews, and retrospectives.
Required Qualifications
- At least 5 years of hands-on experience working with Databricks, Spark SQL, and Delta Lake.
- At least 3 years of experience designing and implementing data lake architectures on Databricks.
- Strong SQL and PL/SQL skills, including experience interpreting and refactoring legacy stored procedures.
- Hands-on experience with dimensional modeling and data warehouse design principles.
- Experience building data pipelines and ETL solutions using Databricks and Amazon S3.
- Proficiency in at least one programming language: Python, Scala, or Java.
- Experience with performance tuning, troubleshooting, job scheduling, and production support.
- Experience using Git or similar source-control platforms.
- Experience working in Agile development environments.
- Bachelor’s degree in Computer Science, Information Technology, Data Engineering, or a related field.
Preferred Qualifications
- Databricks cloud certification.