A Principal Data Engineer is responsible for expanding and optimizing data and data pipeline architecture, as well as optimizing data flow and collection for cross-functional teams. This role involves building and optimizing data systems from the ground up, ensuring that data delivery architecture is consistent throughout ongoing projects.
The Principal Data Engineer supports software developers, database architects, data analysts, and data scientists on data initiatives, ensuring optimal data delivery and architecture. This position requires expertise in data pipeline building and data wrangling, with a focus on enhancing data systems for improved efficiency and performance.
- Develop, construct, test, and maintain data architectures: responsible for creating and maintaining optimal data pipeline architectures necessary for the organization's operational efficiency.
- Collaborate with data scientists and other stakeholders: expected to collaborate with data scientists, business analysts, IT professionals, and management to assist with data-related technical issues and support their data infrastructure needs.
- Assemble and prepare large, complex data sets to meet functional business requirements: required to work with different teams to understand their data needs and ensure those needs are met.
- Build analytics tools: This will involve developing tools that utilize the data pipeline to provide actionable insights into key business performance metrics. 5%
- Maintain data security and compliance: ensure systems meet industry practices and business requirements for data privacy and protection. 5%
- Identify, design, and implement internal process improvements: This includes optimizing data delivery, automating manual processes, and re-designing infrastructure for greater scalability. 3%
- Keep up-to-date with the latest technology trends: Continually learn and adapt to new technologies that could improve existing systems 2%
- Hands on experience with programming languages (e.g. Python, SQL, GoLang)
- Knowledge with SQL database design concepts and Data Models
- Great numerical and analytical skills
- Degree in Computer Science, IT, or similar field; a Master’s is a plus
- Data engineering certification (e.g., Cloud Engineer, Data Engineer) is a plus
- Experience with big data tools: Hadoop, Spark, Kafka, etc.
- Experience with relational SQL and NoSQL databases, including MongoDG, Postgres and/or Cassandra.
- Experience with data pipeline and workflow management tools: Airflow, Dataflow, Composer, etc.
- Experience with GCP, Snowflake, Azure cloud services: BQ, Dataflow, Redshift
- Experience with stream-processing systems: Kafka, Spark-Streaming, etc.
- Good Knowledge of software engineering principles, methodologies and current best practices.
- Able to multitask and shift priorities as needed.