Job Summary:
ThunderYard Solutions is seeking a Data Engineer to support the U.S. Department of Veterans Affairs in designing, developing, and maintaining scalable data solutions that support mission-critical healthcare and business operations. The ideal candidate will have expertise in data architecture, ETL/ELT processes, cloud-based data platforms, and analytics technologies, with a strong commitment to delivering secure, high-quality data solutions in a federal environment.
This role will collaborate with cross-functional teams, including data analysts, software developers, and government stakeholders, to optimize data pipelines, improve data accessibility, and ensure compliance with federal security and privacy standards. The successful candidate will demonstrate strong problem-solving skills, technical leadership, and the ability to work effectively in an agile environment supporting veteran-focused initiatives.
Required Qualifications:
- 3+ years of experience as a Data Engineer or in a similar data focused role
- Hands on experience with Databricks
- Strong experience building ETL/ELT pipelines
- Proficiency in Python and SQL
- Experience with Apache Spark / PySpark
- Experience with Azure, Azure Synapse Analytics, and Azure Data Factory
- Solid understanding of data modeling, data warehousing, and analytics use cases
- Design, develop, and maintain ETL/ELT pipelines to ingest, transform, and load data from multiple sources such as APIs, relational databases, cloud storage, and streaming platforms
- Build scalable batch and near real time data pipelines using Databricks and Apache Spark (PySpark / SQL)
- Implement data transformation logic following best practices for performance, reliability, and reusability
- Support schema evolution, data validation, deduplication, and error handling in ETL workflows
Databricks Platform Development
- Develop and optimize pipelines using Delta Lake and medallion (Bronze / Silver / Gold) architecture patterns
- Use Databricks Workflows / Jobs or similar orchestration tools to schedule and monitor pipelines
- Optimize Spark jobs for performance and cost (partitioning, caching, file sizing, query tuning)
- Collaborate on data governance initiatives using Unity Catalog, access controls, and lineage where applicable
Collaboration & Operations
- Work closely with data architects, analytics teams, and downstream consumers to define data requirements
- Troubleshoot pipeline failures and data quality issues and implement long term fixes
- Produce documentation for pipelines, datasets, and operational runbooks
- Participate in CI/CD practices using Git based version control for notebooks and code deployments
Preferred Qualifications (Nice to Have):
Preferred / Nice to Have