Our Client
Our client provides the leading professional development platform purpose-built exclusively for accounting and advisory firms. They offer the only comprehensive solution for managing learning compliance, L&D initiatives, attendance tracking, and training content. Trusted by over 130 of the top 200 firms, their platform boasts over 99% client retention by helping administrators and compliance specialists save countless hours on manual tasks and eliminate risky errors.
Requirements
Responsibilities
We are seeking a self-directed, hands-on Data Engineer to build the data foundation that powers the entire organization. Today, our data lives across 250+ transactional tables refreshed weekly. You will own consolidating this into ~10 well-defined analytic tables, transitioning from full refreshes to daily (or better) incremental loads, and standing up a centralized analytics layer that BI, product, and AI can draw from with confidence.
Responsibilities
- Own data pipeline design and implementation; consolidate 250+ transactional tables into ~10 analytic tables; move from weekly full-refresh to daily/near-real-time incremental loads via CDC (e.g., AWS DMS). Build and own the centralized analytics layer.
- Establish unified data definitions, metrics, KPIs, and benchmarks across the organization. Deliver a Master Data Management (MDM) framework with unified definitions and governance standards.
- Drive migration to the centralized layer while deprecating direct raw table access.
- Partner with internal departments to drive self-service analytics and build aggregations/data marts for benchmarking and experimentation.
- Conduct exploratory and statistical analysis for data validation, and implement data observability/monitoring.
- Partner with the Applied AI Engineer to design AI-ready data models (including retrieval- and embedding-ready structures to ground RAG and LLM applications). Expand the data lake to ingest third-party sources (e.g., HubSpot, Jira, Zendesk).
Must Haves
- 5+ years of data engineering experience building/operating production pipelines and analytics layers.
- Strong SQL and Python , with deep hands-on PostgreSQL expertise.
- Deep experience in dimensional modeling and consolidating transactional schemas into analytic tables.
- Hands-on experience with Change Data Capture (CDC) and incremental loading strategies.
- Hands-on experience with the AWS Data Stack (Athena, Glue, DMS, S3, RDS, Lambda).
- Proven experience establishing Master Data Management (MDM) , metrics definitions, and data governance.
- Hands-on experience constructing AI-ready pipelines (retrieval/embedding-ready data for RAG/LLM applications).
Nice to Haves