Data Architect
Principal Data Architect Location: Glasgow or Remote Workstyle: Hybrid Reports to: CTO About Chemify: Chemify is revolutionising chemistry.
Principal Data Architect
Location: Glasgow or Remote
Workstyle: Hybrid
Reports to: CTO
About Chemify:
Chemify is revolutionising chemistry. We are creating a future where the synthesis of previously unimaginable molecules, drugs, and materials is instantly accessible. By combining AI, robotics, and the worlds largest continually expanding database of chemical programs, we are accelerating chemical discovery to improve quality of life and extend the reach of humanity.
The Role
At Chemify, robots run real chemical experiments around the clock, and every spectrum, video frame, and sensor reading is gathered to decide what to synthesise next. We're bringing in a Principal Data Architect on a focused six-month engagement to set the architecture that turns a firehose of scientific literature and robotic lab data into a clean, governed, ML-ready ecosystem that our scientists and AI models can build on. You'll define the blueprint and stand up the foundations during the engagement, working closely with our ML and data engineering teams so they can carry it forward. We're looking for someone senior enough to make the right calls quickly and pragmatic enough to deliver real systems in the time allocated.
Scope of the Engagement
Over six months you'll own how data flows across Chemify, including how it's stored, synchronised, governed, and shared, within real regulatory and contractual limits. Concretely, you'll aim to:
Set the AI-ready foundation. Partner with our ML and data engineering teams to design ingestion and orchestration that captures scientific and operational data in a form models can use immediately. You'll define and stand up the initial data Lakehouse on AWS for the raw data coming off our robots, and design a versioned feature store that standardises chemical descriptors so there's a tight loop between lab execution and the discovery AI.
Model chemistry as data. Architect graph models for complex reaction networks, define the semantic ontologies that let AI reason across different chemical data types, and design vector search (e.g. pgvector, Pinecone) for similarity lookups across large molecular sets.
Design the telemetry pipeline. Specify streaming ingestion for high-frequency robot and sensor data, with zero loss of the "negative data" from failed experiments that's so valuable for training, plus a distributed pattern that keeps local lab "edge" data in sync with central training clusters without sacrificing consistency.
Set the governance guardrails. Define the tenancy and partitioning models that guarantee strict IP isolation for enterprise clients, secure sharing patterns for research partners, and the architectural groundwork for SOC 2 and ISO 27001 readiness.
We'll agree concrete deliverables and milestones with you up front, so success is clearly defined on both sides.
What you will bring:
Posted July 31, 2026