Job Summary
We are seeking a highly skilled Clinical Data Engineer to join our dynamic team. The successful candidate will be responsible for designing, developing, and maintaining robust data pipelines and architectures to support clinical research and healthcare data analysis. This role offers an exciting opportunity to work with cutting-edge big data technologies and contribute to innovative health solutions. The ideal applicant will possess strong technical expertise in Big Data, database management, data warehousing, and scripting, with a keen eye for detail and analysis.
You will create monitoring tools for tracking live data out in the field that can alert the research associates to any issues, work closely with ML/AI to ensure that incoming data are stored in formats that are easily ingestible and clearly labeled, and work to align datasets with multiple incoming sources of data for analysis by our team.
Duties
- Develop, implement, and optimise scalable data pipelines using tools such as Apache Spark, Hadoop, and Informatica.
- Design and maintain efficient data warehouses and databases, including Oracle, Microsoft SQL Server, and other relational systems.
- Utilise programming languages such as Java, Python, VBA, Bash (Unix shell), and Shell Scripting to automate data processing tasks.
- Manage cloud-based data storage solutions on AWS to ensure secure and reliable data access.
- Collaborate with clinical teams to understand data requirements and translate them into technical solutions.
- Perform data validation, cleansing, and transformation processes to ensure high-quality datasets for analysis.
- Monitor system performance and troubleshoot issues related to data pipelines or storage systems.
- Document architecture designs, workflows, and procedures for compliance and knowledge sharing purposes.
- Stay abreast of emerging technologies in big data analytics and healthcare informatics to recommend improvements.
Data Engineering & Infrastructure
- Build and maintain scalable ETL pipelines using Python, SQL, and APIs to ingest and process large-scale biometric and sensor data
- Design data models and workflows that support clinical studies, internal tools, and downstream analytics
- Manage data storage, retrieval, and archival systems in AWS, including handling long-term access and data restore workflows
- Ensure data integrity, reproducibility, and proper versioning across evolving datasets and analyses
- Leverage AI-assisted tools to accelerate data analysis, debugging, and code development, improving iteration speed and reducing manual effort
Clinical Analytics & Algorithm Validation
- Analyze sleep, physiological, and behavioral datasets to evaluate product performance and validate new features
- Perform statistical analyses (e.g., correlation, error metrics, bootstrapping, validation frameworks) to assess algorithm accuracy and clinical outcomes
- Develop evaluation pipelines for metrics like HR/HRV accuracy, presence detection, and sleep staging
- Build tools and structured datasets to support training and validation of machine learning models, integrating multiple data sources for supervised learning
- Investigate edge cases, sensor issues, and data anomalies to improve model robustness
Internal Tooling & Visualization
- Maintain and extend Python-based applications for visualizing and annotating biometric data
- Develop interactive tools for researchers and engineers to inspect sessions, validate signals, and debug algorithms
- Streamline workflows for clinical teams to reduce manual effort and improve reproducibility
Cross-Functional Collaboration & Communication
- Partner with Machine Learning, Hardware, Firmware, and Product teams to build algorithms and test prototypes
- Work with Growth and Product teams to explore user behavior and inform feature development
- Synthesize findings into reports, dashboards, and presentations for internal teams and external audiences
- Contribute to abstracts, posters, and conference presentations; communicate uncertainty, methodology, and tradeoffs clearly to guide decision-making
What You’ll Need To Succeed
- 2+ years of data engineering experience with health/physiology data in a research context — you’ve built ETL pipelines around messy, real-world biometric or sensor datasets, not just clean CSVs
- Advanced Python and SQL proficiency — Pandas, NumPy, time-series analysis, and production-quality scripting are daily tools, not occasional ones
- Intermediate-to-advanced signal processing and biometric data experience — you’ve worked directly with heart rate, HRV, sleep staging, or similar physiological signals from wearable or embedded sensors
- Intermediate-to-advanced statistical modeling and validation skills — you can design and execute correlation analyses, error metrics, bootstrapping, and validation frameworks independently
- Working proficiency with AWS and Snowflake — you’ve built or maintained cloud-based data storage, retrieval, and archival systems, not just queried them
Bonus Points
- Experience with clinical or regulatory trial data, familiarity with GCP/ICH guidelines.
- Background in ML model validation or building structured training datasets for supervised learning
- Fluency with AI-assisted development tools (Claude, Cursor, ChatGPT, Copilot) as part of your daily workflow
- Domain knowledge in sleep science, biometrics, or wearable/embedded sensor data
- Experience integrating internal and third-party APIs into unified data pipelines
- Strong cross-functional communication skills — ability to translate complex analyses into clear insights for non-technical stakeholders
Skills
- Extensive experience with AWS cloud services for scalable data management.
- Proficiency in Java, Python, VBA, Bash (Unix shell), and Shell Scripting for automation tasks.
- Strong understanding of big data frameworks such as Hadoop, Apache Hive, Spark, and related ecosystems.
- Hands-on experience with relational databases, including Oracle and Microsoft SQL Server; expertise in database design is essential.
- Knowledge of data warehousing concepts and tools to support large-scale clinical datasets.
- Familiarity with Informatica or similar ETL tools for efficient data integration workflows.
- Excellent analysis skills with the ability to interpret complex datasets accurately.
- Experience working within clinical or healthcare environments is desirable but not essential.
- Strong organisational skills with the ability to manage multiple projects simultaneously while maintaining attention to detail.
This role provides an excellent platform for professionals eager to advance their careers within health informatics by leveraging their technical expertise in a meaningful context that impacts patient care outcomes positively.
Pay: £35,000.00-£60,000.00 per year
Benefits:
- Company pension
- Free parking
- On-site parking
Work Location: In person