Adur District, England
Job Summary
Data Engineer – GCP Data Products Team (IDEA Lab)
Location: UK (Edinburgh)
Lab: IDEA – Innovation, Data Engineering & Artificial Intelligence
Summary
The IDEA Lab is expanding its GCP Data Products capability and is looking for a highly skilled Data Engineer to build, optimise, and scale data pipelines and large‑scale data processing workloads in a cloud‑native environment. You will work with modern distributed data systems, contribute to data platform modernisation, and support large‑scale ingestion, transformation, and analytics workloads across the Business Transaction Banking platform.
Key Responsibilities
Key Responsibilities
Design and deliver end‑to‑end data pipelines on cloud platforms (GCP preferred).
Build scalable data ingestion, transformation, and processing workflows using distributed technologies such as Spark, Flink, Storm, or similar.
Develop robust ELT/ETL processing and migration pipelines, including support for legacy Datastage decommissioning and modernisation.
Work with a variety of database technologies including relational, NoSQL, MPP and columnar stores (BigQuery, Redshift, Azure SQLDW, HBase, MongoDB).
Implement streaming and messaging‑based pipelines using Kafka, Pulsar or Pub/Sub.
Build optimised, scalable data models to support diverse consumption patterns, applying partitioning, sharding, bucketing and aggregation strategies.
Apply performance tuning and optimisation across storage, compute and query layers.
Ensure secure handling of data including authentication, authorisation, encryption in transit/at rest, and cloud‑native security controls.
Implement monitoring, alerting and observability for large‑scale distributed data workloads.
Use orchestration tools such as Cloud Composer, Airflow or equivalent to operationalise pipelines.
Contribute to CI/CD, containerisation, Kubernetes-based deployments, and automated testing practices.
Participate in data governance, metadata, catalogue and lineage processes as needed.
Collaborate with engineers, architects and SMEs to deliver stable, high‑quality data products.
Skill Requirements
Required Skills & Experience
Strong programming skills in Java (preferred) , Python , or Scala .
Hands‑on experience with cloud data services (GCP preferred; Azure/AWS acceptable).
Practical experience with distributed data processing frameworks such as Spark (Core/SQL/Streaming), Flink, or Storm.
Experience across multiple database technologies—Relational, NoSQL, MPP, columnar.
Strong knowledge of data ingestion, transformation and messaging systems : Kafka, Pulsar, Pub/Sub, etc.
Understanding of designing scalable data models for varied access patterns.
Experience with performance tuning , cost‑optimisation and scaling strategies.
Experience delivering large‑scale big data solutions in batch and/or streaming environments, on cloud or on‑premise.
Good familiarity with the wider data ecosystem and open‑source frameworks.
Experience with orchestration (Airflow/Composer) and workflow automation.
Understanding of DevOps for data systems: CI/CD, containers, Kubernetes and automated testing.
Knowledge of security for big‑data systems including IAM, encryption, and cluster‑level controls.
Basic understanding of monitoring and alerting for distributed systems.
Knowledge of dimensional modelling (star, snowflake, normalized/denormalized).
Awareness of data governance, cataloguing and lineage tools.
Other Requirements
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-