Skip to main content

What does a Data Engineer do?

By Gemma Richards on

 A data engineer designs and builds systems that collect, store and move data. Their role involves creating pipelines that convert raw data into neat, usable formats. 

 In this blog, Gemma Richards, Data Engineer at Genomics England, shares more about her role and why it’s important for genomics. 

What does a Data Engineer do at Genomics England?

As a Data Engineer at Genomics England, I design, build and maintain the systems that move clinical data from programmes such as the 100,000 Genomes Project and the Genomic Medicine Service, into the Research Environment, where it can be accessed by approved researchers.

A large part of my role involves developing data pipelines, which are automated processes that move data from source to storage, transforming it into a usable format along the way. I also develop ETL (extract, transform, and load) workflows, which combine data from multiple sources into a large central repository. I do this using cloud technologies and orchestration tools to process data from both internal and external sources.

Alongside building new solutions, I’m also involved in testing, optimisation, quality assurance and deploying the AWS infrastructure that supports our platforms.

What I enjoy most is the variety of the role. One day I might be building a new data pipeline, and the next I could be investigating data quality issues, improving performance, or helping prepare a dataset for release. That variety means there is always something new to learn and a new challenge to tackle.

The road to data engineering

I studied Mathematics at the University of Nottingham, specialising in statistics and computational modules. Here, I discovered a real enjoyment for programming, and the logical thinking and problem-solving that comes with it.

At first, I wasn't entirely convinced I had chosen the right modules, but before long something clicked, and I realised it was exactly the kind of challenge I enjoyed. Now I get to spend my working days doing just that!

After graduating, I joined a data consultancy called Kubrick, where I completed 4 months of training in data engineering. I was then placed at Genomics England as a consultant, and spent 2 years working with the data release team before joining permanently.

In theory, I could have ended up working in almost any industry, so I feel very fortunate to have been placed at Genomics England. It has given me the opportunity to apply my technical skills to work that has a genuine impact on people's lives, which is something I find incredibly rewarding.

Impact for patients

While I don’t work directly with patients, the data and systems we build help researchers to access high-quality clinical and genomic datasets. These datasets allow them to make findings that lead to diagnoses for patients and families, as well as improving treatments.

Knowing that our work contributes, even indirectly, to better outcomes for people, is one of the most rewarding aspects of the role. Every improvement we make to the quality, accessibility and reliability of data helps researchers focus on generating insights that can advance our understanding of rare conditions and cancer.

What makes me passionate?

What makes me most passionate about my role is the combination of challenge and variety. I enjoy digging into complex problems, understanding how systems work and finding solutions that make our platforms more reliable and efficient.

There’s no feeling quite like seeing a difficult problem finally solved, or watching a complex data pipeline complete successfully after days of investigation and development! 

I'm also fortunate to work alongside talented and supportive colleagues, which makes Genomics England a great place to develop both technically and professionally.

What motivates me most is remembering that every piece of data we work with represents a real person and their healthcare journey. That responsibility drives the care we take in ensuring our data is secure, accurate and of the highest quality.

Knowing that our work helps enable research that could lead to better diagnoses, treatments and outcomes for patients gives a real sense of purpose to the challenges we tackle every day.

 What's coming up next for me?

Something particularly exciting for me now, is the work we’re doing to modernise our data release pipelines in preparation for the next generation of our Research Environment platform. Having spent more than 4 years working on different aspects of the data release process, it’s incredibly rewarding to see such significant progress and investment in this area.

The changes we’re making will help us deliver data more efficiently and at greater scale, ensuring researchers continue to have access to the high-quality data they need. It’s exciting to be involved in a period of such transformation, and to see how the work we’re doing today will support Genomics England’s future ambitions, including scaling to 500,000 genomes and beyond.

If you want to learn more about the work happening at Genomics England, check out our other blogs. You can also visit our careers page.