NHS England Engagement Series – Cohort 4: Linking Health and Non-Health Data – part 2 of 2

In Part 1 of this series, we looked at what it takes to earn public trust in data linkage: purpose, public benefit, transparency, safeguards, meaningful engagement and accountability. Get those right, and people are willing to say yes. But governance answers only half the question. It tells people their data will be used responsibly – it doesn’t tell them whether the linkage itself actually worked.

Participants generally understood why linking health and non-health data could be useful. Their concerns became more pronounced when considering how linked data might represent different people and communities, and what could happen when those data are used to inform research, services or decisions about individuals. These concerns matter because linkage is never perfect. Some records will be missed or linked incorrectly, and those errors may be more common for some groups than others. What matters, then, is not simply whether linkage is “accurate”, but whether it is good enough and fair enough for the purpose for which it is being used. A small amount of linkage error may have one set of consequences when estimating population trends for research, and quite another when linked data are used to plan services, identify people for support or feed into predictive systems. Public engagement therefore needs to focus not only on how linkage works, but on the questions people are more likely to care about: who might be missing, who might be misrepresented, what the consequences could be, and what is being done about it.

Linkage bias: Who gets missed or misrepresented

Linkage is never perfect. The important question for public trust is whether errors are distributed evenly. If some people or groups are systematically more difficult to link than others, the resulting linked data may represent some parts of the population better than others. The Ministry of Justice highlights two broad points in the linkage pipeline where bias can arise:

Source: Bias in Data Linking – Ministry of Justice

Bias can arise before linkage even begins. Missing, inconsistent or outdated information can make certain records harder to match; naming conventions, changes of address and differences in how information is recorded can all matter. The linkage process itself can introduce further differences, because decisions about which records are compared and what counts as sufficient evidence for a match may work better for some records than others.

Why does this matter? Because the consequences depend on what the linked data are being used for. In research, differential linkage may distort estimates or produce findings from a population that is less representative than it appears. In service planning, it may affect which communities appear to have particular needs. And where linked data are used operationally – for example, to identify individuals for support, assess risk or inform predictive systems – errors may have more direct consequences for the people affected.

That distinction also matters for public engagement. Cohort 4 participants raised particular concerns about profiling, discrimination and the possibility that some communities could be disproportionately affected when linked data are used to predict behaviour or inform decisions. Their concerns were shaped not simply by whether linkage was technically accurate, but by what the data would be used for, who could be affected if it went wrong, and what safeguards were in place.#

There is a difficult tension here. Greater transparency about linkage limitations may make some people more cautious about particular uses of their data. But avoiding those conversations is not the answer. Meaningful engagement means being open not only about the benefits of linkage, but also about uncertainty, differential impacts and what is being done to identify and reduce them.

The SPRINT workshops explored what happens when these technical issues are made accessible to non-specialist audiences. Across three UK-wide workshops, participants were introduced to linkage quality, uncertainty, bias and the potential consequences of getting linkage wrong, using plain language and practical examples. The aim was not to turn people into linkage experts, but to give them enough understanding to ask more informed questions about who may be excluded or misrepresented, how quality is assessed, and what safeguards should be expected.

The goal, then, is not to promise that linkage works perfectly or identically for everyone. It is to demonstrate that differences in linkage quality are being looked for, understood and taken seriously – and that the level of assurance and safeguards reflects what the linked data will actually be used to do.

Path to building trust and capability, for the long run

Looking at data linkage through a quality lens changes the question. It is no longer just why is my data being used? but also: who might be missing or misrepresented, what happens if the linkage is wrong, and does that matter for this particular use?

Trustworthy governance is necessary, but people also need confidence that that linkage is sufficiently accurate, representative, and appropriate for what it is being used to do.

That does not mean expecting the public to become experts in linkage methods. It means giving people enough information to understand where uncertainty and bias can arise, what their consequences might be, and how they are being assessed and managed.

Greater understanding does not necessarily produce either greater trust or greater scepticism. It can instead enable a more informed conversation about the conditions under which linkage is acceptable – including transparency about limitations, appropriate quality assurance, independent oversight and the ability to challenge how linked data are used.

And if greater transparency makes people more questioning, that is not necessarily a failure of trust. It may mean that people are better equipped to make meaningful judgements about when and how their data should be linked.

An informed public is not simply more likely to trust data linkage; it is better equipped to decide when that trust is warranted.

Learn more about SPRINT