On June 8, 2026, the SAFETEXT working group held its third community workshop as a pre-conference event at the 9th Healthcare Text Analytics conference HealTAC 2026 in Brighton.

The workshop brought together over 80 attendees, including researchers, clinicians, industry representatives, members of the public, and TRE and information governance professionals. A key topic of the workshop was to provide feedback on the current draft of the SAFETEXT de-identification protocol, which has been prepared following the previous two community events. It followed on discussions earlier on the day on three DARE UK Next-Gen catalysts that also focused on handling free-text data within TREs.

Pre-conference workshop, Brighton 2026

SAFETEXT is a DARE UK community working group that focuses on safe and responsible access to free text-data within TREs. A special focus is on guidance for de-identification of healthcare free text, balancing between being sufficiently precise to be useful as a guide to good practice but also flexible enough to apply across different settings. Key challenges also include balancing between protecting privacy and preserving the research value of the data, and between what is technically feasible and what is operationally realistic for TREs with finite resources. 

Discussions in the third community workshop covered a wide range of themes: the governance and legal basis for de-identification across the UK’s different nations, transparency and accountability when things go wrong, and the practical implications and resource costs the protocol would place on TREs. Groups also examined the level of specificity needed in the protocol, how to handle variability in clinical free-text data and the challenges of indirect (or quasi) identifiers. Public and patient involvement in reviewing identifiable data was another key topic, alongside questions about terminology and how metadata can be made both technically useful and accessible to the public.

Participants gathered in Brighton for the third SAFETEXT community workshop

In the second session, attention turned to synthetic clinical text. While producing synthetic text is easy, participants discussed how evaluation needs depend on the purpose of the synthetic data, the current lack of mature metrics for assessing authenticity and usefulness, concerns about clinical coherence and hallucination, and the importance of carrying over governance principles – particularly around purpose and accountability – from de-identification work to synthetic text generation.

Mixed groups worked through different sections of the draft protocol
Handwritten notes as feedback

The SAFETEXT team has collected the points from discussions and presented the current state of the play at the panel at the main HealTAC conference. Chaired by Jaya Chaturvedi (from King’s College London) and featuring Arlene Casey (University of Edinburgh), Elizabeth Ford (Brighton and Sussex Medical School), Jackie Caldwell (Public Health Scotland), Sarah Markham (public contributor) and Goran Nenadic (University of Manchester), the panel addressed some of the most critical questions facing the field: do current de-identification approaches for unstructured data strike the right balance between privacy and utility? Who should approve emerging SAFETEXT protocols for them to scale and be more widely adopted? Are AI researchers being equipped with sufficient governance and ethics training?

SAFETEXT panel discussion at the main HealTAC conference

Over the coming weeks, the working group will incorporate the feedback into a revised draft. By September, the team will circulate the next version to everyone on the SAFETEXT mailing list, with a call to provide comments. The aim is to have a first community-developed guide on providing safe and responsible access to free-text data by the end of September, and present it to the wider community as a living document that describes good practice to handling sensitive free-text data within TREs.

Learn more about DARE UK Community Groups