RP34 - Development of Automated Detection of Reportable Events in Cancer Cases and Automated Extraction of Items for Cancer Registry Reports

Patients must be actively screened in everyday clinical practice to determine whether they meet the inclusion criteria for ongoing clinical trials. Automatic detection of patients who could be included in ongoing clinical trials at the university hospital would increase the recruitment rate of patients in clinical trials. Furthermore, automatically recognising occasions for reporting to disease registries could improve the workflow from structured data collection to submitting reports to disease registries. Nationwide, care providers are legally obligated to report cancer cases, including diagnosis, therapy and course of disease. For legal reasons, cancer registry notifications in North Rhine-Westphalia can only be made digitally. It has been regulated which variables with permitted characteristics are subject to structured reporting (nationally standardised oncological basic data set, oBDS). Hospital information systems have been optimised neither for the automatic detection of potential study patients in clinical studies nor for automated extraction of items relevant for reporting to disease registries such as cancer registries. For example, for the automated extraction in the context of reporting to cancer registries, some reportable items of the oBDS are only available in plain text, i.e. unstructured text form. Furthermore, important items for generating a cancer registry report are not necessarily at the same level or category of data storage in the hospital (e.g. laboratory information system, pathology, etc.). As a result, the unstructured or scattered data must be searched for in the hospital information system with the help of medical documentalists and then converted into a compliant structured form. This effort is sometimes high and means that the remuneration of the reporters associated with the reports may not be adequate. The situation is very similar for the collection of structured information, which requires an assessment of patients concerning the fulfilment of inclusion criteria in clinical trials. This project aims to use AI methods, e.g. NLP methods (RP02) based on transformers, to (1) recognise that it is a cancer patient or a patient eligible for an ongoing clinical trial [1], (2) recognise that there is a reason for cancer registry reporting and (3) prepare data in such a way that reporting to the cancer registry works with little manual effort. Steps (1)-(3) will initially be developed using the “cutaneous malignant melanoma” use case. Even if a major risk of this project is the potentially poor data quality, this project should explore the quality with which steps (1)-(3) are feasible with the aim of possibly achieving a reduction in the workload involved in generating cancer registry reports in the future through a combination of AI and human post-processing.


[1] H. Schäfer, A. Idrissi-Yaghir, W. Galetzka, M. Bexte, and C. M. Friedrich, “WisPerMed Text at TREC Clinical Trials Track 2021,” in Proceedings of the Thirtieth Text REtrieval Conference, TREC 2021, online, November 15-19, 2021, I. Soboroff and A. Ellis, Eds., in NIST Special Publication, vol. 500–335. National Institute of Standards and Technology (NIST), 2021. [Online]. Available: https://trec.nist.gov/pubs/trec30/papers/wispermedtxt-CT.pdf

Next
Previous