
This project builds upon the multicentric clinical database and digital pathology infrastructure established during previous project phases through the collaboration between AZ Delta/RADar Learning & Innovation Center and UZ Leuven. The project focuses on developing and validating a large language model (LLM)-based pipeline for automated extraction of structured clinical data from routine medical records. By transforming unstructured clinical documentation into high-quality research datasets, it provides a scalable data infrastructure that will facilitate future multimodal AI research combining clinical, radiological and pathological data for personalised prognostic prediction in colon cancer.
Introduction
High-quality clinical datasets are essential for developing and validating artificial intelligence (AI) models in oncology. However, extracting structured clinical information from electronic health records remains a labour-intensive and time-consuming process. This project aims to overcome this challenge by developing a large language model (LLM)-based pipeline capable of automatically extracting and structuring clinical, pathological and radiological information from routine medical records. Building on the manually curated multicentric database established during the first phases of this research programme, the project seeks to accelerate future AI development through automated, reproducible database generation.
Objective
The primary objective of this project is to develop and validate an LLM-based pipeline capable of automatically generating high-quality structured clinical datasets from heterogeneous medical records in the domain of colon cancer.
- Developing an LLM tailored to colorectal cancer terminology and clinical documentation.
- Automatically extracting structured clinical, pathology and radiology variables from unstructured medical reports in the domain of colon cancer.
- Benchmarking the LLM-generated database against the manually curated multicentric database developed by AZ Delta and UZ Leuven.
- Establishing a scalable data-generation workflow that supports future multimodal AI research and prospective clinical studies.
Methodology
The project uses retrospective clinical documentation, pathology reports and radiology descriptions collected from participating Belgian hospitals. De-identified reports are used to train and fine-tune an LLM capable of recognising, extracting and harmonising clinically relevant information across different reporting styles and institutions.
The extracted variables are automatically converted into a structured research database and validated against the existing manually curated gold-standard dataset. Performance is evaluated based on extraction accuracy, completeness, reproducibility and robustness across participating centres. The resulting pipeline provides the data foundation for future multimodal AI models integrating clinical, radiological and pathological information to improve prognostic prediction in patients with colon cancer.
Impact and future directions
By automating one of the most labour-intensive steps in clinical AI research, this project has the potential to substantially accelerate the development of high-quality multicentric datasets while improving standardisation and reproducibility.
General info and contact
RADar project research lead: Dr. Liesbeth De Bruecker
RADar project researchers: M.Sc. Louise Berteloot, Ir. Dimitrios Daskalakis, Prof. Dr. Ir. Peter De Jaeger, Dr. Nathalie Mertens
Principal investigator: Prof. MD Jeroen Dekervel (UZ Leuven)
Site investigators: MD Sofie De Meulder (AZ Delta), MD Pieter-Jan Cuyle (Imeldaziekenhuis Bonheiden), MD Antoon Billiet (AZ Oostende)
Timeline: 2026-2027
Status: Ongoing
Partners: AZ Delta/RADar Learning & Innovation Center, UZ Leuven, Imeldaziekenhuis Bonheiden, AZ Oostende
Funding: Kom op Tegen Kanker, Roche
