Project

Medical AI Model Development

Modern medicine generates large amounts of imaging and complex multimodal data. AI has high potential to support clinicians by reducing diagnostic errors and streamline routine hospital workflows. However, several barriers limit its safe adoption in clinical practice. Many deep-learning systems lack interpretability, cannot communicate diagnostic uncertainty, produce hallucinations, or exploit unintended shortcuts in training datasets. Scaling these systems also requires large, annotated datasets, whose creation is constrained by the time and cost of expert review. In addition, clinical deployment requires automated quality-assurance mechanisms to identify inconsistent scans and protect patients from unnecessary radiation exposure.

The Medical AI Model Development project addressed these challenges across several medical imaging applications, including chest and musculoskeletal radiography, computed tomography (CT), magnetic resonance imaging (MRI), and radiogenomic analysis. High-performance computing infrastructure was essential throughout the project. The research involved hundreds of thousands of high-resolution medical images, whole-exome sequencing data from cancer cohorts, large-scale model training, and repeated inference and evaluation cycles. Without distributed GPU clusters, these computationally intensive workflows would not have been feasible.

Project Details

Project term

July 1, 2024–August 1, 2026

Affiliations

Uniklinik RWTH Aachen

Institute

Lab for Artificial Intelligence in Medicine

Principal Investigator

Dr. Daniel Truhn

Methods

The research focused on four areas:

Self-explainable architectures and shortcut prevention: We developed an inherently interpretable architecture for chest radiograph classification as an alternative to post hoc visualization methods such as heatmaps. The architecture divides each image into non-overlapping spatial patches, processes the patches independently using a neural network, and combines their logits using an unweighted mean. This design makes it possible to visualize explicitly how individual image regions contribute evidence for or against specific conditions.

Uncertainty quantification and hallucination filtering: We used stochastic sampling to measure model disagreement and reduce diagnostic overconfidence. In visual question answering, the system generated multiple answers for each scan, grouped them according to their core clinical meaning, and used the resulting level of disagreement to identify and reject questions associated with unreliable responses or false findings. In multitask chest radiograph classification, repeated stochastic inference with active dropout quantified predictive uncertainty across tens of thousands of images, allowing uncertain predictions to be routed to downstream decision agents.

Automated label extraction and multi-agent reasoning: To reduce reliance on manual annotation, large language models extracted structured labels from routine clinical reports. These labels were then used to train multilabel classification networks for three anatomical joint regions. For visually ambiguous diagnostic tasks, like distinguishing melanoma from benign nevi or pulmonary edema from pneumonia, we developed a multi-agent system in which two opposing agents presented competing interpretations and a judge agent evaluated their arguments directly against the image.

Large-scale anatomical morphometry and quality assurance: We developed deep-learning pipelines, to quantify anatomical structures automatically. Applications included measuring knee indices on lateral radiographs, assessing lower-limb torsion on MRI, and screening low-dose CT examinations for scan-boundary and image-noise irregularities through automated segmentation of lung and aortic volumes. We also combined foundation-model-based genomic sequence analysis with automated three-dimensional tumor segmentation to investigate associations between somatic gene disruption and radiomic tumor phenotypes.

Results

Explainability without loss of performance: The patch-based network achieved classification performance comparable to that of conventional deep-learning models across 14 chest conditions in large benchmark cohorts. It also localized thoracic abnormalities more accurately than traditional post hoc visualization methods.

Reduction in hallucinations and diagnostic errors: Semantic uncertainty filtering substantially improved the accuracy of vision-language models. On the combined multimodal datasets, selectively rejecting high-uncertainty questions increased GPT-4o accuracy from 51.7% to 76.3%. Similarly, providing downstream clinical decision agents with explicit binary error-risk indicators reduced confident diagnostic errors in unreliable cases from 8.5% to 2.7%.

Automated labeling and contrastive adjudication: The automated report-extraction pipeline achieved more than 98% accuracy when parsing complex clinical text. The resulting classifiers were robust to minor linguistic uncertainty in the extracted labels. For visually ambiguous conditions, structured multi-agent adjudication improved diagnostic accuracy in skin-lesion differentiation by 11 percentage points compared with standard baselines.

Population-scale morphometry and quality auditing: The automated CT quality-assurance tool processed 38,834 full-volume screening examinations. It characterized deviations in scan coverage, identifying mean excess coverage of 31.2 mm below the lower lung boundary and 14.5 mm above the upper boundary. In the radiogenomic analysis, combining genomic foundation models with CT imaging recovered established renal cancer drivers and identified statistically significant associations between tumor volume and 46 candidate genes not included in existing reference lists.

Discussion

Taken together, the project’s findings demonstrate how high-performance computing can support the development of medical AI systems that are more transparent, uncertainty-aware, and capable of automated quality control. These advances represent important steps toward moving medical AI beyond opaque and overconfident prototypes and toward robust clinical decision-support systems.

Several limitations remain. First, patch-based explainability may not adequately represent diseases whose diagnosis depends on integrating information across distant anatomical regions. Second, the uncertainty methods primarily measure response consistency rather than factual correctness. A model may therefore produce a confident and consistently repeated hallucination that is not detected by consistency-based filtering. Third, most evaluations were conducted using retrospective cohorts, simplified two-dimensional representations of volumetric anatomy, or isolated diagnostic tasks without access to complete clinical histories and real-world electronic health-record data.

Future work will focus on integrating additional multimodal patient data, improving computational efficiency for hospital deployment, and conducting prospective clinical studies to evaluate the safety and effectiveness of AI-based decision support in routine clinical practice.

Additional Project Information

DFG classification: 205-30 Radiology and Nuclear Medicine
Software: PyTorch
Cluster: CLAIX

Publications

Large language model-based uncertainty-adjusted label extraction for artificial intelligence model development in upper extremity radiograph,
Hanna Kreutzer, Anne-Sophie Caselitz, Thomas Dratsch, Daniel Pinto dos Santos, Christiane Kuhl, Daniel Truhn, Sven Nebelung,
https://dx.doi.org/10.1007/s00330-025-12102-1, November 2025

Low-Dose CT Quality Assurance at Scale: Automated Detection of Overscanning, Underscanning, and Image Noise,
Patrick Wienholt, Alexander Hermans, Robert Siepmann, Christiane Kuhl, Daniel Pinto dos Santos, Sven Nebelung, Daniel Truhn,
https://dx.doi.org/10.3390/life16010152, January 2026

MedicalPatchNet: a patch-based self-explainable AI architecture for chest X-ray classification,
Patrick Wienholt, Christiane Kuhl, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn,
https://dx.doi.org/10.1038/s41598-026-40358-0, February 2026

Hallucination filtering in radiology vision-language models using discrete semantic entropy,
Patrick Wienholt, Sophie Caselitz, Robert Siepmann, Philipp Bruners, Keno Bressem, Christiane Kuhl, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn,
https://dx.doi.org/10.1007/s00330-026-12384-z, February 2026

Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study,
Zihao Zhao, Frederik Hauke, Juliana De Castilhos, Sven Nebelung, Daniel Truhn,
https://dx.doi.org/10.48550/arXiv.2602.22959, February 2026

Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans,
Frederik Hauke, Jeremias Krause, Patrick Wienholt, Christiane Kuhl, Ingo Kurth, Sikander Hayat, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn,
https://dx.doi.org/10.48550/arXiv.2607.20583, July 2026

Bayesian uncertainty estimation improves clinical decision making in medical AI agents,
Frederik Hauke, Patrick Wienholt, Christiane Kuhl, Dyke Ferber, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn,
https://dx.doi.org/10.48550/arXiv.2607.20582, July 2026

Large-Scale AI-based Morphometry Challenges Traditional Reference Values for Patellofemoral Indices on Lateral Knee Radiographs,
Dennis Eschweiler, Eneko Cornejo Merodio, Felix Barajas Ordonez, Hanna Kreutzer, Aleksandar Lichev, Daniel Yeong-Jun Geiger, Christian David Weber, Christiane Katharina Kuhl, Daniel Truhn, Sven Nebelung, 2026

External Validation and Large-Scale Application of AI-Based Multi-Parametric Lower-Limb Morphometry from Torsional MRI,
Simon D. Westfechtel, Marc S. von der Stück, Felix M. Barajas Ordonez, Jonathan Lemessa, Dennis Eschweiler, Wolfgang Fischer, Juliana de Castilhos, Christiane K. Kuhl, Daniel Truhn, Sven Nebelung, 2026