Project

Extract, Design and study digital histological biomarkers to investigate the underlying effects and mechanisms of kidney diseases

Kidney diseases remain difficult to diagnose and stratify because conventional histopathology relies on qualitative, visually assessed criteria that vary considerably between observers. Whole-slide imaging now makes it possible to capture tissue morphology at a resolution and scale that human assessment cannot match, but turning gigapixel images into quantitative, reproducible biomarkers requires substantial computational infrastructure. Our project addresses this gap by building a deep-learning-based computational pathology pipeline that automatically segments kidney tissue into its relevant sub-compartments and extracts high-dimensional Pathomics features from thousands of digitized slides. The long-term aim is to move beyond descriptive pathology toward objective, data-driven diagnostics by identifying novel visual biomarkers, refining patient risk stratification, and extending established diagnostic frameworks such as the Banff and MEST-C classification systems with quantitative digital counterparts. Processing cohorts of this size and expanding to include Renal Cell Carcinoma (RCC) cases with their high nuclear density is only feasible with access to high-performance computing. Segmentation, feature extraction, and the downstream statistical modeling of thousands of high-resolution images each demand parallelized GPU and CPU resources far beyond what a local workstation could provide within a realistic timeframe.

Project Details

Project term

March 18, 2025–June 17, 2026

Affiliations

Uniklinik RWTH Aachen

Institute

Institut für Pathologie

Principal Investigator

Prof. Dr. Dr. Peter Boor

Methods

The workflow is organized as a multi-stage pipeline covering data curation, tissue segmentation, feature extraction, and statistical analysis. Whole-slide images pass through a preprocessing step to filter out low-quality slides and artifacts, followed by foreground–background segmentation to isolate the tissue regions. Nuclei segmentation is performed using CellViT, a vision-transformer architecture that we fine-tuned on PAS-stained kidney tissue. For segmentation of kidney compartments, we use Segmenter, which our benchmark study showed had the highest Dice score. Each segmented compartment is then passed through a feature extraction pipeline consisting of PyRadiomics, HistomicsTK, NetworkX-based graph features, and pathologically designed features. All resulting descriptors are compiled into dense feature matrices for downstream statistical analysis; we trained 2 classifiers to detect inflammatory nuclei and intra-glomerular cell subtypes from nuclei features, and 2 more classifiers for glomerular disease patterns and tubular atrophy classification. For the RCC extension, the same CellViT-based pipeline will be applied to tumor nuclei, feeding into regularized survival models. Our cohort comprises roughly 1,000 patients, each contributing multiple slides and tissue sections, with individual tissues generating on the order of 10E5-10E6 nuclei. Given this scale, parallelization is necessary; we run the segmentations on a single GPU per tissue, and feature extraction is distributed across 64-core CPU jobs to keep pace with the data volume.

Results

In the first phase, we completed development of the core Pathomics pipeline, MEGATRON, and showed that pathomics features can classify diagnostic lesions such as inflammatory and intra-glomerular cell subtypes. We also used tubular and interstitial features to separate atrophic from non-atrophic tubules, define scarred interstitium, and characterize inflamed and scarred-inflamed interstitial foci. Applying MEGATRON to IgA nephropathy cohorts, we found that these features were independent predictors of long-term outcomes, outperforming the established MEST-C score as well as our earlier FLASH framework. The IgAN cohort also let us explore additional clinical questions beyond the original scope, linking tubular features to eGFR decline slope in an external cohort.
In the second phase, we are extending MEGATRON to additional international cohorts, with a focus on kidney transplantation. Incorporating data from new centers has introduced staining and scanner variability beyond what we saw internally, which we addressed by extending our normalization steps to keep features comparable across sites. Once harmonized, we aggregate compartment-level features to the tissue and patient level and apply MEGATRON across internal and international multi-center transplant cohorts. In this setting, pathomics features have captured a smooth morphological transition from healthy tissue through borderline changes to acute T-cell-mediated rejection, and we are evaluating their performance as independent predictors of long-term outcomes relative to the established Banff score. To date, roughly 25 percent of the total planned dataset has been processed through the full pipeline, consuming approximately 2.49 of the 12 million allocated core-hours.

Discussion

These results show that MEGATRON’s features carry independent, clinically meaningful prognostic information beyond existing scoring systems. In IgA nephropathy, outperforming MEST-C and our earlier FLASH framework suggests that lesion-targeted features capture disease-relevant heterogeneity that other methods miss, and the link to eGFR decline slope in an external cohort supports that this signal generalizes beyond the development cohort. We are now testing whether this advantage extends to kidney transplantation, where preliminary application of MEGATRON has captured a smooth transition from healthy through borderline to acute T-cell-mediated rejection; evaluating these features against the Banff score across our international multi-center transplant cohorts is a central focus of the current phase.
The main bottlenecks encountered were logistical rather than scientific: curating expert ground truth annotations across cohorts and harmonizing heterogeneous multi-center staining and scanner variability required more time than planned, leaving a substantial share of allocated core-hours unused. Our priority for the remaining period is to complete the transplant cohort analysis now underway, apply the validated pipeline to the newly finalized multi-center cohorts to confirm that features generalize beyond the datasets used for development, and extend it to the planned Renal Cell Carcinoma application, where nuclear grading is highly prognostic but suffers from marked inter-observer variability and an automated alternative could improve reproducibility. This continued expansion, together with the scale of nuclear-level feature extraction across millions of cells, makes sustained HPC access essential to completing the project’s final validation stage.

Additional Project Information

DFG classification: 205-06 Pathology
Software: PyTorch, PyRadiomics, HistomicsTK, NetworkX, CellViT, Segmenter
Cluster: CLAIX