SYNTHSEG-DIFF: Conditional Denoising Diffusion Probabilistic Model For Cross-Modal Medical Image Synthesis and Multi-Class Brain Tumour Segmentation
DOI:
https://doi.org/10.51483/IJAIML.6.3.2026.84-99Keywords:
denoising diffusion probabilistic model, brain tumour segmentation, cross-modal image synthesis, multi-modal MRI, spectral attention, uncertainty estimation, external validation.Abstract
Brain tumour segmentation and cross-modal MRI synthesis remain challenging because of limited annotated multi-modal data, tumor heterogeneity, class imbalance, and variability in acquisition protocols. We present SynthSeg-Diff, a conditional denoising diffusion probabilistic model (C-DDPM) that combines learnable multi-modal cross-attention, a Multi-scale Adaptive Stochastic Guidance (MASG) noise schedule, and a Feature-Space Spectral Attention (FSSA) module for brain MRI synthesis and multi-class tumor segmentation. For T1ce synthesis, T1, T2, and FLAIR are used as conditioning inputs and T1ce is treated only as the synthesis target; the segmentation pathway uses the available multi-modal MRI inputs. A tri-planar 2.5D perceptual consistency loss and mixed-precision FP16 training are used to improve structural consistency and computational efficiency. The method is evaluated using stratified 5-fold cross-validation on BraTS 2021 (n = 1251) and zero-shot testing on BraTS 2023 Adult Glioma (n = 1251) and UPenn-GBM (n = 611) to assess generalization under domain shift. On BraTS 2021, SynthSeg-Diff reports a case-averaged DSC of 0.918 ± 0.011, IoU of 0.861 ± 0.013, HD95 of 4.1 ± 0.4 mm, sensitivity of 92.7 ± 0.6%, and T1ce synthesis PSNR of 35.6 ± 0.7 dB. External testing yields DSC values of 0.886 ± 0.014 on BraTS 2023 and 0.851 ± 0.019 on curated UPenn-GBM data, with a larger decline on raw clinical UPenn-GBM scans. Reported improvements over re-trained baselines are statistically significant using paired Wilcoxon signed-rank tests with Bonferroni correction. Stochastic inference produces voxel-wise predictive-uncertainty maps that are reported to correlate with annotation-disagreement regions. Single-pass deterministic inference is reported as 21.4 ms per volume, while the 10-pass uncertainty pipeline requires 214 ms. These findings support further investigation of frequency-aware conditional diffusion for multimodal brain MRI analysis while highlighting the need for prospective validation, missing-modality robustness, and rigorous verification of uncertainty and latency measurements.





