Masked Autoencoders based Multi-Modal Occlusion-Robust Attention-Guided MRI Fusion for Brain Tumor Detection

Main Article Content

Kaluva Jaya Deepthi, M Sirish Kumar, B. Narendra Kumar Rao

Abstract

Early and an accurate detection of the brain tumors from a multi-parametric magnetic resonance imaging (MRI) which includes T1, contrast-enhanced T1 (T1ce), T2, and FLAIR is considered critical for selecting timely treatment and which has improved patient outcomes. Many of the high-performing deep models are trained on single modalities or curated datasets and it therefore gets underperformed on a heterogeneous, occluded, or noisy clinical scans. Existing fusion schemes often use fixed or naive aggregation, and it fails to learn to compensate for missing or corrupted information. This leads to brittle behavior when faced with the motion artifacts, inter-center acquisition variability, or missing slices, which limits the clinical adoption. We propose an occlusion-invariant, attention-guided multi-modal fusion framework. Each MRI modality is standardized, patch-embedded, and encoded by a modality-specific backbone. An attention fusion module computes adaptive modality weights to emphasize the complementary lesion cues. A Masked Autoencoder (MAE) is trained in a self-supervised way by masking 10–40% of patches across the modalities and reconstructing them, enabling the encoder to predict the missing content and learn occlusion invariance. The fused representation is fine-tuned with the lightweight dual-head network that jointly has performed early-stage tumor detection/classification and coarse segmentation using a has combined cross-entropy + Dice loss. Training uses mixed supervised and self-supervised objectives and curriculum masking to progressively increase the mask difficulty. Evaluated on a public multi-modal benchmark, the method attains classification accuracy 96.4%, AUC 0.987, sensitivity 95.1%, specificity 96.9%, and mean Dice 0.823. Under 30% patch occlusion, accuracy falls only 2.1% (to 94.3%) versus a 9.8% drop for a conventional fusion baseline. Relative to single-modality models, we observe +4.8% accuracy and +6.7% Dice improvements, while reducing inference FLOPs by ≈28% and achieving inference latency ≈18 ms per slice on a single GPU, which supports the near-real-time clinical use.

Article Details

How to Cite
Kaluva Jaya Deepthi, M Sirish Kumar, B. Narendra Kumar Rao. (2026). Masked Autoencoders based Multi-Modal Occlusion-Robust Attention-Guided MRI Fusion for Brain Tumor Detection. International Journal of Special Education, 41(10s), 396–421. Retrieved from https://internationalsped.com/index.php/ijse/article/view/3670
Section
General