A Dual-Stream Multi-Axis Vision Transformer and Feature-Attention Fusion Framework for Multimodal Heart Disease Classification from ECG Images and Clinical Data
Main Article Content
Abstract
Cardiovascular disease remains the leading cause of mortality worldwide, and its reliable diagnosis depends on the joint interpretation of electrocardiogram (ECG) waveforms and routine clinical measurements. Most existing computer-aided diagnosis systems, however, exploit only a single modality and therefore discard complementary diagnostic evidence. This paper proposes a dual-stream multimodal deep-learning framework that couples a Multi-Axis Vision Transformer (MaxViT) branch for ECG image understanding with a feature-attention multilayer perceptron (MLP) branch for tabular clinical data, and combines their learned embeddings through a dedicated late-fusion head. Both unimodal encoders are pre-trained independently and then frozen, so that the 832-dimensional fusion head learns purely cross-modal interactions and cannot collapse onto a single dominant modality. The framework classifies patients into four clinically meaningful categories—Normal, Abnormal Heartbeat, Myocardial Infarction and Post-MI History—on a label-aligned corpus of 928 ECG images and a SMOTE-balanced clinical cohort. Experimental results show that the clinical-only branch attains 68.75% accuracy, the ECG-only branch attains 96.76%, and the proposed fusion model reaches 98.75% accuracy with a macro F1-score of 0.987, outperforming both unimodal baselines and a set of representative single-modality and early-fusion methods reported in the literature. The results confirm that principled multimodal fusion of ECG images and clinical variables yields a more accurate and clinically trustworthy cardiac classifier than any individual modality alone.


