Adaptive Multi-head Attention and Residual Gated Linear Units: An Ablation Analysis of Transformer Enhancements for Multimodal Clinical Tabular Data
DOI:
https://doi.org/10.64321/jcr.v3i4.03Keywords:
Transformer ablation, multimodal clinical data, cardiovascular risk prediction, adaptive attention, gated linear units, tabular deep learningAbstract
Background: Transformer architectures for tabular data have shown promise, but their optimal design for multimodal cardiovascular risk prediction remains unclear. This proof-of-concept study systematically ablates two architectural innovations using a synthetically aligned multimodal dataset derived from NHANES, Framingham, and Kaggle sources.
Objective: To systematically ablate two architectural innovations, Adaptive Multi-Head Attention (with learnable scaling parameters α, β) and Residual Gated Linear Units (ResidualGLU) within a TabTransformer framework for multitask prediction of coronary heart disease (CHD) and stroke.
Methods: Controlled ablation on fused multimodal dataset (N=10,410, 21 features from physiological, environmental, demographic, behavioral modalities). Four configurations: (1) Baseline TabTransformer, (2) +Adaptive Attention only, (3) +ResidualGLU only, (4) Full EnTabTransformer (both). Evaluation: 5-fold cross-validation with accuracy, F1-score, AUC-ROC. Statistical significance via paired t-tests with Bonferroni correction. Mechanistic analysis: gradient flow tracking (30 runs) and attention entropy.
Results: Adaptive Attention alone improved CHD AUC by +1.31pp (p<0.001) and stroke AUC by +1.75pp (p<0.001). ResidualGLU alone yielded +0.76pp (CHD, p=0.012) and +1.02pp (stroke, p=0.008). Full EnTabTransformer achieved super-additive gains: +2.71pp (CHD) and +3.77pp (stroke) exceeding sum of individual improvements by 31% (CHD) and 36% (stroke). Gradient analysis showed ResidualGLU reduced vanishing gradient probability by 34% (p=0.003). Adaptive Attention increased attention entropy by 0.42 bits (p<0.001), indicating broader feature coverage. Computational overhead: +18% parameters, +10% inference time. All results are based on a proof‑of‑concept dataset that includes synthetic ECG and simulated patient alignment; prospective validation on authentic multimodal clinical data is required before deployment.
Conclusion: Adaptive Attention and ResidualGLU address orthogonal aspects of feature representation, cross-feature interaction weighting vs. non-linear transformation capacity. Their combination yields synergetic gains. For resource-constrained deployment, Adaptive Attention alone is recommended; for maximum accuracy, full EnTabTransformer is justified.
References
Alam, F., Ananbeh, O., Malik, K. M., Odayani, A. Al, Hussain, I. Bin, Kaabia, N., Aidaroos, A. Al, & Saudagar, A. K. J. (2023). Towards Predicting Length of Stay and Identification of Cohort Risk Factors Using Self-Attention-Based Transformers and Association Mining: COVID-19 as a Phenotype. Diagnostics, 13(10). https://doi.org/10.3390/diagnostics13101760
Amirahmadi, A., Etminani, F., & Ohlsson, M. (2025). Adaptive noise-augmented attention for enhancing Transformer fine-tuning on longitudinal medical data. Frontiers in Artificial Intelligence, 8. https://doi.org/10.3389/frai.2025.1663484
Chikumo, O. T., & Ndlovu, B. (2026). Transformer-based Models for Cardiovascular Disease Predictions from Electronic Health Records: A Systematic Review Article history. In Journal of Applied Informatics and Computing (JAIC) (Vol. 10, Number 1). http://jurnal.polibatam.ac.id/index.php/JAIC
Dauphin, Y. N., Fan, A., Auli, M., & Grangier, D. (2017). Language Modeling with Gated Convolutional Networks.
Dubey, P., Dubey, P., & Bokoro, P. N. (2025). Advancing CVD Risk Prediction with Transformer Architectures and Statistical Risk Factor Filtering. Technologies, 13(5). https://doi.org/10.3390/technologies13050201
Gorishniy, Y., Rubachev, I., Khrulkov, V., & Babenko, A. (2023). Revisiting Deep Learning Models for Tabular Data. https://doi.org/10.48550/arXiv.2106.11959
Gu, X., Tang, W., Han, J., Sangha, V., Liu, F., Gowda, S. N., Ribeiro, A. H., Schwab, P., Branson, K., Clifton, L., Ribeiro, A. L. P., Liu, Z., & Clifton, D. A. (2026). Cardiac health assessment across scenarios and devices using a multimodal foundation model pretrained on data from 1.7 million individuals. Nature Machine Intelligence, 8(2), 220–233. https://doi.org/10.1038/s42256-026-01180-5
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep Residual Learning for Image Recognition. http://image-net.org/challenges/LSVRC/2015/
Huang, X., Khetan, A., Cvitkovic, M., & Karnin, Z. (2020). TabTransformer: Tabular Data Modeling Using Contextual Embeddings. http://arxiv.org/abs/2012.06678
Kline, A., Wang, H., Li, Y., Dennis, S., Hutch, M., Xu, Z., Wang, F., Cheng, F., & Luo, Y. (2022). Multimodal machine learning in precision health: A scoping review. In npj Digital Medicine (Vol. 5, Number 1). Nature Research. https://doi.org/10.1038/s41746-022-00712-8
Koo, B., Sung, I., Lee, S., & Kim, S. (2025). Transcriptome Transformer: improving patient survival prediction via multitask learning of transcriptomic and clinical features. Briefings in Bioinformatics, 26(6). https://doi.org/10.1093/bib/bbaf628
Le, T.-D., Macabiau, C., Albert, K., Chatzinotas, S., Jouvet, P., & Noumeir, R. (2025). Transformer Meets Gated Residual Networks To Enhance Photoplethysmogram Artifact Detection Informed by Mutual Information Neural Estimation. http://arxiv.org/abs/2405.16177
Lim, B., Arık, S., Loeff, N., & Pfister, T. (2021). Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4), 1748–1764. https://doi.org/10.1016/j.ijforecast.2021.03.012
Madan, S., Lentzen, M., Brandt, J., Rueckert, D., Hofmann-Apitius, M., & Fröhlich, H. (2024). Transformer models in biomedicine. In BMC Medical Informatics and Decision Making (Vol. 24, Number 1). BioMed Central Ltd. https://doi.org/10.1186/s12911-024-02600-5
Moshawrab, M., Adda, M., Bouzouane, A., Ibrahim, H., & Raad, A. (2023). Smart Wearables for the Detection of Cardiovascular Diseases: A Systematic Literature Review. In Sensors (Vol. 23, Number 2). MDPI. https://doi.org/10.3390/s23020828
Noor, N., Bilal, M., Abbasi, S. F., Pournik, O., & Arvanitis, T. N. (2025). A novel transformer-based approach for cardiovascular disease detection. Frontiers in Digital Health, 7. https://doi.org/10.3389/fdgth.2025.1548448
Roy, S., Koehler, G., Baumgartner, M., Ulrich, C., Petersen, J., Isensee, F., & Maier-Hein, K. (2023). Transformer Utilization in Medical Image Segmentation Networks. http://arxiv.org/abs/2304.04225
Shaheenur Islam Sumon, Sakib Bin Islam, Sohanur Rahman, Sakib Abrar Hossain, Khandakar, A., Hasan, A., Murugappan, M., & H Chowdhury, M. E. (2025). CardioTabNet: A Novel Hybrid Transformer Model for Heart Disease Prediction using Tabular Medical Data. https://doi.org/10.1007/s13755-025-00361-7
Shazeer, N. (2020). GLU Variants Improve Transformer. http://arxiv.org/abs/2002.05202
Somepalli, G., Goldblum, M., Schwarzschild, A., Bruss, C. B., & Goldstein, T. (2021). SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training. http://arxiv.org/abs/2106.01342
Talaat, F. M., & Aly, W. F. (2025). Toward precision cardiology: a transformer-based system for adaptive prediction of heart disease. Neural Computing and Applications, 37(19), 13547–13571. https://doi.org/10.1007/s00521-025-11172-y
Teoh, J. R., Dong, J., Zuo, X., Lai, K. W., Hasikin, K., & Wu, X. (2024). Advancing healthcare through multimodal data fusion: a comprehensive review of techniques and applications. In PeerJ Computer Science (Vol. 10). PeerJ Inc. https://doi.org/10.7717/PEERJ-CS.2298
Vaswani, A., Brain, G., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention Is All You Need.
Ye, J., Zhang, W., Li, Z., Li, J., & Tsung, F. (2026). MedSpaformer: a Transferable Transformer with Multi-granularity Token Sparsification for Medical Time Series Classification. www.aaai.org
Zhang, J., Zang, X., Chen, H., Yan, X., & Tang, B. (2025). Highlights DispFormer: A Dual Attention Transformer with Denoising for Irregular Clinical Time Series Classification DispFormer: A Dual Attention Transformer with Denoising for Irregular Clinical Time Series Classification. https://github.com/junjzhang7/DispFormer.
Zhang, Y., & Li, S. (2025). ChronoFormer: Time-Aware Transformer Architectures for Structured Clinical Event Modeling. http://arxiv.org/abs/2504.07373
Zhao, F., Zhang, C., & Geng, B. (2024). Deep Multimodal Data Fusion. ACM Computing Surveys, 56(9). https://doi.org/10.1145/3649447
Zheng, Y., Ma, Y., & Tian, C. (2022). TMRN-GLU: A Transformer-Based Automatic Classification Recognition Network Improved by Gate Linear Unit. Electronics (Switzerland), 11(10). https://doi.org/10.3390/electronics11101554
Zheng, Z., Zhang, F., Liu, S., Xia, T., Liu, X., Hu, D., & Zhou, H. (2026). Multi-Gate Residuals. http://arxiv.org/abs/2605.23259
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Muhammed Kuliya

This work is licensed under a Creative Commons Attribution 4.0 International License.