Adaptive Multi-head Attention and Residual Gated Linear Units: An Ablation Analysis of Transformer Enhancements for Multimodal Clinical Tabular Data

Authors

  • Muhammed Kuliya
  • Zaharaddeen Salele Iro

DOI:

https://doi.org/10.64321/jcr.v3i4.03

Keywords:

Transformer ablation, multimodal clinical data, cardiovascular risk prediction, adaptive attention, gated linear units, tabular deep learning

Abstract

Background: Transformer architectures for tabular data have shown promise, but their optimal design for multimodal cardiovascular risk prediction remains unclear. This proof-of-concept study systematically ablates two architectural innovations using a synthetically aligned multimodal dataset derived from NHANES, Framingham, and Kaggle sources.

Objective: To systematically ablate two architectural innovations, Adaptive Multi-Head Attention (with learnable scaling parameters α, β) and Residual Gated Linear Units (ResidualGLU) within a TabTransformer framework for multitask prediction of coronary heart disease (CHD) and stroke.

Methods: Controlled ablation on fused multimodal dataset (N=10,410, 21 features from physiological, environmental, demographic, behavioral modalities). Four configurations: (1) Baseline TabTransformer, (2) +Adaptive Attention only, (3) +ResidualGLU only, (4) Full EnTabTransformer (both). Evaluation: 5-fold cross-validation with accuracy, F1-score, AUC-ROC. Statistical significance via paired t-tests with Bonferroni correction. Mechanistic analysis: gradient flow tracking (30 runs) and attention entropy.

Results: Adaptive Attention alone improved CHD AUC by +1.31pp (p<0.001) and stroke AUC by +1.75pp (p<0.001). ResidualGLU alone yielded +0.76pp (CHD, p=0.012) and +1.02pp (stroke, p=0.008). Full EnTabTransformer achieved super-additive gains: +2.71pp (CHD) and +3.77pp (stroke) exceeding sum of individual improvements by 31% (CHD) and 36% (stroke). Gradient analysis showed ResidualGLU reduced vanishing gradient probability by 34% (p=0.003). Adaptive Attention increased attention entropy by 0.42 bits (p<0.001), indicating broader feature coverage. Computational overhead: +18% parameters, +10% inference time. All results are based on a proof‑of‑concept dataset that includes synthetic ECG and simulated patient alignment; prospective validation on authentic multimodal clinical data is required before deployment.

Conclusion: Adaptive Attention and ResidualGLU address orthogonal aspects of feature representation, cross-feature interaction weighting vs. non-linear transformation capacity. Their combination yields synergetic gains. For resource-constrained deployment, Adaptive Attention alone is recommended; for maximum accuracy, full EnTabTransformer is justified.

Author Biographies

Muhammed Kuliya

Department of Information Technology Science, Federal University Dutse, Dutse, Jigawa State, Nigeria

Zaharaddeen Salele Iro

Department of Information Technology Science, Federal University Dutse, Dutse, Jigawa State, Nigeria

References

Alam, F., Ananbeh, O., Malik, K. M., Odayani, A. Al, Hussain, I. Bin, Kaabia, N., Aidaroos, A. Al, & Saudagar, A. K. J. (2023). Towards Predicting Length of Stay and Identification of Cohort Risk Factors Using Self-Attention-Based Transformers and Association Mining: COVID-19 as a Phenotype. Diagnostics, 13(10). https://doi.org/10.3390/diagnostics13101760

Amirahmadi, A., Etminani, F., & Ohlsson, M. (2025). Adaptive noise-augmented attention for enhancing Transformer fine-tuning on longitudinal medical data. Frontiers in Artificial Intelligence, 8. https://doi.org/10.3389/frai.2025.1663484

Chikumo, O. T., & Ndlovu, B. (2026). Transformer-based Models for Cardiovascular Disease Predictions from Electronic Health Records: A Systematic Review Article history. In Journal of Applied Informatics and Computing (JAIC) (Vol. 10, Number 1). http://jurnal.polibatam.ac.id/index.php/JAIC

Dauphin, Y. N., Fan, A., Auli, M., & Grangier, D. (2017). Language Modeling with Gated Convolutional Networks.

Dubey, P., Dubey, P., & Bokoro, P. N. (2025). Advancing CVD Risk Prediction with Transformer Architectures and Statistical Risk Factor Filtering. Technologies, 13(5). https://doi.org/10.3390/technologies13050201

Gorishniy, Y., Rubachev, I., Khrulkov, V., & Babenko, A. (2023). Revisiting Deep Learning Models for Tabular Data. https://doi.org/10.48550/arXiv.2106.11959

Gu, X., Tang, W., Han, J., Sangha, V., Liu, F., Gowda, S. N., Ribeiro, A. H., Schwab, P., Branson, K., Clifton, L., Ribeiro, A. L. P., Liu, Z., & Clifton, D. A. (2026). Cardiac health assessment across scenarios and devices using a multimodal foundation model pretrained on data from 1.7 million individuals. Nature Machine Intelligence, 8(2), 220–233. https://doi.org/10.1038/s42256-026-01180-5

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep Residual Learning for Image Recognition. http://image-net.org/challenges/LSVRC/2015/

Huang, X., Khetan, A., Cvitkovic, M., & Karnin, Z. (2020). TabTransformer: Tabular Data Modeling Using Contextual Embeddings. http://arxiv.org/abs/2012.06678

Kline, A., Wang, H., Li, Y., Dennis, S., Hutch, M., Xu, Z., Wang, F., Cheng, F., & Luo, Y. (2022). Multimodal machine learning in precision health: A scoping review. In npj Digital Medicine (Vol. 5, Number 1). Nature Research. https://doi.org/10.1038/s41746-022-00712-8

Koo, B., Sung, I., Lee, S., & Kim, S. (2025). Transcriptome Transformer: improving patient survival prediction via multitask learning of transcriptomic and clinical features. Briefings in Bioinformatics, 26(6). https://doi.org/10.1093/bib/bbaf628

Le, T.-D., Macabiau, C., Albert, K., Chatzinotas, S., Jouvet, P., & Noumeir, R. (2025). Transformer Meets Gated Residual Networks To Enhance Photoplethysmogram Artifact Detection Informed by Mutual Information Neural Estimation. http://arxiv.org/abs/2405.16177

Lim, B., Arık, S., Loeff, N., & Pfister, T. (2021). Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4), 1748–1764. https://doi.org/10.1016/j.ijforecast.2021.03.012

Madan, S., Lentzen, M., Brandt, J., Rueckert, D., Hofmann-Apitius, M., & Fröhlich, H. (2024). Transformer models in biomedicine. In BMC Medical Informatics and Decision Making (Vol. 24, Number 1). BioMed Central Ltd. https://doi.org/10.1186/s12911-024-02600-5

Moshawrab, M., Adda, M., Bouzouane, A., Ibrahim, H., & Raad, A. (2023). Smart Wearables for the Detection of Cardiovascular Diseases: A Systematic Literature Review. In Sensors (Vol. 23, Number 2). MDPI. https://doi.org/10.3390/s23020828

Noor, N., Bilal, M., Abbasi, S. F., Pournik, O., & Arvanitis, T. N. (2025). A novel transformer-based approach for cardiovascular disease detection. Frontiers in Digital Health, 7. https://doi.org/10.3389/fdgth.2025.1548448

Roy, S., Koehler, G., Baumgartner, M., Ulrich, C., Petersen, J., Isensee, F., & Maier-Hein, K. (2023). Transformer Utilization in Medical Image Segmentation Networks. http://arxiv.org/abs/2304.04225

Shaheenur Islam Sumon, Sakib Bin Islam, Sohanur Rahman, Sakib Abrar Hossain, Khandakar, A., Hasan, A., Murugappan, M., & H Chowdhury, M. E. (2025). CardioTabNet: A Novel Hybrid Transformer Model for Heart Disease Prediction using Tabular Medical Data. https://doi.org/10.1007/s13755-025-00361-7

Shazeer, N. (2020). GLU Variants Improve Transformer. http://arxiv.org/abs/2002.05202

Somepalli, G., Goldblum, M., Schwarzschild, A., Bruss, C. B., & Goldstein, T. (2021). SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training. http://arxiv.org/abs/2106.01342

Talaat, F. M., & Aly, W. F. (2025). Toward precision cardiology: a transformer-based system for adaptive prediction of heart disease. Neural Computing and Applications, 37(19), 13547–13571. https://doi.org/10.1007/s00521-025-11172-y

Teoh, J. R., Dong, J., Zuo, X., Lai, K. W., Hasikin, K., & Wu, X. (2024). Advancing healthcare through multimodal data fusion: a comprehensive review of techniques and applications. In PeerJ Computer Science (Vol. 10). PeerJ Inc. https://doi.org/10.7717/PEERJ-CS.2298

Vaswani, A., Brain, G., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention Is All You Need.

Ye, J., Zhang, W., Li, Z., Li, J., & Tsung, F. (2026). MedSpaformer: a Transferable Transformer with Multi-granularity Token Sparsification for Medical Time Series Classification. www.aaai.org

Zhang, J., Zang, X., Chen, H., Yan, X., & Tang, B. (2025). Highlights DispFormer: A Dual Attention Transformer with Denoising for Irregular Clinical Time Series Classification DispFormer: A Dual Attention Transformer with Denoising for Irregular Clinical Time Series Classification. https://github.com/junjzhang7/DispFormer.

Zhang, Y., & Li, S. (2025). ChronoFormer: Time-Aware Transformer Architectures for Structured Clinical Event Modeling. http://arxiv.org/abs/2504.07373

Zhao, F., Zhang, C., & Geng, B. (2024). Deep Multimodal Data Fusion. ACM Computing Surveys, 56(9). https://doi.org/10.1145/3649447

Zheng, Y., Ma, Y., & Tian, C. (2022). TMRN-GLU: A Transformer-Based Automatic Classification Recognition Network Improved by Gate Linear Unit. Electronics (Switzerland), 11(10). https://doi.org/10.3390/electronics11101554

Zheng, Z., Zhang, F., Liu, S., Xia, T., Liu, X., Hu, D., & Zhou, H. (2026). Multi-Gate Residuals. http://arxiv.org/abs/2605.23259

Downloads

Published

2026-07-10

How to Cite

Muhammed Kuliya, & Zaharaddeen Salele Iro. (2026). Adaptive Multi-head Attention and Residual Gated Linear Units: An Ablation Analysis of Transformer Enhancements for Multimodal Clinical Tabular Data. Journal of Current Research and Studies, 3(4), 25–40. https://doi.org/10.64321/jcr.v3i4.03