Input-Dependent Feature Fusion in CNN–Transformer Networks for Image Classification

Authors

  • Mrs Komal Sharma
  • Dr Monika Sainger

Keywords:

convolutional neural networks, vision Transformers, adaptive fusion, feature weighting, explainable AI, image classification, medical imaging.

Abstract

Modern image classifiers often face a trade-off between detailed spatial representation and broad contextual reasoning. Convolutional networks are particularly effective at preserving local visual structure, while Transformer encoders can relate information across distant image regions. This paper develops a research framework in which these two representations are not merged with a fixed rule. Instead, a learnable fusion gate assigns input-dependent importance to the convolutional and attention branches. The framework also records branch weights and attention information so that the source of a prediction can be inspected. The experimental plan covers CIFAR-10, ImageNet and a proposed MRI/CT medical-imaging setting for lung-cancer classification. Performance is to be assessed through accuracy, precision, recall and F1-score, together with computational measures such as floating-point operations and inference time. Because the supplied study material does not contain completed experimental measurements, this revised paper deliberately avoids claiming numerical improvements that have not been observed. The contribution is therefore presented as a reproducible experimental design rather than as a fabricated performance claim.

Downloads

Published

2026-07-19

How to Cite

Sharma, M. K., & Sainger, D. M. (2026). Input-Dependent Feature Fusion in CNN–Transformer Networks for Image Classification . International Journal of Artificial Intelligence and Machine Learning, 6(7s), 1029–1036. Retrieved from https://svedbergopen.com/index.php/ijaiml/article/view/1146