Web Element Analysis: A Multi-Modal AI Framework Utilizing Structural Visual and Interaction Embeddings

Authors

  • Ibrar Ahmed
  • Richa Gupta
  • Ihtiram Raza Khan
  • Kanika Singhal
  • Deepak Chandra Uprety
  • Vrinda Sachdeva

DOI:

https://doi.org/10.51483/IJAIML.6.8s.2026.437-448

Keywords:

multi-modal embeddings, web element representation, DOM structure encoding, visual feature extraction, interaction modeling, graph neural networks, transformer architectures, automated web testing.

Abstract

Understanding of Web elements is essential for many applications such as automated User Interface (UI) testing, Web accessibility checking, semantic classification of Web elements and cross-browser anomaly detection. The current methods mostly focus on a single source of information, either using Document Object Model (DOM) structures, appearance or interaction logs, and thus provide only partial contextual awareness and lack generalizability in heterogeneous web contexts. The fragmented treatment of web elements is a big research gap since it does not consider the complementary relationship among structural, visual and behavioral elements of web interfaces as a whole, which, together, define the user interaction with modern web interfaces. To overcome this constraint, this study introduces a novel multi-modal feature extraction framework called MMFeX-Web that extracts complete and coherent features about web elements. The key goal is to improve the understanding of web elements through collaborative fusion of three orthogonal modalities: (i) structural embeddings from DOM tree structures, (ii) visual embeddings extracted from rendered pixel regions using a convolutional vision encoder, and (iii) interaction-level embeddings representing sequences over time and user interactions. A learned gated fusion operator (Ψ) learns to fuse these modalities adaptively and build a composite embedding vector (φ(e) ∈ ℝᵈ) given the type of the element and the page context. From a methodological perspective, the proposed framework is based on a multi-modal learning architecture and adaptive fusion mechanisms to retain the modality-specific information while learning the dependencies between modalities. Theoretical results set bounds on the amount of mutual information preserved by the fusion process and show that the fused representation subsumes single modality approaches in terms of Rademacher complexity. The effectiveness of MMFeX-Web is empirically tested using five benchmark datasets for web accessibility testing, automated UI testing, classification of semantic elements, and cross browser anomaly detection. Experimental results show that MMFeX-Web outperforms the best current baselines by 8.3% to up to 14.7% on the F1 score. The results validate that the fusion of structural, visual and interaction-level information greatly improves the representation learning of web elements. This research opens the door to more intelligent, reliable, and context-aware Web automation systems, and has practical applications in software testing, accessibility testing, browser compatibility testing, and next generation human computer interaction technologies.

Downloads

Published

2026-08-01

How to Cite

Ahmed, I., Gupta, R., Khan, I. R., Singhal, K., Uprety, D. C., & Sachdeva, V. (2026). Web Element Analysis: A Multi-Modal AI Framework Utilizing Structural Visual and Interaction Embeddings. International Journal of Artificial Intelligence and Machine Learning, 6(8s), 437–448. https://doi.org/10.51483/IJAIML.6.8s.2026.437-448