Deep Learning-Based Feature Learning for Drug Discovery and Biomedical Knowledge Extraction

Authors

  • Mohammad Shahid
  • Dr. Vimal Bibhu
  • Dr. Quaisar Alam

DOI:

https://doi.org/10.51483/IJAIML.6.11s.2026.1556-1563

Keywords:

Predicting drug-target interactions, drug chemical, convolutional neural network, Deep learning architectures.

Abstract

Deep Learning (DL) models rely on feature learning rather than manual feature extraction. This is a promising approach for a variety of applications including drug discovery and biomedical knowledge extraction. Although DL is similar to Machine Learning, they differ in the complexity of the learned representation. Consequently, ML is often not able to learn enough nonlinear representations to fit complicated maps out of the data which DL easily can. Furthermore, the end-to-end pipeline of DL models provides a reducing number of steps and so, errors in the pipeline.The flexibility and adaptability of DL models allows for transferring a model to new data or conducting new experiments easily. Such properties of reproducibility, causal reasoning and the process standardization together with the drastic computational cost drop for feature generation makes DL such a successful model in science.In [AdaBoost-DNN model] we adopt DL for drug-target interaction (DTI) predictive modelling, which allows for training by minimizing a loss function in an end-to-end fashion while aggregating low-level features from 3D molecular structures (molecular fingerprints) and high-level features of target proteins (protein-descriptor).Diversity of data often is a challenge arising with deep neural networks in that the networks get hard to train. In contrast to 2D molecular fingerprints, 3D descriptors are less common and hard to generate, which explains the limited number of protein-descriptor datasets. Our proposed AdaBoost-DNN model improves the DNN by aggregating weak models (Random Forest, Logistic Regression) into a strong model.In the goal of single embedding of weak classifiers, It has proven to capture the advantages of different input modalities and to make associations between them. It takes various input data types and accomplishes a method of transferring representations. In our architecture, the weak classifiers, encode the data from different domains and the DNN learns a common feature representation using the learned molecular descriptors and protein descriptors. In a systematic performance evaluation we match the resulting performance (accuracy, precision, recall, F1, ROC-AUC) with the performance of the well-established task learning methods Random Forest and Logistic Regression. With geometric feature precision, we evaluate the clusters of molecule-protein pairings that are predicted to be likely binders. The clear boundary between the two types of molecular targets may be a step for sending stronger evidence of the model’s correctness instead of just indiscriminative performance.

Downloads

Published

2026-09-22

How to Cite

Shahid, M., Bibhu, D. V., & Alam, D. Q. (2026). Deep Learning-Based Feature Learning for Drug Discovery and Biomedical Knowledge Extraction. International Journal of Artificial Intelligence and Machine Learning, 6(11s), 1556–1563. https://doi.org/10.51483/IJAIML.6.11s.2026.1556-1563