Chemical Language Models for Activity-Guided De Novo Drug Discovery: A Hybrid Transfer Learning Framework for Intelligent Molecular Design and Virtual Screening

Authors

  • Maniteja Gorikapudi Department of Chemistry, Malla Reddy College of Engineering for Women, Maisammaguda, Hyderabad,Telangana 500100. Author
  • K. Sampath Kumar Department of Pharmaceutics, Malla Reddy Institute of Pharmaceutical Sciences, Malla Reddy Vishwavidyapeeth, Suraram, Hyderabad -500055, Telangana, India Author
Search on Google Scholar

Keywords:

  • Chemical Language Models; De novo Drug Design; Transfer Learning; SMILES; Virtual Screening.

Abstract

The identification of novel therapeutic molecules remains one of the most resource-intensive stages of pharmaceutical research owing to the enormous size of chemical space and the low probability of discovering compounds possessing favourable biological activity and acceptable pharmacokinetic characteristics. Recent advances in artificial intelligence have transformed computational drug discovery by enabling deep learning models to understand molecular representations and generate chemically valid structures. Among these approaches, Chemical Language Models (CLMs) have emerged as a powerful class of generative algorithms capable of learning molecular grammar directly from textual molecular representations such as Simplified Molecular Input Line Entry System (SMILES) strings. Unlike conventional quantitative structure–activity relationship models that depend primarily on predefined descriptors, CLMs autonomously capture complex structural relationships and generate novel molecular entities with desirable biological properties. The present study proposes a hybrid transfer-learning Chemical Language Model integrating molecular language modelling, perplexity-guided molecular prioritization, physicochemical property optimization, virtual screening, molecular docking, and explainable artificial intelligence for activity-guided de novo drug discovery. A transformer-based generative architecture was initially pretrained on a comprehensive molecular database to acquire generalized chemical knowledge and subsequently fine-tuned using experimentally validated phosphoinositide 3-kinase gamma (PI3Kγ) inhibitors. Molecular perplexity scoring was incorporated to rank generated compounds according to model confidence while simultaneously minimizing dataset bias. Newly generated molecules underwent sequential filtering using Lipinski's Rule of Five, Veber criteria, PAINS exclusion, synthetic accessibility assessment, ADMET prediction, molecular docking, and molecular dynamics simulation. Explainable artificial intelligence methods were further incorporated to interpret structural features contributing to predicted biological activity.The proposed computational framework demonstrated high molecular validity, structural novelty, scaffold diversity, and favourable predicted pharmacokinetic characteristics while substantially reducing computational screening efforts. Integration of transfer learning with perplexity-based prioritization improved candidate selection and reduced false-positive predictions compared with conventional generative models. The framework illustrates the growing potential of hybrid Chemical Language Models as intelligent molecular design platforms capable of accelerating early-stage drug discovery, particularly in therapeutic areas where experimentally validated datasets remain limited. The proposed workflow provides a scalable computational strategy for precision molecular design and supports future implementation of autonomous artificial intelligence systems in medicinal chemistry.

1

Downloads

Published

2026-08-10

How to Cite

Gorikapudi, M. G., & Kumar, K. S. K. (2026). Chemical Language Models for Activity-Guided De Novo Drug Discovery: A Hybrid Transfer Learning Framework for Intelligent Molecular Design and Virtual Screening. Interconnected Journal of Chemistry and Pharmaceutical Sciences (IJCPS), 2(2), 1-15. https://ijcps.nknpub.com/1/article/view/21