2606004459
  • Open Access
  • Article

Efficient Adaptation of Chemical Language Models for Molecular Property Prediction

  • Kavya Agrawal 1,†,   
  • Harsha Harod 1,†,   
  • Ines Simeone 2,   
  • Michele Ceccarelli 3,   
  • Halima Bensmail 4,   
  • Sukrit Gupta 1,   
  • Raghvendra Mall 4,5,*

Received: 05 May 2026 | Revised: 20 Jun 2026 | Accepted: 30 Jun 2026 | Published: 14 Jul 2026

Abstract

Chemical language models (CLMs) like ChemBERTa and Molformer enable compound property prediction, but their computational demands limit adoption in resource-constrained settings. We integrate Low-Rank Adapters (LoRA) with CLMs to significantly reduce trainable parameters required for finetuning. Results: We observe 3–5% area under the receiver operating curve (AUC) improvement across classification tasks for molecule toxicity, blood-brain barrier permeability, and flavor prediction over Molformer-XL. Our approach achieves Matthews correlation coefficient (MCC) scores of 0.80–0.90 across three tasks while reducing model parameters by 75–95%. By comparing embeddings from zero-shot and finetuned CLMs combined with molecular physicochemical properties, we attain optimal performance across four datasets. This lightweight adaptation retains performance efficiency while reducing over-parameterization.

Graphical Abstract

References 

  • 1.

    Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008.

  • 2.

    Ahmad, W.; Simon, E.; Chithrananda, S.; et al. ChemBERTa-2: Towards Chemical Foundation Models. arXiv 2022, arXiv:2209.01712.

  • 3.

    Huang, K.; Fu, T.; Glass, L.M.; et al. DeepPurpose: A Deep Learning Library for Drug–Target Interaction Prediction. Bioinformatics 2020, 36, 5545–5547. https://doi.org/10.1093/bioinformatics/btaa1005.

  • 4.

    Ross, J.; Belgodere, B.; Chenthamarakshan, V.; et al. Large-Scale Chemical Language Representations Capture Molecular Structure and Properties. Nat. Mach. Intell. 2022, 4, 1256–1264. https://doi.org/10.1038/s42256-022-00580-7.

  • 5.

    Chithrananda, S.; Grand, G.; Ramsundar, B. ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction. arXiv 2020, arXiv:2010.09885.

  • 6.

    Likhosherstov, V.; Choromanski, K.; Dubey, A.; et al. Favor#: Sharp Attention Kernel Approximations via New Classes of Positive Random Features. arXiv 2023, arXiv:2302.00787.

  • 7.

    Rong, Y.; Bian, Y.; Xu, T.; et al. Self-Supervised Graph Transformer on Large-Scale Molecular Data. Adv. Neural Inf. Process. Syst. 2020; 33, 12559–12571.

  • 8.

    Liu, S.; Wang, H.; Liu, W.; et al. Pre-Training Molecular Graph Representation with 3D Geometry. In Proceedings of the International Conference on Learning Representations, online, 25–29 April 2022.

  • 9.

    Zhou, G.; Gao, Z.; Ding, Q.; et al. Uni-Mol: A Universal 3D Molecular Representation Learning Framework. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023.

  • 10.

    Akiba, T.; Sano, S.; Yanase, T.; et al. Optuna: A Next-Generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 4–8 August 2019; pp. 2623–2631. https://doi.org/10.1145/3292500.3330701.

  • 11.

    Martins, I.F.; Teixeira, A.L.; Pinheiro, L.; et al. A Bayesian Approach to In Silico Blood-Brain Barrier Penetration Modeling. J. Chem. Inf. Model. 2012, 52, 1686–1697. https://doi.org/10.1021/ci300124c.

  • 12.

    Wu, Z.; Ramsundar, B.; Feinberg, E.N.; et al. MoleculeNet: A Benchmark for Molecular Machine Learning. Chem. Sci. 2018, 9, 513–530. https://doi.org/10.1039/C7SC02664A.

  • 13.

    Novick, P.A.; Ortiz, O.F.; Poelman, J.; et al. SweetLead: An In Silico Database of Approved Drugs, Regulated Chemicals, and Herbal Isolates for Computer-Aided Drug Discovery. PLoS One 2013, 8, e79568. https://doi.org/10.1371/journal.pone.0079568.

  • 14.

    Zimmermann, Y.; Sieben, L.; Seng, H.; et al. A Chemical Language Model for Multi-Class Molecular Taste Prediction. NPJ Sci. Food 2025, 9, 122. https://doi.org/10.1038/s41538-025-00474-z.

  • 15.

    Ramsundar, B.; Eastman, P.;Walters, P.; et al. Deep Learning for the Life Sciences; O’ReillyMedia: Sebastopol, CA, USA, 2019.

  • 16.

    Kim, S.; Chen, J.; Cheng, T.; et al. PubChem 2023 Update. Nucleic Acids Res. 2023, 51, D1373–D1380. https://doi.org/10.1093/nar/gkac956.

  • 17.

    Han, X.; Jia, M.; Chang, Y.; et al. Directed Message Passing Neural Network (D-MPNN) with Graph Edge Attention (GEA) for Property Prediction of Biofuel-Relevant Species. Energy AI 2022, 10, 100201. https://doi.org/10.1016/j.egyai.2022.100201.

  • 18.

    Pan, Z. Large Language Model for Molecular Chemistry. Nat. Comput. Sci. 2023, 3, 5.

  • 19.

    Heo, B.; Park, S.; Han, D.; et al. Rotary Position Embedding for Vision Transformer. In European Conference on Computer Vision; Springer: Cham, Switzerland, 2024; pp. 289–305.

  • 20.

    Isard, M.; Yu, Y. Distributed Data-Parallel Computing Using a High-Level Programming Language. In Proceedings of the 2009 ACM SIGMOD International Conference on Management of Data, Providence, RI, USA, 29 June–2 July 2009; pp. 987–994.

  • 21.

    Irwin, J.J.; Tang, K.G.; Young, J.; et al. ZINC20—A Free Ultralarge-Scale Chemical Database for Ligand Discovery. J. Chem. Inf. Model. 2020, 60, 6065–6073. https://doi.org/10.1021/acs.jcim.0c00675.

  • 22.

    Ruddigkeit, L.; van Deursen, R.; Blum, L.C.; et al. Enumeration of 166 Billion Organic Small Molecules in the Chemical Universe Database GDB-17. J. Chem. Inf. Model. 2012, 52, 2864–2875. https://doi.org/10.1021/ci300415d.

  • 23.

    Schuh, M.G.; Boldini, D.; Sieber, S.A. Synergizing Chemical Structures and Bioassay Descriptions for Enhanced Molecular Property Prediction in Drug Discovery. J. Chem. Inf. Model. 2024, 64, 4640–4650.

  • 24.

    Nowakowska, S. ChemBERTa-2: Fine-Tuning for Molecule’s HIV Replication Inhibition Prediction. ChemRxiv 2023. https://doi.org/10.26434/chemrxiv-2023-b57vx.

  • 25.

    Lee, D.; Cho, Y. Fine-Tuning Pocket-Conditioned 3D Molecule Generation via Reinforcement Learning. In Proceedings of the ICLR 2024Workshop on Generative and Experimental Perspectives for Biomolecular Design, Vienna, Austria, 7–11 May 2024.

  • 26.

    Balne, C.C.S.; Bhaduri, S.; Roy, T.; et al. Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications. arXiv 2024, arXiv:2404.13506.

  • 27.

    Zhou, H.; Wan, X.; Vuli´c, I.; et al. AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning. Trans. Assoc. Comput. Linguist. 2024, 12, 525–542.

  • 28.

    Fu, Z.; Yang, H.; So, A.M.C.; et al. On the Effectiveness of Parameter-Efficient Fine-Tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA,7–14 February 2023; pp. 12799–12807.

  • 29.

    Xu, L.; Xie, H.; Qin, S.Z.J.; et al. Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 6107–6126.

  • 30.

    Hu, E.J.; Shen, Y.; Wallis, P.; et al. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the 2022 International Conference on Learning Representations (ICLR), Online, 25 April 2022.

  • 31.

    Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794.

  • 32.

    Ke, G.; Meng, Q.; Finley, T.; et al. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017; pp.3146–3154.

  • 33.

    Prokhorenkova, L.; Gusev, G.; Vorobev, A.; et al. CatBoost: Unbiased Boosting with Categorical Features. In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, QC, Canada, 3–8 December 2018; Volume 31, pp. 6638–6648.

  • 34.

    Chang, J.; Ye, J.C. Bidirectional Generation of Structure and Properties Through a Single Molecular Foundation Model. Nat. Commun. 2024, 15, 2323.

  • 35.

    Mall, R.; Elbasir, A.; Almeer, H.; et al. A Modelling Framework for Embedding-Based Predictions for Compound-Viral Protein Activity. Bioinformatics 2021, 37, 2544–2555.

  • 36.

    Alhosani, H.; Mall, R.; Singh, A.; et al. Feature-PHLA: Physico-Chemical Features Efficiently Predict Peptide-HLA Binding Affinity. In Proceedings of the 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Lisbon, Portugal, 3–6 December 2024; pp. 4–11.

  • 37.

    Mall, R.; Singh, A.; Patel, C.N.; et al. VISH-Pred: An Ensemble of Fine-Tuned ESM Models for Protein Toxicity Prediction. Brief. Bioinform. 2024, 25, bbae243.

  • 38.

    Khurana, S.; Rawi, R.; Kunji, K.; et al. DeepSol: A Deep Learning Framework for Sequence-Based Protein Solubility Prediction. Bioinformatics 2018, 34, 2605–2613.

  • 39.

    Elbasir, A.; Mall, R.; Kunji, K.; et al. BCrystal: An Interpretable Sequence-Based Protein Crystallization Predictor. Bioinformatics 2020, 36, 1429–1438.

  • 40.

    Jethalia, M.; Jani, S.P.; Ceccarelli, M.; et al. PanCancer Network Analysis Reveals Key Master Regulators for Cancer Invasiveness. J. Transl. Med. 2023, 21, 558.

  • 41.

    Mall, R.; Kaushik, R.; Martinez, Z.A.; et al. Benchmarking Protein Language Models for Protein Crystallization. Sci. Rep. 2025, 15, 2381.

  • 42.

    Liu, S.; Nie, W.; Wang, C.; et al. Multi-Modal Molecule Structure–Text Model for Text-Based Retrieval and Editing. Nat. Mach. Intell. 2023, 5, 1447–1457.

  • 43.

    Edwards, C.; Lai, T.; Ros, K.; et al. Translation Between Molecules and Natural Language. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022.

  • 44.

    Xia, J.; Zhao, C.; Hu, B.; et al. Mole-BERT: Rethinking Pre-Training Graph Neural Networks for Molecules. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023.

  • 45.

    Livne, M.; Miftahutdinov, Z.; Tutubalina, E.; et al. Nach0: Multimodal Natural and Chemical Languages Foundation Model. Chem. Sci. 2024, 15, 8380–8389.

  • 46.

    Hu, C.; Li, H.; Yuan, Y.; et al. Omni-Mol: Multitask Molecular Model for Any-to-Any Modalities. Adv. Neural Inf. Process. Syst. 2026, 38, 66945–66988.

  • 47.

    Maleki, S.; Huetter, J.C.; Chuang, K.V.; et al. Efficient Fine-Tuning of Single-Cell Foundation Models Enables Zero-Shot Molecular Perturbation Prediction. arXiv 2024, arXiv:2412.13478.

  • 48.

    Ren, W.; Li, X.; Wang, L.; et al. Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning. arXiv 2024, arXiv:2402.18865.

Share this article:
How to Cite
Agrawal, K.; Harod, H.; Simeone, I.; Ceccarelli, M.; Bensmail, H.; Gupta, S.; Mall, R. Efficient Adaptation of Chemical Language Models for Molecular Property Prediction. Translational Insights 2026, 1 (1), 13. https://doi.org/10.53941/ti.2026.100013.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.