UNCERTAINTY-AWARE LEARNING OF SITES OF METABOLISM FROM LC-MS DATA USING GRAPH NEURAL NETWORKS

June 2026, ISSX 16th European Meeting

Ismael Zamora1; Ramon Adàlia1,2,3

1Mass Analytica, S.L., Sant Cugat del Vallés, Spain; 2Lead Molecular Design, SL, Sant Cugat del Vallès, Spain; 3Universitat Autònoma de Barcelona, Cerdanyola del Vallès, Spain;

 

Abstract

Predicting sites of metabolism (SoMs) is critical for guiding the design of metabolically stable drug candidates [1], yet early drug discovery relies heavily on liquid chromatography-mass spectrometry (LC-MS) data that often cannot unambiguously resolve exact atomic positions [2]. Conventional predictive models depend on definitive ground-truth labels, leading to the systematic exclusion of structurally ambiguous but experimentally informative observations. We hypothesized that explicitly modeling this uncertainty, rather than discarding it, would enable robust and practically useful SoM prediction.

To address this, we developed an uncertainty-aware framework in which candidate SoMs derived from LC-MS metabolite identification experiments (MassMetaSite/Oniro platform [3]) are represented as soft atom-level plausibility scores reflecting the frequency of atom involvement across equally scoring structural hypotheses. These scores were aggregated at the parent compound level following expert curation to remove clear artifacts while preserving ambiguity.

To control for label noise, metabolite observations annotating more than 40% of the atoms were excluded, with sensitivity analysis confirming optimal performance near this threshold. The resulting labels were used to train graph neural networks with four graph attention layers [4], incorporating atom-level chemical descriptors and MetaSite7 [5] reactivity features.

Training was formulated as a ranking problem using a pairwise RankNet [6] objective to preserve graded relevance, alongside a top-k-oriented LambdaGap-S+ [7] objective derived from thresholded labels.

On a soft-label human liver microsome dataset, the uncertainty-aware RankNet model with MetaSite7 features achieved strong performance (Spearman correlation 51.23%, NDCG 81.25%), substantially outperforming random and rule-based baselines.

Importantly, models trained exclusively on ambiguous LC-MS data generalized to an external benchmark of experimentally confirmed SoMs, where the LambdaGap-S+ objective achieved up to 75.0% top-2 accuracy on the XenoSite [8] dataset and significantly exceeded random performance.

These results demonstrate that the structural ambiguity inherent to LC-MS metabolite identification can be leveraged as a meaningful training signal rather than treated as a limitation. By aligning machine learning models with the realities of discovery-stage data, this framework enables continuous model development directly from routine metabolite identification workflows without requiring definitive structural elucidation, providing actionable insight into metabolic liability and supporting more efficient drug design.

CITATIONS

  1. Zhang, Z., & Tang, W. (2018). Drug metabolism in drug discovery and development. Acta Pharmaceutica Sinica B, 8(5), 721-732. https://doi.org/10.1016/j.apsb.2018.04.003
  2. Castro-Perez, J., & Prakash, C. (2020). Recent advances in mass spectrometric and other analytical techniques for the identification of drug metabolites. In Identification and Quantification of Drugs, Metabolites, Drug Metabolizing Enzymes, and Transporters (pp. 39-71). Elsevier. https://doi.org/10.1016/b978-0-12-820018-6.00002-8
  3. Zamora, I., Fontaine, F., Serra, B., & Plasencia, G. (2013). High-throughput, computer assisted, specific MetID. A revolution for drug discovery. Drug Discovery Today: Technologies, 10(1), e199-e205. https://doi.org/10.1016/j.ddtec.2012.10.015
  4. Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., & Bengio, Y. (2017). Graph Attention Networks (Version 3). arXiv. https://doi.org/10.48550/ARXIV.1710.10903
  5. Cruciani, G., Desantis, J., Palomba, T., Baroni, M., Valeri, A., Siragusa, L., Venturi, L., Zamora, I., Meyer, C., Laboureur, L., Emre, I., & Goracci, L. (2024). Decoding phase I & II human drug metabolism using the prediction tool MetaSite for chemists, medicinal chemists, and metID experts (Manuscript in preparation). Mass Analytica. Retrieved February 4, 2026, from https://mass-analytica.com/products/metasite7/
  6. Burges, C., Shaked, T., Renshaw, E., Lazier, A., Deeds, M., Hamilton, N., & Hullender, G. (2005). Learning to rank using gradient descent. In Proceedings of the 22nd International Conference on Machine Learning – ICML’05 (pp. 89-96). ACM Press. The 22nd International Conference. https://doi.org/10.1145/1102351.1102363
  7. Adalia, R., Sanjuan, G., Margalef, T., & Zamora, I. (2025). The LambdaGap Framework for Precision-Oriented Ranking. ACM Transactions on Information Systems, 43(4), 1-39. https://doi.org/10.1145/3733235
  8. Zaretzki, J., Matlock, M., & Swamidass, S. J. (2013). XenoSite: Accurately Predicting CYP-Mediated Sites of Metabolism with Neural Networks. Journal of Chemical Information and Modeling, 53(12), 3373-3383. https://doi.org/10.1021/ci400518g

 

You must be logged in to access this content. Not yet registered? Create a new account

 

 

Tags -