REVENG: 3D-DRIVEN REVERSE ENGINEERING OF METABOLIC BIOLOGICAL PATHWAYS

REVENG: 3D-DRIVEN REVERSE ENGINEERING OF METABOLIC BIOLOGICAL PATHWAYS

September 2026, 25th EuroQsar

Tommaso Palomba1,2, Savannah Mason1, Paula Cifuentes1, Ismael Zamora1, Gabriele Cruciani2

1Mass Analytica, S.L., Sant Cugat del Vallés, Spain; 2Molecular Discovery Ltd, Kinetic Business Centre, Borehamwood, UK

Abstract

Elucidating the specific enzymes responsible for metabolite formation is a cornerstone of drug discovery; it is essential for mapping metabolic pathways, assessing drug-drug interactions, driving lead optimization to enhance metabolic properties, and mitigating toxicity risk. While liquid chromatography-mass spectrometry (LC-MS/MS) excels at structural Metabolite Identification (MetID), experimental reaction phenotyping remains a major resource-intensive bottleneck. To bridge this gap, we present RevEng, an innovative computational workflow integrated within MetaSite 71. Taking the parent drug and a target metabolite (for example, an experimentally detected LC-MS/MS metabolite) as inputs, RevEng structurally reverse-engineers the metabolic space to reconstruct all potential metabolic pathways, including complex multi-step biotransformations, and assigns a probability score to the enzymes involved in each step. This inverse pathway prediction relies entirely on structure-based 3D docking and intrinsic chemical reactivity, using dedicated enzyme-specific predictive models for 23 Phase I (CYP and non-CYP) and 20 Phase II major human drug-metabolizing enzymes2,3. These models are interaction- and reaction-based and are therefore fully independent of training-set-derived substrate knowledge. Enzyme active sites are characterized using GRID Molecular Interaction Fields (MIFs)4, while the docking procedure is specifically designed to mimic the biological and chemical determinants of each enzymatic reaction by evaluating substrate–residue interactions within the catalytic site. The resulting interaction energy, combined with an assessment of the molecule’s intrinsic reactivity toward a specific biotransformation, enables prediction of the most likely site of metabolism, related metabolites structures and the likelihood of metabolism by the enzyme under consideration. The workflow was rigorously validated against proprietary in-house experimental datasets, and the predicted inverse pathways were benchmarked against established literature data for diverse chemical entities. The 3D-docking engine successfully unraveled complex metabolic fates, accurately backtracking both single-step biotransformations and sequential, multi-step cascades (such as Phase I functionalization followed by Phase II conjugation). Beyond pathway elucidation, the prediction of the enzymes responsible for the formation of observed metabolites can directly drive experimental design, enabling the detection of metabolites using targeted enzymatic assays rather than complex biological matrices, thus leading to a better characterization of the metabolic profile of a drug. Furthermore, these predictions provide invaluable guidance for targeted metabolite synthesis and for anticipating potential toxicity liabilities. In conclusion, RevEng, in combination with the other advanced tools integrated within MetaSite 7, successfully bridges analytical chemistry and mechanism-based 3D docking, offering a chemically accurate and structurally-driven solution for automated reaction phenotyping and intelligent experimental design in early drug discovery.

  1. Isoherranen, N., & Zanger, U. M. (2025). Cytochrome P450 enzymes in drug metabolism and precision medicine: Current challenges and future directions. ACS Pharmacology & Translational Science, 8(2), 114-131.
  2. Cruciani, G., Carosati, E., De Boeck, B., Ethirajulu, K., Mackie, C., Howe, T., & Vianello, R. (2005). MetaSite: Understanding Metabolism in Human Cytochromes from the Perspective of the Chemist. Journal of Medicinal Chemistry, 48(22), 6970-6979.
  3. Di, L., & Obach, R. S. (2024). The evolving landscape of non-CYP enzymes in drug metabolism: Current status and clinical relevance. Drug Metabolism Reviews, 56(2), 145-168.
  4. Artese, A., Cross, S., Costa, G., Distinto, S., Parrotta, L., Alcaro, S., Ortuso, F., & Cruciani, G. (2013). Molecular interaction fields in drug discovery: recent advances and future perspectives. WIREs Computational Molecular Science, 3(6), 594-613.

You must be logged in to access this content. Not yet registered? Create a new account

 

 

PREDICTION OF PHASE I AND PHASE II METABOLISM

PREDICTION OF PHASE I AND PHASE II METABOLISM

June 2026, ISSX 16th European Meeting

Tommaso Palomba1,2, Ludovico Venturi1,2, Massimo Baroni1, Paolo Benedetti1, Gabriele Cruciani1,2,3

1Mass Analytica, S.L. Sant Cugat del Vallès, Spain; 2Molecular Discovery Ltd, Kinetic Business Centre, Borehamwood, UK; 3Department of Chemistry, Biology, and Biotechnology, University of Perugia, Italy

Abstract

Many drugs are metabolized in the human body to form metabolites for elimination. This metabolism is categorized into two classes of reactions. Phase I reactions are mediated by Cytochrome P450 enzymes, in addition to approximately thirty other enzymes. An additional twenty other enzymes are involved in Phase II reactions, which often result in conjugation products.

While the cytochrome P450 enzymes have long been studied, there is a growing need to better understand the role of non-CYP enzymes in drug metabolism. To facilitate this, we developed MetaSite 7 to provide the prediction of drug metabolism using 23 non-CYP Phase I enzymes including eight families and their related isoforms. Twenty enzymes, including 12 different families, responsible for Phase II reactions were also considered.

The prediction determines the probability of a molecule to be a substrate for each enzyme by identifying reactive atoms in the molecule structure. In MetaSite 7, we used the enzyme’s 3D structure to identify reactive atoms of the molecule which may be exposed to the cofactor or active binding site of the enzyme. The GRID force field is used to calculate flexible interaction fields between the catalytic residues of the enzyme and the molecule.

Multiple docking poses are considered, and the energy contributions are summed. The poses are ranked, and the lowest energy pose is used to generate a [1,2,3]This prediction can be repeated for each of the enzymes in MetaSite 7, resulting in a more comprehensive view of the site of metabolism (SOM) for both Phase I and Phase II enzymes.

Furthermore, the predictions may be combined by subjecting the potential first-generation metabolites to further predictions, allowing for a more complete view of second-generation metabolites.

In conclusion, we utilized the 3D structures of the different CYP isoforms and 43 non-CYP enzymes to predict both Phase I and Phase II metabolism, along with the probability of a drug molecule to be a substrate for each enzyme. These predictions provide an increased understanding of the potential metabolic pathway in the human body.

You must be logged in to access this content. Not yet registered? Create a new account

 

 

GRAPHORMER BASED PREDICTION OF PEPTIDE METABOLIC STABILITY AND PHARMACOKINETIC RELATED PROPERTIES

GRAPHORMER BASED PREDICTION OF PEPTIDE METABOLIC STABILITY AND PHARMACOKINETIC RELATED PROPERTIES

June 2026, ISSX 16th European Meeting

Ismael Zamora1, Albert Garriga1, Luca Morettoni1, Paula Cifuentes2,3, Ramon Adàlia2,4

1Mass Analytica, S.L., Sant Cugat del Vallés, Spain; 2Lead Molecular Design, SL, Sant Cugat del Vallès, Spain 3Universitat Pompeu Fabra, Barcelona, 08003, Spain, 4Universitat Autònoma de Barcelona, Cerdanyola del Vallès, Spain

 

Abstract

Peptides represent a promising class of therapeutic agents due to their high target specificity, favorable safety profiles, and efficient tissue penetration. However, their clinical development remains limited by poor in vivo stability, rapid enzymatic degradation, and short plasma half-lives. Structural flexibility and high solvent accessibility further increase susceptibility to proteolysis, resulting in rapid clearance and reduced therapeutic efficacy. While chemical strategies such as cyclization and incorporation of non-natural amino acids can improve peptide stability, it remains challenging prior to synthesis to identify peptides that would benefit from such modifications based on metabolic stability and pharmacokinetic properties.

Here, we present a Graphormer-based machine learning framework for predicting key properties relevant to peptide metabolic stability, including proteolytic cleavage sites and peptide blood stability, as well as pharmacokinetic properties such as permeability. In addition, we demonstrate the ability of the approach to predict other peptide properties, including solvent accessibility, secondary structure, and post-translational modifications, which are included to assess the model’s capacity to capture structural information relevant for predicting metabolism- and pharmacokinetics-related properties.

The models employ transformer architecture with added mechanisms to encode graph structural information. Importantly, unlike most existing state-of-the-art tools, our approach is not restricted to natural amino acids and supports cyclic peptides. Depending on the target property, the Graphormer architecture was adapted for different downstream tasks, including binary or multiclass classification, ranking, and regression.

The Graphormer site of cleavage ranking model was evaluated under 28 datasets collected from MEROPS, which compiled experimentally verified cleavage sites from proteases involved in peptide-drug degradation. The site-of-cleavage ranking model was evaluated on 28 MEROPS datasets of experimentally verified protease cleavage sites involved in peptide-drug degradation, achieving an average MAP of 0.59 and Precision@1 of 0.47, outperforming ProsperousPlus.¹

The peptide blood stability model, trained on 635 peptides up to 50 amino acids including cyclic and modified peptides, achieved comparable performance to PepMSND², with accuracy 0.84, precision 0.85, recall 0.81, F1 score 0.81, and MCC 0.70.

For permeability, trained and tested on 6,888 cyclic peptides from the CycPeptMP database³, the model achieved R² = 0.77, MAE = 0.36, and MSE = 0.26, showing good generalization to structurally diverse peptides.

For solvent accessibility, formulated as buried versus exposed residues and validated on protein sequences up to 1,000 amino acids, the model achieved recall 0.81 and F1 score 0.78, comparable to existing methods.⁴

For secondary structure prediction evaluated using Jensen-Shannon distance across helix, strand, and coil states, the model achieved the lowest JSD for coil (0.207), outperforming all compared methods⁵, and improved over JPred4 and PSIPRED for helix and strand while narrowing the gap with PEP2D.

Finally, in predicting post-translational modifications across seven PTM types, the binary model outperformed state-of-the-art method⁶ for lysine acetylation, SUMOylation, ubiquitination, and arginine methylation, and achieved comparable performance for the remaining classes, demonstrating robustness across diverse biochemical modifications.

You must be logged in to access this content. Not yet registered? Create a new account

 

 

REVENG: REVERSE ENGINEERING THE METABOLIC BIOLOGICAL PATHWAY

REVENG: REVERSE ENGINEERING THE METABOLIC BIOLOGICAL PATHWAY

June 2026, ISSX 16th European Meeting

Tommaso Palomba1, Savannah Mason1, Paula Cifuentes1,3, Ludovico Venturi1,2, Ismael Zamora1

1Mass Analytica, S.L., Sant Cugat del Vallés, Spain; 2Molecular Discovery Ltd, Kinetic Business Centre, Borehamwood, UK; 3Universitat Pompeu Fabra, Barcelona, 08003, Spain

Abstract

Most drugs undergo chemical transformations in the body, known as biotransformations, to generate metabolites that are more readily eliminated. These reactions are primarily mediated by metabolic enzymes, mainly in the liver, and are highly specific, with each enzyme favoring particular substrates. Identifying the enzymes responsible for metabolite formation is essential for elucidating metabolic pathways, predicting metabolic behavior, and assessing potential toxicity risks. Metabolite Identification (MetID) studies, typically conducted in vitro or in vivo, rely heavily on LC-MS/MS for detecting and structurally characterizing metabolites. However, most discovery studies provide limited insight into the enzymes involved, and experimental approaches for reaction phenotyping, such as recombinant enzyme incubations or chemical inhibition studies, are often time- and resource-intensive, limiting comprehensive pathway characterization.

To address this gap, we developed a workflow that integrates LC-MS/MS MetID data from in vitro incubations with computational predictions to identify enzymatic pathways responsible for metabolite formation, including both Phase I and Phase II reactions. The computational approach simulates interactions between xenobiotic compounds and human metabolic enzymes using their 3D structures, evaluating the exposure of reactive atoms to catalytic residues. Multiple docking poses are generated and scored based on energy contributions, and the best pose is normalized to provide a probability ranking.

MetID data were processed using MassMetaSite on the ONIRO server with LC-MS/MS datasets from Sciex and Thermo instruments, and predictions were performed using MetaSite 7 within Oniro. The model was applied to three compounds and validated against experimental CYP phenotyping data from the literature.

For dextromethorphan, Phase I metabolism studies using recombinant enzymes and human liver microsomes identified six metabolites formed through N- and O-dealkylation and hydroxylation after 30 minutes. Experimentally, CYP2D6 mediated O-dealkylation to dextrorphan, while CYP3A4 catalyzed N-dealkylation to 3-methoxymorphinan. The model accurately predicted both pathways with high probability, consistent with the experimental results.¹

Phase II metabolism of dextromethorphan was further evaluated in hepatocytes, and revealed a metabolite formed via O-dealkylation followed by glucuronidation after 140 minutes; the model correctly suggested CYP2D6 for the initial step and UGT enzymes for conjugation, in agreement with literature reports.²

Two additional compounds were also assessed for Phase I metabolism: thioridazine metabolism in human liver microsomes produced metabolites via S-oxidation and N-demethylation, with the model identifying CYP3A4 for 5-sulfoxide formation and suggested CYP1A2 as one enzyme involved in N-demethylation. The model also predicted CYP2D6 involvement in the formation of mesoridazine and its further conversion to sulphoridazine, consistent with published data.³ ⁴

For ethoxyresorufin, incubated with recombinant enzymes for 30 minutes, O-dealkylation to resorufin was observed in CYP1A2 incubations, and the model predicted involvement of both CYP1A2 and CYP1A1.⁵

Together, these results demonstrate the capability of the in silico functionality integrated within a MetID platform to predict Phase I and Phase II enzymatic pathways for LC-MS/MS identified metabolites.

 

You must be logged in to access this content. Not yet registered? Create a new account

 

 

UNCERTAINTY-AWARE LEARNING OF SITES OF METABOLISM FROM LC-MS DATA USING GRAPH NEURAL NETWORKS

UNCERTAINTY-AWARE LEARNING OF SITES OF METABOLISM FROM LC-MS DATA USING GRAPH NEURAL NETWORKS

June 2026, ISSX 16th European Meeting

Ismael Zamora1; Ramon Adàlia1,2,3

1Mass Analytica, S.L., Sant Cugat del Vallés, Spain; 2Lead Molecular Design, SL, Sant Cugat del Vallès, Spain; 3Universitat Autònoma de Barcelona, Cerdanyola del Vallès, Spain;

 

Abstract

Predicting sites of metabolism (SoMs) is critical for guiding the design of metabolically stable drug candidates [1], yet early drug discovery relies heavily on liquid chromatography-mass spectrometry (LC-MS) data that often cannot unambiguously resolve exact atomic positions [2]. Conventional predictive models depend on definitive ground-truth labels, leading to the systematic exclusion of structurally ambiguous but experimentally informative observations. We hypothesized that explicitly modeling this uncertainty, rather than discarding it, would enable robust and practically useful SoM prediction.

To address this, we developed an uncertainty-aware framework in which candidate SoMs derived from LC-MS metabolite identification experiments (MassMetaSite/Oniro platform [3]) are represented as soft atom-level plausibility scores reflecting the frequency of atom involvement across equally scoring structural hypotheses. These scores were aggregated at the parent compound level following expert curation to remove clear artifacts while preserving ambiguity.

To control for label noise, metabolite observations annotating more than 40% of the atoms were excluded, with sensitivity analysis confirming optimal performance near this threshold. The resulting labels were used to train graph neural networks with four graph attention layers [4], incorporating atom-level chemical descriptors and MetaSite7 [5] reactivity features.

Training was formulated as a ranking problem using a pairwise RankNet [6] objective to preserve graded relevance, alongside a top-k-oriented LambdaGap-S+ [7] objective derived from thresholded labels.

On a soft-label human liver microsome dataset, the uncertainty-aware RankNet model with MetaSite7 features achieved strong performance (Spearman correlation 51.23%, NDCG 81.25%), substantially outperforming random and rule-based baselines.

Importantly, models trained exclusively on ambiguous LC-MS data generalized to an external benchmark of experimentally confirmed SoMs, where the LambdaGap-S+ objective achieved up to 75.0% top-2 accuracy on the XenoSite [8] dataset and significantly exceeded random performance.

These results demonstrate that the structural ambiguity inherent to LC-MS metabolite identification can be leveraged as a meaningful training signal rather than treated as a limitation. By aligning machine learning models with the realities of discovery-stage data, this framework enables continuous model development directly from routine metabolite identification workflows without requiring definitive structural elucidation, providing actionable insight into metabolic liability and supporting more efficient drug design.

CITATIONS

  1. Zhang, Z., & Tang, W. (2018). Drug metabolism in drug discovery and development. Acta Pharmaceutica Sinica B, 8(5), 721-732. https://doi.org/10.1016/j.apsb.2018.04.003
  2. Castro-Perez, J., & Prakash, C. (2020). Recent advances in mass spectrometric and other analytical techniques for the identification of drug metabolites. In Identification and Quantification of Drugs, Metabolites, Drug Metabolizing Enzymes, and Transporters (pp. 39-71). Elsevier. https://doi.org/10.1016/b978-0-12-820018-6.00002-8
  3. Zamora, I., Fontaine, F., Serra, B., & Plasencia, G. (2013). High-throughput, computer assisted, specific MetID. A revolution for drug discovery. Drug Discovery Today: Technologies, 10(1), e199-e205. https://doi.org/10.1016/j.ddtec.2012.10.015
  4. Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., & Bengio, Y. (2017). Graph Attention Networks (Version 3). arXiv. https://doi.org/10.48550/ARXIV.1710.10903
  5. Cruciani, G., Desantis, J., Palomba, T., Baroni, M., Valeri, A., Siragusa, L., Venturi, L., Zamora, I., Meyer, C., Laboureur, L., Emre, I., & Goracci, L. (2024). Decoding phase I & II human drug metabolism using the prediction tool MetaSite for chemists, medicinal chemists, and metID experts (Manuscript in preparation). Mass Analytica. Retrieved February 4, 2026, from https://mass-analytica.com/products/metasite7/
  6. Burges, C., Shaked, T., Renshaw, E., Lazier, A., Deeds, M., Hamilton, N., & Hullender, G. (2005). Learning to rank using gradient descent. In Proceedings of the 22nd International Conference on Machine Learning – ICML’05 (pp. 89-96). ACM Press. The 22nd International Conference. https://doi.org/10.1145/1102351.1102363
  7. Adalia, R., Sanjuan, G., Margalef, T., & Zamora, I. (2025). The LambdaGap Framework for Precision-Oriented Ranking. ACM Transactions on Information Systems, 43(4), 1-39. https://doi.org/10.1145/3733235
  8. Zaretzki, J., Matlock, M., & Swamidass, S. J. (2013). XenoSite: Accurately Predicting CYP-Mediated Sites of Metabolism with Neural Networks. Journal of Chemical Information and Modeling, 53(12), 3373-3383. https://doi.org/10.1021/ci400518g

 

You must be logged in to access this content. Not yet registered? Create a new account