Structure-to-Shift Prediction for 19F NMR: Comparison of Graph Neural Networks, HOSE Code, and Large Language Models

O'Neill, Nathan, Jones, Adam P., Huber, Katharina T. ORCID: https://orcid.org/0000-0002-6368-7511, Kemsley, E. Kate ORCID: https://orcid.org/0000-0003-0669-3883, Cobas, Carlos, Williamson, David, Sharman, Gary and Hollerton, John (2026) Structure-to-Shift Prediction for 19F NMR: Comparison of Graph Neural Networks, HOSE Code, and Large Language Models. In: SMASH – Small molecule NMR conference 2026, 2026-09-13 - 2026-09-17, Hotel Alpenrock.

[thumbnail of SMASH_Poster_2026]
Preview
PDF (SMASH_Poster_2026)
Available under License Creative Commons Attribution.

Download (510kB) | Preview

Abstract

Accurate prediction of 19F NMR chemical shifts underpins structural verification of fluorine-containing compounds, of importance in the pharmaceutical and agrichemical industries. Here we compare three distinct approaches to structure-to-shift prediction: • Supervised machine learning using Graph Convolutional Neural Networks (GCNNs) of the message-passing type • Empirical database-based prediction using Mnova’s Modgraph NMRPredict, a commercially-available tool from SciY-Mestrelab • A zero-shot reasoning workflow generated by Anthropic's Claude-Fable large language model (LLM) The GCNN models were trained on a proprietary, curated collection of ~15,000 structures assigned with 19F chemical shifts, using an approach analogous to that described by Williamson et al [1]. The Mnova Modgraph predictor draws upon the same underlying data collection but employs a HOSE-code-based prediction strategy [2]. In contrast, Claude-Fable was prompted to devise and implement a prediction methodology based on analysis of molecular environments and chemical reasoning. All three methods predict chemical shifts directly from molecular structure and can therefore be applied to previously unseen compounds, a key requirement for discovery chemistry and quantitative downstream applications such as structural verification. Each method was applied to fluorine-containing compounds in the publicly available database, nmrshiftdb2. These were filtered to exclude any structures present in the GCNN/Modgraph training set, yielding a test set of 795 structures and 4599 19F chemical shift annotations. The respective Mean and Median Absolute Errors in prediction are summarized below: The GCNN model delivered the best overall performance, achieving the lowest Mean Absolute Error (5.98 ppm, approximately 25% lower than the other methods), while exhibiting a Median Absolute Error comparable to that of the established ModGraph predictor (2.96 versus 2.89 ppm). The GCNN and ModGraph predictors both achieved substantially lower Median Absolute Errors than Claude-Fable. These results indicate that, notwithstanding the impressive general capabilities of modern LLMs, dedicated NMR prediction engines trained on curated data collections remain advantageous for quantitative prediction tasks in NMR spectroscopy. References [1] D. Williamson, S. Ponte, I. Iglesias, N. Tonge, C. Cobas, and E. K. Kemsley, "Chemical shift prediction in 13C NMR spectroscopy using ensembles of message passing neural networks (MPNNs)" Journal of Magnetic Resonance, vol. 368, p. 107795, 2024/11/01/2024 [2] Kuhn, Stefan & Johnson, Sean. (2019). Stereo-Aware Extension of HOSE Codes. ACS Omega. 4. 10.1021/acsomega.9b00488.

Item Type: Conference or Workshop Item (Poster)
Faculty \ School: Faculty of Science > School of Chemistry, Pharmacy and Pharmacology
Faculty of Science
Faculty of Science > School of Computing Sciences
UEA Research Groups: Faculty of Science > Research Groups > Computational Biology
Depositing User: LivePure Connector
Date Deposited: 17 Sep 2026 13:16
Last Modified: 17 Sep 2026 13:16
URI: https://ueaeprints.uea.ac.uk/id/eprint/104577
DOI:

Downloads

Downloads per month over past year

Actions (login required)

View Item View Item