Cook, Frederika, Masala, Giovanni Luca (2026) Trustworthy Molecular AI: Multi-Dimensional Validation of LLM Generated Chemical Descriptions. In: Paszynski, Maciej and Barnard, Amanda S. and Zhang, Yongjie Jessica, eds. Computational Science – ICCS 2026 Workshops. 26th International Conference, ICCS 2026, Proceedings Part I. Lecture Notes in Computer Science Springer ISBN 978-3-032-29911-6. (doi:10.1007/978-3-032-29912-3_26) (KAR id:114278)
|
PDF
Author's Accepted Manuscript
Language: English
This work is licensed under a Creative Commons Attribution 4.0 International License.
|
|
|
Download this file (PDF/1MB) |
Preview |
| Request a format suitable for use with assistive technology e.g. a screenreader | |
| Official URL: https://doi.org/10.1007/978-3-032-29912-3_26 |
|
| Additional URLs: |
|
Abstract
Large Language Models (LLMs) offer unprecedented opportunities for molecular property prediction, yet exhibit concerning capabilities for generating plausible-sounding but factually incorrect chemical descriptions. Whilst traditional evaluation metrics measure linguistic similarity, they cannot detect chemically impossible claims. Domain benchmarks focus on entity-level verification rather than systematically validating whether claimed structures are computationally possible. We propose a multi-dimensional validation framework, integrating grammatical fluency and lexical validity checking, RDKit-based computational chemistry validation, and DrugBank verification. Evaluating three architecturally distinct LLMs (MolT5, LLaMA, TxGemma) across multiple decoding strategies on 451 antibiotic compounds, we demonstrate that multi-dimensional validation identifies distinct error types invisible to individual metrics. Our interactive dashboard enables inspection at multiple granularities, transforming validation from post-hoc comparison into evidence-based analysis. Our findings reveal systematic blind spots: perfect grammar accompanies fundamental chemical errors, aggregate metrics mask substantial performance differences, and domain-specific pre-training fails to transfer across chemistry tasks. This work sets foundations for trustworthy AI in chemistry, providing quality assurance infrastructure necessary for deploying LLMs in high-stakes applications that carry material consequences.
| Item Type: | Conference proceeding |
|---|---|
| DOI/Identification number: | 10.1007/978-3-032-29912-3_26 |
| Additional information: | For the purpose of open access, the author(s) has applied a Creative Commons Attribution (CC BY) licence to any Author Accepted Manuscript version arising. |
| Subjects: | Q Science > QA Mathematics (inc Computing science) > QA 76 Software, computer programming, |
| Institutional Unit: | Schools > School of Computing |
| Former Institutional Unit: |
There are no former institutional units.
|
| Funders: | University of Kent (https://ror.org/00xkeyj56) |
| Depositing User: | Giovanni Masala |
| Date Deposited: | 01 May 2026 12:17 UTC |
| Last Modified: | 11 Aug 2026 11:44 UTC |
| Resource URI: | https://kar.kent.ac.uk/id/eprint/114278 (The current URI for this page, for reference purposes) |
- Link to SensusAccess
- Export to:
- RefWorks
- EPrints3 XML
- BibTeX
- CSV
- Depositors only (login required):

https://orcid.org/0000-0001-6734-9424
Altmetric
Altmetric