2609005143
  • Open Access
  • Article

Auditable Symbolic Candidate Generation for Biomedical Ontology Matching in Medical AI Systems

  • Ruichao Xia 1,   
  • Xingsi Xue 2,   
  • Haonan Li 3,   
  • Jianfeng Wang 4,*

Received: 01 Jul 2026 | Revised: 02 Sep 2026 | Accepted: 10 Sep 2026 | Published: 28 Sep 2026

Abstract

Biomedical ontology matching is a key step for semantic interoperability in medical AI systems, supporting clinical terminology mapping, cross-resource search, biomedical knowledge graph construction, and phenotype-aware data integration. This paper presents ARB, an auditable symbolic candidate-generation module designed for settings where external biomedical resources, learned embeddings, or reference-alignment calibration are unavailable or undesirable. ARB starts from a token inverted-index blocker, rescues missed candidates through hierarchy-pooled and neighbour-label evidence under a purity gate, and applies same-source structural dominance pruning to reduce lexically similar false siblings. The module is evaluated on seven OAEI 2013 biomedical ontology-matching tasks using candidate-pool metrics and a fixed downstream Hungarian matcher, and on five Bio-ML 2024 localranking tasks as a matcher-independent scoring check. Across the OAEI tasks, the rescue mechanism improves candidate recall by 2.0–10.8 percentage points while limiting gated pool growth to 5–24%. On Anatomy, structural dominance pruning alone raises downstream F1 from 0.197 to 0.366 under the fixed matcher, whereas the full ARB configuration (rescue plus pruning) reaches 0.375. Pruning is less reliable on structurally heterogeneous SNOMED-related tasks. On Bio-ML 2024, ARB achieves a meanMRRof 0.813 and mean Hits@1 of 0.760 using only ontologyinternal symbolic evidence. The module is most useful when lexical blocking leaves meaningful recall headroom; recall-saturated phenotype tasks instead require conservative rescue. These results support ARB as an auditable candidate-generation and cleaning component for biomedical knowledge integration, rather than as a complete ontology matcher or a final clinical mapping tool.

References 

  • 1.

    Abdaoui, H.; Barki, C.; Dergaa, I.; et al. Accurate Clinical Entity Recognition and Code Mapping of Anatomopathological Reports Using BioClinicalBERT Enhanced by Retrieval-Augmented Generation: A Hybrid Deep Learning Approach. Bioengineering 2026, 13, 30. https://doi.org/10.3390/bioengineering13010030.

  • 2.

    Jimenez-Ruiz, E.; Cuenca Grau, B. LogMap: Logic-Based and Scalable Ontology Matching. In Proceedings of the 10th International Semantic Web Conference (ISWC), Bonn, Germany, 23–27 October 2011; pp. 273–288.

  • 3.

    Faria, D.; Pesquita, C.; Santos, E.; et al. The AgreementMakerLight Ontology Matching System. In Proceedings of the OTM Confederated International Conferences, Graz, Austria, 9–13 September 2013; pp. 527–541.

  • 4.

    He, Y.; Chen, J.; Antonyrajah, D.; et al. BERTMap: A BERT-Based Ontology Alignment System. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual Event, 22 February–1 March 2022; Vol. 36, pp. 5684–5691.

  • 5.

    Faria, D.; Silva, M.C.; Cotovio, P.; et al. Results for Matcha and Matcha-DL in OAEI 2023. In Proceedings of the 18th International Workshop on Ontology Matching Co-Located with the 22nd International Semantic Web Conference (ISWC 2023), CEUR Workshop Proceedings, Athens, Greece, 7 November 2023; Vol. 3591, pp. 164–169.

  • 6.

    Papadakis, G.; Ioannou, E.; Niederee, C.; et al. Efficient Entity Resolution for Large Heterogeneous Information Spaces. In Proceedings of the 4th ACM International Conference on Web Search and Data Mining (WSDM), Hong Kong, China, 9–12 February 2011; pp. 535–544.

  • 7.

    Papadakis, G.; Papastefanatos, G.; Palpanas, T.; et al. Scaling Entity Resolution to Large, Heterogeneous Data with Enhanced Meta-Blocking. In Proceedings of the 19th International Conference on Extending Database Technology (EDBT), Bordeaux, France, 15–18 March 2016; pp. 221–232.

  • 8.

    Cuenca Grau, B.; Dragisic, Z.; Eckert, K.; et al. Results of the Ontology Alignment Evaluation Initiative 2013. In Proceedings of the 8th International Workshop on Ontology Matching Co-Located with the 12th International Semantic Web Conference (ISWC 2013), CEUR Workshop Proceedings, Sydney, Australia, 21 October 2013; Vol. 1111, pp. 61–100.

  • 9.

    Ontology Alignment Evaluation Initiative. Results of the Ontology Alignment Evaluation Initiative 2022. Available online: https://oaei.ontologymatching.org/2022/results/ (accessed on 21 June 2026).

  • 10.

    Ontology Alignment Evaluation Initiative. Results of the Ontology Alignment Evaluation Initiative 2023. Available online: https://oaei.ontologymatching.org/2023/results/ (accessed on 21 June 2026).

  • 11.

    Ontology Alignment Evaluation Initiative. Results of the Ontology Alignment Evaluation Initiative 2024. Available online: https://oaei.ontologymatching.org/2024/results/ (accessed on 21 June 2026).

  • 12.

    He, Y.; Chen, J.; Dong, H.; et al. Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching. In Proceedings of the 21st International Semantic Web Conference (ISWC), Hangzhou, China, 23–27 October 2022; pp. 575–591.

  • 13.

    Kebsi, D.; Barki, C.; Dergaa, I.; et al. Personalized Hearing Loss Care Using SNOMED CT-Aligned Ontology and Random Forest Machine Learning: A Hybrid Decision-Support Framework. Audiol. Res. 2026, 16, 37. https://doi.org/10.3390/audiolres16020037.

  • 14.

    Teymurova, S.; Jimenez-Ruiz, E.; Weyde, T.; et al. OWL2Vec4OA: Tailoring Knowledge Graph Embeddings for Ontology Alignment. In Proceedings of the International Knowledge Graphs and Semantic Web: 6th International Conference, KGSWC 2024, Lecture Notes in Computer Science, Paris, France, 11–13 December 2024; Vol. 15459, pp. 168–182.

  • 15.

    Hertling, S.; Paulheim, H. OLaLa: Ontology Matching with Large Language Models. In Proceedings of the 12th Knowledge Capture Conference 2023, Pensacola, FL, USA, 5–7 December 2023; pp. 131–139.

  • 16.

    Qiang, Z.; Wang, W.; Taylor, K. Agent-OM: Leveraging LLM Agents for Ontology Matching. In Proceedings of the VLDB Endowment, Guangzhou, China, 26–30 August 2024; Vol. 18, pp. 516–529.

  • 17.

    Babaei Giglou, H.; D’Souza, J.; Engel, F.; et al. LLMs4OM: Matching Ontologies with Large Language Models. In Proceedings of the Semantic Web: ESWC 2024 Satellite Events, Lecture Notes in Computer Science, Crete, Greece, 26–30 May 2024; Vol. 15344, pp. 25–35.

  • 18.

    Nguyen, L.; Barcelos, E.I.; French, R.H.; et al. KROMA: Ontology Matching with Knowledge Retrieval and Large Language Models. In Proceedings of the Semantic Web—ISWC 2025, Lecture Notes in Computer Science, Nara, Japan, 2–6 November 2025; Vol. 16140, pp. 629–649.

  • 19.

    Song, Y.; Chen, J.; Schmidt, R.A. GenOM: Ontology Matching with Description Generation and Large Language Models. World Wide Web 2026, 29, 29. https://doi.org/10.1007/s11280-026-01413-y.

  • 20.

    Qiang, Z.; Taylor, K.; Wang, W.; et al. OAEI-LLM: A Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching. In Proceedings of the Special Session on Harmonising Generative AI and Semantic Web Technologies (HGAIS 2024), CEUR Workshop Proceedings, Baltimore, MD, USA, 13 November 2024; Vol. 3953.

  • 21.

    Qiang, Z.; Taylor, K.; Wang, W.; et al. OAEI-LLM-T: A TBox Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching. arXiv 2025. arXiv:2503.21813.

  • 22.

    Boussi Rahmouni, H.; Hassine, N.B.E.H.; Chouchen, M.; et al. Healthcare 5.0-Driven Clinical Intelligence: The Learn-Predict-Monitor-Detect-Correct Framework for Systematic Artificial Intelligence Integration in Critical Care. Healthcare 2025, 13, 2553. https://doi.org/10.3390/healthcare13202553.

  • 23.

    Cheng, B.; Furst, J.; Jacobs, T.; et al. Interactive Ontology Matching with Cost-Efficient Learning. arXiv 2024. arXiv:2404.07663.

  • 24.

    Hernandez, M.A.; Stolfo, S.J. The Merge/Purge Problem for Large Databases. ACM SIGMOD Rec. 1995, 24, 127–138. https://doi.org/10.1145/568271.223807.

  • 25.

    Papadakis, G.; Skoutas, D.; Thanos, E.; et al. Blocking and Filtering Techniques for Entity Resolution: A Survey. ACM Comput. Surv. 2021, 53, 31:1–31:42. https://doi.org/10.1145/3377455.

  • 26.

    Zhang, W.; Wei, H.; Sisman, B.; et al. AutoBlock: A Hands-off Blocking Framework for Entity Matching. In Proceedings of the 13th ACM International Conference on Web Search and Data Mining (WSDM), Houston, TX, USA, 3–7 February 2020; pp. 744–752.

  • 27.

    Brinkmann, A.; Shraga, R.; Bizer, C. SC-Block: Supervised Contrastive Blocking within Entity Resolution Pipelines. In Proceedings of the Semantic Web: 21st International Conference, ESWC 2024, Hersonissos, Greece, 26–30 May 2024; pp. 121–142.

  • 28.

    Wang, T.; Lin, H.; Han, X.; et al. Towards Universal Dense Blocking for Entity Resolution. arXiv 2024. arXiv:2404.14831.

  • 29.

    Zhu, X.; Xie, M.; Deng, T.; et al. HyperBlocker: Accelerating Rule-Based Blocking in Entity Resolution Using GPUs. In Proceedings of the VLDB Endowment, Guangzhou, China, 26–30 August 2024; Vol. 18, pp. 308–321. https://doi.org/10.14778/3705829.3705847.

  • 30.

    Wang, R.; Li, Y.; Wang, J. Sudowoodo: Contrastive Self-supervised Learning for Multi-purpose Data Integration and Preparation. In Proceedings of the 39th IEEE International Conference on Data Engineering (ICDE), Anaheim, CA, USA, 3–7 April 2023; pp. 1502–1515.

  • 31.

    Melnik, S.; Garcia-Molina, H.; Rahm, E. Similarity Flooding: A Versatile Graph Matching Algorithm and Its Application to Schema Matching. In Proceedings of the 18th International Conference on Data Engineering (ICDE), San Jose, CA, USA, 26 February–1 March 2002; pp. 117–128.

  • 32.

    Efeoglu, S. GraphMatcher: A Graph Representation Learning Approach for Ontology Matching. In Proceedings of the 17th International Workshop on Ontology Matching, CEUR Workshop Proceedings, Hangzhou, China, 23 October 2022; Vol. 3324, pp. 174–180.

  • 33.

    Charikar, M.S. Similarity Estimation Techniques from Rounding Algorithms. In Proceedings of the 34th Annual ACM Symposium on Theory of Computing (STOC), Montreal, QC, Canada, 19–21 May 2002; pp. 380–388.

  • 34.

    Otsu, N. A Threshold Selection Method from Gray-Level Histograms. IEEE Trans. Syst. Man Cybern. 1979, 9, 62–66. https://doi.org/10.1109/TSMC.1979.4310076.

  • 35.

    Kuhn, H.W. The Hungarian Method for the Assignment Problem. Nav. Res. Logist. Q. 1955, 2, 83–97. https://doi.org/10.1002/nav.3800020109.

Share this article:
How to Cite
Xia, R.; Xue, X.; Li, H.; Wang, J. Auditable Symbolic Candidate Generation for Biomedical Ontology Matching in Medical AI Systems. AI Medicine 2026, 3 (2), 10. https://doi.org/10.53941/aim.2026.100010.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.
Article Metrics
10
Article Views
0
Citations