2607004571
  • Open Access
  • Article

LLM Wiki as a Living Constraint Oracle for Safe Offline RL: Evidence from Type 1 Diabetes Glucose Management

  • Bailing Zhang

Received: 28 Apr 2026 | Revised: 01 Jun 2026 | Accepted: 08 Jul 2026 | Published: 22 Jul 2026

Abstract

Offline reinforcement learning (RL) for high-stakes clinical domains faces a persistent gap: the constraint knowledge encoded in clinical guidelines is rarely available as structured training signal. Existing approaches either hard-code constraints as fixed thresholds or retrieve them ad hoc at query time via retrieval-augmented generation (RAG), neither of which supports systematic maintenance as guidelines evolve. I propose treating a persistently maintained LLM wiki as a living constraint oracle (LCO) for safe offline RL policy evaluation and training. The LCO operates through three components: (1) a Wiki-to-Constraint (W2C) extraction pipeline that converts wiki pages with evidence scores and contradiction flags into a weighted constraint set; (2) WikiConstraint-IQL, which injects wiki-derived penalty terms into IQL training without architectural change; and (3) Dynamic Constraint Update (DCU), which propagates guideline revisions to policy evaluation without retraining. Applying this framework to type 1 diabetes (T1D) glucose management on the OhioT1DM dataset, I find that wiki-constrained policy reduces dangerous insulin administration during hypoglycemia from 98.6% to 81.3% of affected timesteps compared to unconstrained IQL, and that a simulated wiki update from ADA 2021 to ADA 2024 standards widens the safety gap between policies by 7.3× without any model retraining. Analysis of the extracted constraint set reveals that only 10.6% of the 1907 constraints extracted from 17 ADA guideline sections are clinically numeric, and only 1.3% can be mapped to step-level executable checks—an empirical characterisation of the semantic gap between narrative clinical guidelines and computational safety constraints.

References 

  • 1.

    Kostrikov, I.; Nair, A.; Levine, S. Offline Reinforcement Learning with Implicit Q-Learning. In Proceedings of the International Conference on Learning Representations (ICLR), Online, 25–29 April 2022.

  • 2.

    Kumar, A.; Zhou, A.; Tucker, G.; et al. Conservative Q-Learning for Offline Reinforcement Learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 6–12 December 2020; Volume 33, pp. 1179–1191.

  • 3.

    American Diabetes Association Professional Practice Committee. Standards of Care in Diabetes—2025. Diabetes Care 2025, 48, S1–S352.

  • 4.

    Lewis, P.; Perez, E.; Piktus, A.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 6–12 December 2020; Volume 33, pp. 9459–9474.

  • 5.

    Karpathy, A. LLMWiki: A Pattern for Building Personal Knowledge Bases Using LLMs. Available online: https://gist.github.com/karpathy (accessed on 28 April 2026).

  • 6.

    Raghu, A.; Komorowski, M.; Celi, L.A.; et al. Continuous State-Space Models for Optimal Sepsis Treatment: A Deep Reinforcement Learning Approach. In Proceedings of the Machine Learning for Healthcare Conference (MLHC), Boston, MA, USA, 18–19 August 2017; pp. 147–163.

  • 7.

    Komorowski, M.; Celi, L.A.; Badawi, O.; et al. The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care. Nat. Med. 2018, 24, 1716–1720.

  • 8.

    Peine, A.; Hallawa, A.; Bickenbach, J.; et al. Development and Validation of a Reinforcement Learning Algorithm to Dynamically Optimize Mechanical Ventilation in Critical Care. npj Digit. Med. 2021, 4, 32.

  • 9.

    Achiam, J.; Held, D.; Tamar, A.; et al. Constrained Policy Optimization. In Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, NSW, Australia, 6–11 August 2017; pp. 22–31.

  • 10.

    Fang, N.; Gong, W.; Gu, C. Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare. IEEE Trans. Artif. Intell. 2026, 7, 2037–2046.

  • 11.

    Battelino, T.; Danne, T.; Bergenstal, R.M.; et al. Clinical Targets for Continuous Glucose Monitoring Data Interpretation: Recommendations from the International Consensus on Time in Range. Diabetes Care 2019, 42, 1593–1603.

  • 12.

    Battelino, T.; Alexander, C.M.; Amiel, S.A.; et al. Continuous Glucose Monitoring and Metrics for Clinical Trials: An International Consensus Statement. Lancet Diabetes Endocrinol. 2023, 11, 42–57.

  • 13.

    Marling, C.; Bunescu, R. The OhioT1DM Dataset for Blood Glucose Level Prediction: Update 2020. In Proceedings of the CEUR Workshop Proceedings, Online, 29–30 August 2020; Volume 2675, pp. 71–74.

  • 14.

    Microsoft. MarkItDown: Convert Files to Markdown. Available online: https://github.com/microsoft/markitdown (accessed on 8 April 2026).

Share this article:
How to Cite
Zhang, B. LLM Wiki as a Living Constraint Oracle for Safe Offline RL: Evidence from Type 1 Diabetes Glucose Management. Transactions on Artificial Intelligence 2026, 2 (1), 178–189. https://doi.org/10.53941/tai.2026.100011.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.