2608004851
  • Open Access
  • Article

Provenance-Native Audit Infrastructure for LLM-Maintained Wiki Knowledge Systems

  • Bailing Zhang

Received: 27 Apr 2026 | Revised: 31 Jul 2026 | Accepted: 07 Aug 2026 | Published: 01 Sep 2026

Abstract

LLM-maintained wiki systems accumulate structured knowledge by having a language model incrementally build and revise a persistent wiki layer over raw evidence sources. When such systems support decision-making in regulated domains, each output must be traceable to grounded evidence through an auditable chain. This paper presents a provenance-native architecture in which a dedicated provenance layer is embedded in the wiki’s operational loop, recording evidence bindings, claim versions, rule compilations, and decision traces as side effects of normal system operations. The architecture is organized around four technical components: (1) a typed provenance graph that binds claims to evidence passages with confidence-weighted edges; (2) a dual-layer diff mechanism combining textual and semantic comparison to detect silent LLM edits; (3) three confidence propagation strategies over multi-hop evidence–claim–rule–decision chains, compared on their suitability for regulatory communication; and (4) a retraction propagation engine that identifies all downstream dependents when upstream evidence is invalidated. We evaluate the system on a 112-document corpus producing 3174 graph nodes and 4398 edges. In controlled retraction experiments, the engine achieves 100% downstream recall with zero false propagation in under 600 ms per event. A sensitivity analysis over the human-review threshold parameter γ reveals a sharp phase transition in the fraction of flagged decisions, providing a concrete basis for negotiating review policies with regulators. An automated gap analysis shows the architecture demonstrates alignment with five of eight requirements derived from the EU AI Act and FDA AI/ML guidance, and we report module-level evaluations of claim extraction, evidence binding, and silent-edit detection together with a measured baseline comparison on retraction propagation. Code and experimental protocols will be released publicly upon publication.

References 

  • 1.

    Karpathy, A. LLM Wiki: A Pattern for Building Personal Knowledge Bases Using LLMs. Idea Document, 2026. Available online: https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f (accessed on 1 May 2026).

  • 2.

    U.S. Food and Drug Administration. Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) Action Plan; Technical Report; FDA: Silver Spring, MD, USA, 2021.

  • 3.

    European Parliament and Council of the European Union. Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act). Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (accessedon 1 May 2026).

  • 4.

    Cheney, J.; Chiticariu, L.; Tan, W.C. Provenance in Databases: Why, How, and Where. Found. Trends Databases 2009, 1, 379–474.

  • 5.

    Herschel, M.; Diestelk¨amper, R.; Ben Lahmar, H. A Survey on Provenance: What for? What form? What from? VLDB J. 2017, 26, 881–906.

  • 6.

    Freire, J.; Koop, D.; Santos, E.; et al. Provenance for Computational Tasks: A Survey. Comput. Sci. Eng. 2008, 10, 11–21.

  • 7.

    Lebo, T.; Sahoo, S.; McGuinness, D.; et al. PROV-O: The PROV Ontology. W3C Recommendation, 2013. Available online: https://www.w3.org/TR/prov-o/ (accessed on 1 May 2026).

  • 8.

    Bizer, C.; Cyganiak, R. Quality-Driven Information Filtering Using the WIQA Policy Framework. J. Web Semant. 2009, 7, 1–10.

  • 9.

    Gao, T.; Yen, H.; Yu, J.; et al. Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 6465–6488.

  • 10.

    Meng, K.; Bau, D.; Andonian, A.; et al. Locating and Editing Factual Associations in GPT. In Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; pp. 17359–17372.

  • 11.

    Meng, K.; Sharma, A.S.; Andonian, A.; et al. Mass-Editing Memory in a Transformer. In Proceedings of the ICLR 2023 (International Conference on Learning Representations), Kigali, Rwanda, 1–5 May 2023.

  • 12.

    Thorne, J.; Vlachos, A.; Christodoulopoulos, C.; et al. FEVER: A Large-Scale Dataset for Fact Extraction and Verification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, New Orleans, LA, USA, 1–6 June 2018; pp. 809–819.

  • 13.

    Manhaeve, R.; Dumancic, S.; Kimmig, A.; et al. DeepProbLog: Neural Probabilistic Logic Programming. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montreal, QC, Canada, 3–8 December 2018; pp. 3749–3759.

  • 14.

    Packer, C.; Wooders, S.; Lin, K.; et al. MemGPT: Towards LLMs as Operating Systems. arXiv 2023, arXiv:2310.08560.

  • 15.

    Shinn, N.; Cassano, F.; Gopinath, A.; et al. Reflexion: Language Agents with Verbal Reinforcement Learning. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023; pp. 8634–8652.

Share this article:
How to Cite
Zhang, B. Provenance-Native Audit Infrastructure for LLM-Maintained Wiki Knowledge Systems. Artificial Intelligence and Emerging Technologies 2026, 3 (3), 10. https://doi.org/10.53941/aiet.2026.100010.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.