2607004641
  • Open Access
  • Article

NeSyWikiCompiler: A Neural-Symbolic Compiler for Verifiable AI Knowledge Engineering

  • Bailing Zhang

Received: 16 May 2026 | Revised: 30 Jun 2026 | Accepted: 14 Jul 2026 | Published: 27 Jul 2026

Abstract

Knowledge rules in AI engineering carry an awkward double duty. A clinician or a compliance officer has to be able to read and revise them, and a downstream system has to be able to execute and check them. Getting from the natural-language version of such a rule to a symbolic program is rarely the hard part; getting to one that is not merely runnable but logically sound is. We call this the compilation gap and attack it with NeSyWikiCompiler, a four-stage neural-symbolic pipeline. A language-model frontend reads each specification into NeSy-IR, an intermediate representation that holds onto exactly the details downstream code generation needs and that language models routinely drop: predicate types, the direction of numeric comparisons, the modality of each constraint, and a pointer back to the source text. From a single IR, deterministic compilers emit both a Prolog program and a Z3 program. The Z3 side is then checked with Clark’s completion, which surfaces a failure that, in our experiments, plain Prolog execution cannot see at all: rules whose violation condition can never be satisfied, so that the checker silently never fires. A repair loop that is CEGIS-informed but ultimately deterministic for structural errors handles the two kinds of failure differently: an LLM revision step targets semantic slips such as lost numeric thresholds and reversed arguments, while a deterministic reconstruction step—reached once the LLM rounds have failed—is the path taken for the structural contradictions, which the LLM step does not fix. In a cross-domain evaluation covering clinical decision rules, AI course knowledge, and legal compliance specifications—three representative domains rather than an exhaustive sample—all compiled programs achieve syntax validity across both backends. Formal verification exposes a class of structural contradictions that, in this pilot-scale benchmark, is concentrated in the legal specifications, where normative hedging constructs appear to induce rule-constraint conflicts. Deterministic repair recovers most structural failures while LLM-only repair consistently reproduces the same broken rule patterns. These findings characterise a previously unrecognised failure mode in LLM-to-logic compilation and demonstrate a practical engineering toolchain for producing verifiable knowledge-base programs from natural-language specifications. We present the domain-level rates as failure-mode discovery on a small corpus rather than as estimates that generalise without further study.

References 

  • 1.

    Hripcsak, G.; Clayton, P.D.; Pryor, T.A.; et al. The Arden Syntax for Medical Logic Modules. In Proceedings of the 18th Annual Symposium on Computer Applications in Medical Care (SCAMC), Washington, DC, USA, 4–7 November 1990; pp. 200–204.

  • 2.

    Ohno-Machado, L.; Gennari, J.H.; Murphy, S.N.; et al. The GuideLine Interchange Format: A Model for Representing Guidelines. J. Am. Med. Inform. Assoc. 1998, 5, 357–372.

  • 3.

    Yang, F.; Zhao, P.; Wang, Z.; et al. Empower large language model to perform better on industrial domain-specific question answering. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, Singapore, 6–10 December 2023.

  • 4.

    Chen, M.; Tworek, J.; Jun, H.; et al. Evaluating Large Language Models Trained on Code. arXiv 2021. arXiv:2107.03374.

  • 5.

    Karpathy, A. LLMs as Disciplined Wiki Curators. GitHub Gist, April 2026. Available online: https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f (accessed on 10 May 2026).

  • 6.

    Clark, K.L. Negation as Failure. In Logic and Data Bases; Gallaire, H., Minker, J., Eds.; Springer: Boston, MA, USA , 1978; pp. 293–322.

  • 7.

    Zelle, J.M.; Mooney, R.J. Learning to Parse Database Queries Using Inductive Logic Programming. In Proceedings of the 13th National Conference on Artificial Intelligence (AAAI), Portland, OR, USA, 4–8 August 1996; Vol. 2, pp. 1050–1055.

  • 8.

    Zettlemoyer, L.S.; Collins, M. Learning to Map Sentences to Logical Form: Structured Classification with Probabilistic Categorial Grammars. In Proceedings of the 21st Conference on Uncertainty in Artificial Intelligence (UAI), Edinburgh, UK, 26–29 July 2005; pp. 658–666.

  • 9.

    Yu, T.; Zhang, R.; Yang, K.; et al. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), Brussels, Belgium, 31 October–4 November 2018; pp. 3911–3921.

  • 10.

    Wu, M.; Szegedy, C.; Urban, J. Autoformalization with Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 35 (NeurIPS 2022), New Orleans, LA, USA, 28 November–9 December 2022; Vol. 35, pp. 32353–32368.

  • 11.

    Solar-Lezama, A.; Tancau, L.; Bod´ık, R.; et al. Combinatorial Sketching for Finite Programs. In Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), San Jose, CA, USA, 21–25 October 2006; pp. 404–415.

  • 12.

    Gulwani, S. Automating String Processing in Spreadsheets Using Input-Output Examples. ACM SIGPLAN Not. 2011, 46, 317–330. https://doi.org/10.1145/1925844.1926423.

  • 13.

    Alur, R.; Bodik, R.; Juniwal, G.; et al. Syntax-Guided Synthesis. In Proceedings of the IEEE Formal Methods in Computer-Aided Design (FMCAD), Portland, OR, USA, 20–23 October 2013.

  • 14.

    Manhaeve, R.; Dumancic, S.; Kimmig, A.; et al. DeepProbLog: Neural Probabilistic Logic Programming. In Proceedings of the Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Montreal, QC, Canada, 2–8 December 2018; Vol. 31, pp. 3749–3759.

  • 15.

    Yang, Z.; Ishay, A.; Lee, J. NeurASP: Embracing Neural Networks into Answer Set Programming. In Proceedings of the 29th International Joint Conference on Artificial Intelligence (IJCAI), Yokohama, Japan, 11–17 July 2020 ; pp. 1755–1762.

  • 16.

    Riegel, R.; Gray, A.; Luus, F.; et al. Logical Neural Networks. arXiv 2020, arXiv:2006.13155.

  • 17.

    Shortliffe, E.H.; Blois, M.S. The Computer Meets Medicine and Biology: Emergence of a Discipline. In Biomedical Informatics: Computer Applications in Health Care and Biomedicine, 3rd ed.; Shortliffe, E.H., Cimino, J.J., Eds.; Springer: New York, NY, USA, 2006 ; pp. 3–45.

  • 18.

    Palmirani, M.; Governatori, G.; Rotolo, A.; et al. LegalRuleML: XML-Based Rules and Norms. In Proceedings of the Rule-Based Modeling and Computing on the Semantic Web, 5th International Symposium, RuleML 2011-America , Ft. Lauderdale, FL, USA, 3–5 November 2011; Lecture Notes in Computer Science; Vol. 7018, pp. 298–312.

  • 19.

    Bench-Capon, T.; Araszkiewicz, M.; Ashley, K.; et al. A History of AI and Law in 50 Papers: 25 Years of the International Conference on AI and Law. Artif. Intell. Law 2012, 20, 215–319.

  • 20.

    de Moura, L.; Bjørner, N. Z3: An Efficient SMT Solver. In Proceedings of the Tools and Algorithms for the Construction and Analysis of Systems (TACAS), Budapest, Hungary, 29 March–6 April 2008 ; Vol. 4963; pp. 337–340.

  • 21.

    Zhang, B. Beyond RAG: LLM Wikis as Living Semantic Memory: Patterns, Empirical Findings, and Open Problems. Zenodo 2026. Available online: https://zenodo.org/records/20078453 (accessed on 16 May 2026).

  • 22.

    Wielemaker, J.; Schrijvers, T.; Triska, M.; et al. SWI-Prolog. Theory Pract. Log. Program. 2012, 12, 67–96.

Share this article:
How to Cite
Zhang, B. NeSyWikiCompiler: A Neural-Symbolic Compiler for Verifiable AI Knowledge Engineering. AI Engineering 2026, 2 (2), 11. https://doi.org/10.53941/aieng.2026.100011.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.