2608004889
  • Open Access
  • Article

Contextual Bandit Routing for LoRA Experts

  • Miavaka Voninahitra Ambinintsoa Rakotovao Valisoa *,   
  • Hobihery Matio Robinson

Received: 01 Jul 2026 | Revised: 11 Aug 2026 | Accepted: 12 Aug 2026 | Published: 24 Aug 2026

Abstract

The deployment of multiple Low-Rank Adaptation (LoRA) experts for domain-specific tasks requires an efficient routing mechanism. Existing approaches rely on offline-trained routers or static heuristics, which fail to adapt to unseen domains or shifting query distributions. We propose EXP4-LoRA, a contextual bandit framework that treats LoRA expert selection as an online decision problem. Our method combines LinUCB for context-aware reward estimation with the EXP4 algorithm for exploration-exploitation trade-off, enabling per-query routing without any offline training of the router on domain-labeled data. We clarify that our approach requires binary correctness feedback (rather than domain labels), which is often available in practice via self-consistency checks, programmatic verification, or delayed user feedback. We further study proxy reward settings where even correctness labels are unavailable at routing time. On a mixed-domain MMLU benchmark with Gemma-2-2B and four LoRA experts (Mathematics, Code, Biology, Law), EXP4-LoRA achieves 80.1% accuracy after 2000 steps, outperforming EXP3 by 4.1 percentage points and reaching the 70% threshold in 420 steps versus 1800 for EXP3. We replicate the experiments on Phi-3-mini (3.8B) and confirm consistent gains (+4.2 points over EXP3). We also compare against LinTS (Thompson Sampling), which achieves 79.3% accuracy, demonstrating that EXP4-LoRA holds a modest edge. The total inference overhead, including context extraction, is approximately 100% of the base forward pass latency, reducible to 10–15% with layer caching optimizations. Our results demonstrate that lightweight contextual bandits are a viable and principled alternative to learned routers for multi-expert LLM serving.

References 

  • 1.

    Hu, E.J.; Shen, Y.; Wallis, P.; et al. LoRA: Low-rank adaptation of large language models. In Proceedings of the Tenth International Conference on Learning Representations (ICLR 2022), Virtual, 25–29 April 2022.

  • 2.

    Dou, S.; Zhou, E.; Liu, Y.; et al. LoRAMoE: Revolutionizing mixture of experts for maintaining world knowledge in language model alignment. arXiv 2023. arXiv:2312.09979.

  • 3.

    Xu, J.; Lai, J.; Huang, Y. MeteoRA: Multiple-tasks embedded LoRA for large language models. In Proceedings of the International Conference on Learning Representations 2025 (ICLR 2025), Singapore, 24–28 April 2025.

  • 4.

    Li, L.; Chu, W.; Langford, J.; et al. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, Raleigh, NC, USA, 26–30 April 2010; pp. 661–670.

  • 5.

    Auer, P.; Cesa-Bianchi, N.; Freund, Y.; et al. The nonstochastic multiarmed bandit problem. SIAM J. Comput. 2002, 32, 48–77.

  • 6.

    Zhang, Q.; Chen, M.; Bukharin, A.; et al. Ada-LoRA: Adaptive budget allocation for parameter-efficient fine-tuning. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023.

  • 7.

    Liu, S.Y.; Wang, C.Y.; Yin, H.; et al. DoRA: Weight-decomposed low-rank adaptation. arXiv 2024. arXiv:2402.09353.

  • 8.

    Agrawal, S.; Goyal, N. Thompson sampling for contextual bandits with linear payoffs. In Proceedings of the 30th International Conference on Machine Learning (ICML 2013), Atlanta, GA, USA, 16–21 June 2013.

  • 9.

    Chu, W.; Li, L.; Reyzin, L.; et al. Contextual bandits with linear payoff functions. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA, 11–13 April 2011.

  • 10.

    Krishnamurthy, A.; Harris, K.; Foster, D.J.; et al. Can Large Language Models Explore In-Context? In Proceedings of the Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Vancouver, BC, Canada, 10–15 December 2024.

  • 11.

    Dai, X.; Li, J.; Liu, X.; et al. Cost-effective online multi-LLM selection with versatile reward models. arXiv 2024. arXiv:2405.16587.

  • 12.

    Khattab, O.; Singhvi, A.; Maheshwari, P.; et al. DSPy: Compiling declarative language model calls into self-improving pipelines. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024.

  • 13.

    Hendrycks, D.; Burns, C.; Basart, S.; et al. Measuring massive multitask language understanding. In Proceedings of the 9th International Conference on Learning Representations (ICLR 2021), Virtual Event, Austria, 3–7 May 2021.

Share this article:
How to Cite
Valisoa, M. V. A. R.; Robinson, H. M. Contextual Bandit Routing for LoRA Experts. Transactions on Artificial Intelligence 2026, 2 (1), 190–203. https://doi.org/10.53941/tai.2026.100012.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.