2606004211
  • Open Access
  • Article

Towards a Unified Framework for Large–Small Model Collaboration in Cloud–Edge Systems

  • Jing Bi 1,   
  • Yaxi Yang 1,   
  • Haitao Yuan 2,*,   
  • Ziqi Wang 1,   
  • Jiayi Li 1,   
  • Jiahui Zhai 1,   
  • Jia Zhang 3

Received: 09 Mar 2026 | Revised: 20 Jun 2026 | Accepted: 10 Jul 2026 | Published: 17 Jul 2026

Abstract

Foundation models have improved the reasoning and generation ability of artificial intelligence systems. However, they are difficult to deploy in edge environments with limited computation, memory, and data access. Small models are easier to run on edge devices. They support fast and low-latency inference, but they often lack global semantic reasoning and cross-domain generalization. This gap between model ability and deployment cost motivates large–small model collaboration in cloud–edge systems. This survey provides a systematic review and a knowledge-floworiented taxonomy of such collaboration. It focuses on how cloud-side large models and edge-side small models share, update, and coordinate knowledge. We review knowledge distillation, split inference, federated and continual adaptation, and elastic offloading. We also cover lightweight deployment, modular expert design, privacyaware coordination, and agent-driven orchestration. Unlike surveys on edge intelligence, federated learning, model compression, TinyML, or cloud–edge resource scheduling, this survey centers on model collaboration. We treat large–small collaboration as a knowledge-centered problem linked to real deployment constraints. We further discuss trade-offs in accuracy, latency, bandwidth, privacy, energy efficiency, adaptability, and lifecycle management. Finally, we identify open challenges for trustworthy, sustainable, and self-evolving cloud–edge collaborative intelligence.

References 

  • 1.

    Yang, W.; Liu, M.; Wang, Z.; et al. Foundation Models Meet Visualizations: Challenges and Opportunities. Comput. Vis. Media 2024, 10, 399–424.

  • 2.

    Wang, C.; Peng, Y.; Zhang, D.; et al. A Lightweight Cloud–Edge Collaborative Intelligence Inference Framework with Runtime Dynamic Optimization for Resource-Constrained Consumer Electronics. IEEE Trans. Consum. Electron. 2025, 71, 6041–6054.

  • 3.

    An, S.; Jeong, H.; Kim, S.; et al. CINELL: An Energy-Efficient Compute-In/Near-Memory eDRAM Processor for Sparse Transformer-Based Large Language Models. IEEE Trans. Very Large Scale Integr. Syst. 2026, 34, 242–253.

  • 4.

    Sun, L.; Tian, J.; Muhammad, G. FedKC: Personalized Federated Learning with Robustness Against Model Poisoning Attacks in the Metaverse for Consumer Health. IEEE Trans. Consum. Electron. 2024, 70, 5644–5653.

  • 5.

    Adil, M.; Khan, M.K.; Jadoon, M.M.; et al. An AI-Enabled Hybrid Lightweight Authentication Scheme for Intelligent IoMT Based Cyber-Physical Systems. IEEE Trans. Netw. Sci. Eng. 2023, 10, 2719–2730.

  • 6.

    Gu, H.; Zhao, L.; Han, Z.; et al. AI-Enhanced Cloud-Edge-Terminal Collaborative Network: Survey, Applications, and Future Directions. IEEE Commun. Surv. Tutorials 2024, 26, 1322–1385.

  • 7.

    Qu, G.; Chen, Q.; Wei, W.; et al. Mobile Edge Intelligence for Large Language Models: A Contemporary Survey. IEEE Commun. Surv. Tutorials 2025, 27, 3820–3860.

  • 8.

    Wang, Y.; Yang, C.; Lan, S.; et al. End-Edge-Cloud Collaborative Computing for Deep Learning: A Comprehensive Survey. IEEE Commun. Surv. Tutorials 2024, 26, 2647–2683.

  • 9.

    Hoffpauir, K.; Simmons, J.; Schmidt, N.; et al. A Survey on Edge Intelligence and Lightweight Machine Learning Support for Future Applications and Services. J. Data Inf. Qual. 2023, 15, 20.

  • 10.

    Kairouz, P.; McMahan, H.B. Advances and Open Problems in Federated Learning. Found. Trends Mach. Learn. 2021, 14, 1–210.

  • 11.

    Cheng, Y.; Wang, D.; Zhou, P.; et al. A Survey of Model Compression and Acceleration for Deep Neural Networks. arXiv 2017, arXiv:1710.09282.

  • 12.

    Alajlan, N.N.; Ibrahim, D.M. TinyML: Enabling of Inference Deep Learning Models on Ultra-Low-Power IoT Edge Devices for AI Applications. Micromachines 2022, 13, 851.

  • 13.

    Xu, M.; Cai, D.; Yin,W.; et al. Resource-Efficient Algorithms and Systems of Foundation Models: A Survey. ACM Comput. Surv. 2025, 57, 110.

  • 14.

    Song, S.; Li, X.; Li, S.; et al. How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model. IEEE Trans. Knowl. Data Eng. 2025, 37, 5311–5329.

  • 15.

    Wang, X.; Tang, Z.; Guo, J.; et al. Empowering Edge Intelligence: A Comprehensive Survey on On-Device AI Models. ACM Comput. Surv. 2025, 57, 1–39.

  • 16.

    Zheng, Y.; Chen, Y.; Qian, B.; et al. A Review on Edge Large Language Models: Design, Execution, and Applications. ACM Comput. Surv. 2025, 57, 209.

  • 17.

    Zeng, T.; Zhang, X.; Duan, J.; et al. An Offline-Transfer-Online Framework for Cloud–Edge Collaborative Distributed Reinforcement Learning. IEEE Trans. Parallel Distrib. Syst. 2024, 35, 720–731.

  • 18.

    Yan, X.; Li, Z.; Ni, Z.; et al. Traffic-Aware Cloud–Edge Collaborative Offloading for Vehicular Tasks at Complex Intersections. IEEE Trans. Veh. Technol. 2026, 75, 321–334.

  • 19.

    Li, X.; Li, H.; Sun, C.; et al. Edge-Enhanced Intelligence: A Comprehensive Survey of Large Language Models and Edge–Cloud Computing Synergy. IEEE Commun. Surv. Tutorials 2026, 28, 1248–1284.

  • 20.

    Cheng, L.; Zhang, S.; Zhang, H.; et al. Large-Small Model Collaboration in Mobile Edge Networks with Heterogeneous Computational Resources. IEEE J. Sel. Areas Commun. 2026, 44, 2733–2749.

  • 21.

    Rajpal, T.S.; Naithani, A. NeuroCrypt: A Neuro Symbolic AI Ecosystem for Advanced Cryptographic Data Security and Transmission. IEEE Trans. Artif. Intell. 2026, 7, 512–521.

  • 22.

    Li, H.; Li, X.; Fan, Q.; et al. Adaptive Model Partitioning and Pruning for Collaborative DNN Inference in Mobile Edge–Cloud Computing Networks. IEEE Trans. Mob. Comput. 2026, 25, 3744–3759.

  • 23.

    Lim, J.-A.; Lee, J.; Kwak, J.; et al. Cutting-Edge Inference: Dynamic DNN Model Partitioning and Resource Scaling for Mobile AI. IEEE Trans. Serv. Comput. 2024, 17, 3300–3316.

  • 24.

    Bi, J.; Yuan, H.; Duanmu, S.; et al. Energy-Optimized Partial Computation Offloading in Mobile-Edge Computing with Genetic Simulated-Annealing-Based Particle Swarm Optimization. IEEE Internet Things J. 2021, 8, 3774–3785.

  • 25.

    Bi, J.; Wang, Z.; Yuan, H.; et al. Cost-Minimized Partial Computation Offloading in Cloud-Assisted Mobile Edge Computing Systems. In Proceedings of the 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Honolulu, HI, USA, 1–4 October 2023; pp. 5052–5057.

  • 26.

    Zhao, H.; Wang, Z.; Cheng, G.; et al. Online Workload Scheduling for Social Welfare Maximization in the Computing Continuum. IEEE Trans. Serv. Comput. 2025, 18, 2267–2280.

  • 27.

    Wang, Z.; Zhou, Y.; Shi, Y.; et al. Federated Fine-Tuning for Pre-Trained Foundation Models Over Wireless Networks. IEEE Trans. Wirel. Commun. 2025, 24, 3450–3464.

  • 28.

    Ramadan, M.N.A.; Ali, M.A.H.; Khoo, S.Y.; et al. Federated Learning and TinyML on IoT Edge Devices: Challenges, Advances, and Future Directions. ICT Express 2025, 11, 754–768.

  • 29.

    Yang, C.; Zhu, Y.; Lu, W.; et al. Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application. ACM Trans. Intell. Syst. Technol. 2025, 16, 143.

  • 30.

    Shao, J.; Wu, F.; Zhang, X.; et al. Selective Knowledge Sharing for Privacy-Preserving Federated Distillation without a Good Teacher. Nat. Commun. 2024, 15, 349.

  • 31.

    Luo, H.; Chen, T.; Li, X.; et al. KeepEdge: A Knowledge Distillation Empowered Edge Intelligence Framework for Visual Assisted Positioning in UAV Delivery. IEEE Trans. Mob. Comput. 2023, 22, 4729–4741.

  • 32.

    Chen, J.; Xu, W.; Fan, Y.; et al. Fast Multimodal Edge Inference via Selective Feature Distillation. IEEE Trans. Mob. Comput. 2025, 24, 11337–11350.

  • 33.

    Zhang, L.; Ma, K. Structured Knowledge Distillation for Accurate and Efficient Object Detection. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 15706–15724.

  • 34.

    Qiao, D.; Guo, S.; Zhao, J.; et al. ASMAFL: Adaptive Staleness-Aware Momentum Asynchronous Federated Learning in Edge Computing. IEEE Trans. Mob. Comput. 2025, 24, 3390–3406.

  • 35.

    Li, Y.; Liu, Z.; Kou, Z.; et al. Real-Time Adaptive Partition and Resource Allocation for Multi-User End-Cloud Inference Collaboration in Mobile Environment. IEEE Trans. Mob. Comput. 2024, 23, 13076–13094.

  • 36.

    Guo, J.; Fu, Y.; Zhai, Z.; et al. MASA: Multimodal Federated Learning Through Modality-Aware and Secure Aggregation. IEEE Trans. Mob. Comput. 2025, 24, 7328–7344.

  • 37.

    Yang, S.; Yang, J.; Zhou, M.; et al. Learning From Human Educational Wisdom: A Student-Centered Knowledge Distillation Method. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 4188–4205.

  • 38.

    Zhang, S.; Luo, Y.; Lyu, Z.; et al. ShiftKD: Benchmarking Knowledge Distillation under Distribution Shift. Neural Netw. 2025, 192, 107838.

  • 39.

    Li, H.; Li, X.; Fan, Q.; et al. Distributed DNN Inference with Fine-Grained Model Partitioning in Mobile Edge Computing Networks. IEEE Trans. Mob. Comput. 2024, 23, 9060–9074.

  • 40.

    Liang, H.; Sang, Q.; Hu, C.; et al. DNN Surgery: Accelerating DNN Inference on the Edge through Layer Partitioning. IEEE Trans. Cloud Comput. 2023, 11, 3111–3125.

  • 41.

    Wang, Z.; Li, F.; Zhang, Y.; et al. Low-Rate Feature Compression for Collaborative Intelligence: Reducing Redundancy in Spatial and Statistical Levels. IEEE Trans. Multimed. 2024, 26, 2756–2771.

  • 42.

    Yuan, Z.; Rawlekar, S.; Garg, S.; et al. Split Computing with Scalable Feature Compression for Visual Analytics on the Edge. IEEE Trans. Multimed. 2024, 26, 10121–10133.

  • 43.

    Zheng, G.; Wen, M.; Ning, Z.; et al. Computation-Aware Offloading for DNN Inference Tasks in Semantic Communication Assisted MEC Systems. IEEE Trans. Wirel. Commun. 2025, 24, 2693–2706.

  • 44.

    Wang, W.; Zhang, Y.; Huang, R.; et al. Efficient Resource Management and Expansion Scheme for Collaborative Edge-Cloud Computing. IEEE Trans. Mob. Comput. 2024, 23, 2731–2747.

  • 45.

    Hao, Z.; Xu, G.; Luo, Y.; et al. Multi-Agent Collaborative Inference via DNN Decoupling: Intermediate Feature Compression and Edge Learning. IEEE Trans. Mob. Comput. 2023, 22, 6041–6055.

  • 46.

    Sa, C.-C.; Cheng, L.-C.; Chung, H.-H.; et al. Ensuring Bidirectional Privacy on Wireless Split Inference Systems. IEEE Wirel. Commun. 2024, 31, 134–141.

  • 47.

    Zhao, S.; Liu, T.; Jin, H.; et al. SPViT: Accelerate Vision Transformer Inference on Mobile Devices via Adaptive Splitting and Offloading. IEEE Trans. Mob. Comput. 2025, 24, 9303–9318.

  • 48.

    Gong, R.; Liu, X.; Li, Y.; et al. Pushing the Limit of Post-Training Quantization. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 5556–5570.

  • 49.

    Zhu, X.; Li, J.; Liu, Y.; et al. A Survey on Model Compression for Large Language Models. Trans. Assoc. Comput. Linguist. 2024, 12, 1556–1577.

  • 50.

    Cheng, H.; Zhang, M.; Shi, J.Q. A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 10558–10578.

  • 51.

    Liu, H.-I.; Galindo, M.; Xie, H.; et al. Lightweight Deep Learning for Resource-Constrained Environments: A Survey. ACM Comput. Surv. 2024, 56, 267.

  • 52.

    Jin, H.; Bai, D.; Yao, D.; et al. Personalized Edge Intelligence via Federated Self-Knowledge Distillation. IEEE Trans. Parallel Distrib. Syst. 2023, 34, 567–580.

  • 53.

    Wu, Z.; Sun, S.; Wang, Y.; et al. FedICT: Federated Multi-Task Distillation for Multi-Access Edge Computing. IEEE Trans. Parallel Distrib. Syst. 2024, 35, 1107–1121.

  • 54.

    Yao, D.; Pan, W.; Dai, Y.; et al. FedGKD: Toward Heterogeneous Federated Learning via Global Knowledge Distillation. IEEE Trans. Comput. 2024, 73, 3–17.

  • 55.

    Yang, X.; Yu, H.; Gao, X.; et al. Federated Continual Learning via Knowledge Fusion: A Survey. IEEE Trans. Knowl. Data Eng. 2024, 36, 3832–3850.

  • 56.

    Zhang, H.; Tao, M.; Shi, Y.; et al. Federated Multi-Task Learning with Non-Stationary and Heterogeneous Data in Wireless Networks. IEEE Trans. Wirel. Commun. 2024, 23, 2653–2667.

  • 57.

    Tan, A.Z.; Yu, H.; Cui, L.; et al. Towards Personalized Federated Learning. IEEE Trans. Neural Netw. Learn. Syst. 2023, 34, 9587–9603.

  • 58.

    Liu, Y.; Liu, F.; Guo, B.; et al. Sparse-FCL: Sparse Federated Continual Learning for Evolving Mobile Edge Computing Environments. IEEE Trans. Serv. Comput. 2025, 18, 2403–2416.

  • 59.

    Zuo, X.; Luopan, Y.; Han, R.; et al. FedViT: Federated Continual Learning of Vision Transformer at Edge. Future Gener. Comput. Syst. 2024, 154, 1–15.

  • 60.

    Yu, D.; Shen, L.; Hao, H.; et al. MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services. IEEE Trans. Serv. Comput. 2024, 17, 2626–2639.

  • 61.

    Feng, X.; Feng, X.; Du, X.; et al. Adapter-Based Selective Knowledge Distillation for Federated Multi-Domain Meeting Summarization. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 3694–3708.

  • 62.

    Cai, W.; Jiang, J.; Wang, F.; et al. A Survey on Mixture of Experts in Large Language Models. IEEE Trans. Knowl. Data Eng. 2025, 37, 3896–3915.

  • 63.

    Wang, L.; Chen, S.; Jiang, L.; et al. Parameter-Efficient Fine-Tuning in Large Language Models: A Survey of Methodologies. Artif. Intell. Rev. 2025, 58, 227.

  • 64.

    Chen, Z.; Zhu, B.; Wang, J.; et al. Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities. ACM Comput. Surv. 2026, 58, 317. https://doi.org/10.1145/3809166.

  • 65.

    Piccialli, F.; Chiaro, D.; Qi, P.; et al. Federated and Edge Learning for Large Language Models. Inf. Fusion 2025, 117, 102840.

  • 66.

    Husom, E.J.; Goknil, A.; Astekin, M.; et al. Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency. ACM Trans. Internet Things 2025, 6, 28.

  • 67.

    Li, X.; Zhang, Y.; Wang, J.; et al. Cloud–Edge-End Collaborative Intelligent Service Computation Offloading: A Digital Twin-Driven Edge Coalition Approach for Industrial IoT. IEEE Trans. Netw. Serv. Manag. 2024, 21, 6318–6330.

  • 68.

    Du, J.; Jiang, C.; Benslimane, A.; et al. SDN-Based Resource Allocation in Edge and Cloud Computing Systems: An Evolutionary Stackelberg Differential Game Approach. IEEE/ACM Trans. Netw. 2022, 30, 1613–1628.

  • 69.

    Hossein Shokouhi, M.; Hadi, M.; Pakravan, M.R. Mobility-Aware Computation Offloading for Hierarchical Mobile Edge Computing. IEEE Trans. Netw. Serv. Manag. 2024, 21, 3372–3384.

  • 70.

    Cohen, I.; Chiasserini, C.F.; Giaccone, P.; et al. Dynamic Service Provisioning in the Edge-Cloud Continuum with Bounded Resources. IEEE/ACM Trans. Netw. 2023, 31, 3096–3111.

  • 71.

    Zeng, Z.; Zhao, Y.; Chen, X.; et al. Accelerating Federated Codistillation via Adaptive Computation Amount Allocation. IEEE Trans. Mob. Comput. 2025, 24, 5584–5597.

  • 72.

    Lin, Z.; Bi, S.; Lin, X.; et al. Efficient Parallel Split Learning over Resource-Constrained Wireless Edge Networks. IEEE Trans. Mob. Comput. 2024, 23, 9224–9239.

  • 73.

    Yuan, H.; Bi, J.; Wang, Z.; et al. Partial and Cost-Minimized Computation Offloading in Hybrid Edge and Cloud Systems. Expert Syst. Appl. 2024, 250, 123896.

  • 74.

    Purice, D.; Barchi, F.; R¨oder, T.; et al. Edge AI Lifecycle Management. In Advancing Edge Artificial Intelligence–System Contexts; River Publishers: Gistrup, Denmark, 2024; pp. 43–64.

  • 75.

    Hinder, F.; Artelt, A.; Hammer, B. One or Two Things We Know about Concept Drift—A Survey on Monitoring Evolving Environments. Front. Artif. Intell. 2024, 7, 1330257.

  • 76.

    Zhou, Y.; You, C.; Huang, K. Communication Efficient Cooperative Edge AI via Event-Triggered Computation Offloading. IEEE Trans. Commun. 2026, 74, 3190–3205.

  • 77.

    Bhuyan, B.P.; Mohapatra, S.S.; Panda, S.K. Neuro-Symbolic Artificial Intelligence: A Survey. Neural Comput. Appl. 2024, 36, 12809–12844.

  • 78.

    Pan, Y.; Su, Z.; Wang, Y.; et al. Cloud-Edge Collaborative Large Model Services: Challenges and Solutions. IEEE Netw. 2025, 39, 182–191.

  • 79.

    Renggli, C.; Rimanic, L.; G¨urel, N.M.; et al. A Data Quality-Driven View of MLOps. IEEE Data Eng. Bull. 2021, 44, 11–23.

  • 80.

    Wang, Z.; Li, X.; Wang, J.; et al. Energy-Minimized Partial Computation Offloading in Cloud-Assisted Vehicular Edge Computing Systems. In Proceedings of the 2023 IEEE International Conference on Networking, Sensing and Control (ICNSC), Marseille, France, 25–27 October 2023; pp. 1–6.

  • 81.

    Younis, A.; Maheshwari, S.; Pompili, D. Energy-Latency Computation Offloading and Approximate Computing in Mobile-Edge Computing Networks. IEEE Trans. Netw. Serv. Manag. 2024, 21, 3401–3415.

  • 82.

    Bi, J.; Wang, Z.; Yuan, H.; et al. Cost-Minimized Computation Offloading and User Association in Hybrid Cloud and Edge Computing. IEEE Internet Things J. 2024, 11, 16672–16683.

  • 83.

    Fan, W.; Xiao, F.; Pan, Y.; et al. Latency-Aware Joint Task Offloading and Energy Control for Cooperative Mobile Edge Computing. IEEE Trans. Serv. Comput. 2025, 18, 1515–1528.

  • 84.

    Yang, N.; Wen, J.; Zhang, M.; et al. Multi-objective Deep Reinforcement Learning for Mobile Edge Computing. In Proceedings of the 2023 21st International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt), Singapore, 24–27 August 2023; pp. 1–8.

  • 85.

    Shao, J.; Mao, Y.; Zhang, J. Learning Task-Oriented Communication for Edge Inference: An Information Bottleneck Approach. IEEE J. Sel. Areas Commun. 2022, 40, 197–211.

  • 86.

    Wang, Y.; Liu, H.; Cao, H. VOSA: Verifiable and Oblivious Secure Aggregation for Privacy-Preserving Federated Learning. IEEE Trans. Dependable Secur. Comput. 2023, 20, 3601–3616.

  • 87.

    Zhu, Y.; Gong, J.; Zhang, K.; et al. Malicious-Resistant Non-Interactive Verifiable Aggregation for Federated Learning. IEEE Trans. Dependable Secur. Comput. 2024, 21, 5600–5616.

  • 88.

    Zhou, Y.; Wang, R.; Liu, J.; et al. Exploring the Practicality of Differentially Private Federated Learning: A Local Iteration Tuning Approach. IEEE Trans. Dependable Secur. Comput. 2024, 21, 3280–3294.

  • 89.

    Ma, J.; Zhou, Y.; Cui, L.; et al. An Optimized Sparse Response Mechanism for Differentially Private Federated Learning. IEEE Trans. Dependable Secur. Comput. 2024, 21, 2285–2295.

  • 90.

    Liu, J.; Li, X.; Liu, X.; et al. Privacy-Preserving and Verifiable Outsourcing Linear Inference Computing Framework (PPVLC). IEEE Trans. Serv. Comput. 2023, 16, 4591–4604.

  • 91.

    Kim, D.; Guyot, C. Optimized Privacy-Preserving CNN Inference with Fully Homomorphic Encryption. IEEE Trans. Inf. Forensics Secur. 2023, 18, 2175–2187.

  • 92.

    Yamansavascilar, B.; Baktir, A.C.; Sonmez, C.; et al. DeepEdge: A Deep Reinforcement Learning Based Task Orchestrator for Edge Computing. IEEE Trans. Netw. Sci. Eng. 2023, 10, 538–552.

  • 93.

    Li, K.; Wang, X.; He, Q.; et al. Computation Offloading in Resource-Constrained Multi-Access Edge Computing. IEEE Trans. Mob. Comput. 2024, 23, 10665–10677.

  • 94.

    Hsieh, L.-T.; Liu, H.; Guo, Y.; et al. Deep Reinforcement Learning-Based Task Assignment for Cooperative Mobile Edge Computing. IEEE Trans. Mob. Comput. 2024, 23, 3156–3171.

  • 95.

    Sun, C.; Li, X.; Wang, C.; et al. Hierarchical Deep Reinforcement Learning for Joint Service Caching and Computation Offloading in Mobile Edge-Cloud Computing. IEEE Trans. Serv. Comput. 2024, 17, 1548–1564.

  • 96.

    Luo, Z.; Dai, X. Reinforcement Learning-Based Computation Offloading in Edge Computing: Principles, Methods, Challenges. Alex. Eng. J. 2024, 108, 89–107.

  • 97.

    Lv, Z.; Xiao, L.; Du, Y.; et al. Efficient Communications in Multi-Agent Reinforcement Learning for Mobile Applications. IEEE Trans. Wirel. Commun. 2024, 23, 12440–12454.

  • 98.

    Nguyen, D.C.; Ding, M.; Pathirana, P.N.; et al. Cooperative Task Offloading and Block Mining in Blockchain-Based Edge Computing with Multi-Agent Deep Reinforcement Learning. IEEE Trans. Mob. Comput. 2023, 22, 2021–2037.

  • 99.

    Huang, T.; Fang, Z.; Tang, Q.; et al. Dual-Timescales Optimization of Task Scheduling and Resource Slicing in Satellite-Terrestrial Edge Computing Networks. IEEE Trans. Mob. Comput. 2024, 23, 14111–14126.

  • 100.

    Liu, P.; An, K.; Lei, J.; et al. Computation Rate Maximization for SCMA-Aided Edge Computing in IoT Networks: A Multi-Agent Reinforcement Learning Approach. IEEE Trans. Wirel. Commun. 2024, 23, 10414–10429.

  • 101.

    Gao, Z.; Yang, L.; Dai, Y. Large-Scale Computation Offloading Using a Multi-Agent Reinforcement Learning in Heterogeneous Multi-Access Edge Computing. IEEE Trans. Mob. Comput. 2023, 22, 3425–3443.

  • 102.

    Li, X.; Huangfu, W.; Xu, X.; et al. Secure Offloading with Adversarial Multi-Agent Reinforcement Learning against Intelligent Eavesdroppers in UAV-Enabled Mobile Edge Computing. IEEE Trans. Mob. Comput. 2024, 23, 13914–13928.

  • 103.

    Lin, Y.; Zhang, Y.; Luo, C.; et al. Dynamic Expert Pruning and Routing for Efficient On-Device Mixture-of-Experts Inference. IEEE Internet Things J. 2023, 10, 15963–15975.

  • 104.

    Kong, R.; Li, Y.; Wang, W.; et al. Serving MoE Models on Resource-Constrained Edge Devices via Dynamic Expert Swapping. IEEE Trans. Comput. 2025, 74, 2799–2811.

  • 105.

    Guo, B.; Zhou, C.; He, S.; et al. Mixture-of-Experts as Continual Knowledge Adapter for Mobile Vision Understanding. IEEE Trans. Mob. Comput. 2025, 24, 9868–9882.

  • 106.

    Zhang, Z.; Zhao, D.; Liu, R.; et al. ACL: Adaptive Edge-Cloud Collaborative Learning for Heterogeneous Devices with Unlabeled Local Data. IEEE Trans. Mob. Comput. 2025, 24, 8152–8166.

  • 107.

    Hao, Y.; Yang, S.; Li, F.; et al. Learning Adaptive Multi-Timescale Scheduling for Mobile Edge Computing. IEEE Trans. Mob. Comput. 2025, 24, 7297–7311.

  • 108.

    Sharma, N.; Ghosh, A.; Misra, R.; et al. Deep Meta Q-Learning Based Multi-Task Offloading in Edge–Cloud Systems. IEEE Trans. Mob. Comput. 2024, 23, 2583–2598.

  • 109.

    Wang, Z.; Wong, V.W.S. Bayesian Meta-Learning for Adaptive Traffic Prediction in Wireless Networks. IEEE Trans. Mob. Comput. 2024, 23, 6620–6633.

  • 110.

    Dai, P.; Huang, Y.; Hu, K.; et al. Meta Reinforcement Learning for Multi-Task Offloading in Vehicular Edge Computing. IEEE Trans. Mob. Comput. 2024, 23, 2123–2138.

  • 111.

    Liu, Z.; Huang, L.; Gao, Z.; et al. GA-DRL: Graph Neural Network-Augmented Deep Reinforcement Learning for DAG Task Scheduling Over Dynamic Vehicular Clouds. IEEE Trans. Netw. Serv. Manag. 2024, 21, 4226–4242.

  • 112.

    Li, Y.; Xu, W.; Qi, Y.; et al. SR-FDIL: Synergistic Replay for Federated Domain-Incremental Learning. IEEE Trans. Parallel Distrib. Syst. 2024, 35, 1879–1890.

  • 113.

    Yu, T.; Fu, K.; Wang, S.; et al. Prompting Video-Language Foundation Models with Domain-Specific Fine-Grained Heuristics for Video Question Answering. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 1615–1630.

  • 114.

    Zhou, X.; Liu, M.; Yurtsever, E.; et al. Vision Language Models in Autonomous Driving: A Survey and Outlook. IEEE Trans. Intell. Veh. 2024, 9, 1–20.

  • 115.

    Zhao, S.; Liu, S.; Jiang, Y.; et al. Industrial Foundation Models (IFMs) for Intelligent Manufacturing: A Systematic Review. J. Manuf. Syst. 2025, 82, 420–448.

  • 116.

    AlSaad, R.; Abd-alrazaq, A.; Boughorbel, S.; et al. Multimodal Large Language Models in Health Care: Applications, Challenges, and Future Outlook. J. Med. Internet Res. 2024, 26, e59505.

  • 117.

    Liu, G.-P. Digital-Twin Predictive Control of Nonlinear Systems with Time Delays, Unknown Dynamics, and Communication Delays. IEEE Trans. Cybern. 2024, 54, 7198–7210.

  • 118.

    Wang, C.; Deng, Y.; Ning, Z.; et al. Building a Lightweight Trusted Execution Environment for Arm GPUs. IEEE Trans. Dependable Secur. Comput. 2024, 21, 3801–3816.

  • 119.

    Chaudhari, S.; Aggarwal, P.; Murahari, V.; et al. RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs. ACM Comput. Surv. 2025, 58, 53.

  • 120.

    Cheng, Y.; Zhang, W.; Zhang, Z.; et al. Toward Federated Large Language Models: Motivations, Methods, and Future Directions. IEEE Commun. Surv. Tutorials 2025, 27, 2733–2764.

  • 121.

    Gonzalez Barman, K.; Lohse, S.; de Regt, H.W. Reinforcement Learning from Human Feedback in LLMs: Whose Culture, Whose Values, Whose Perspectives? Philos. Technol. 2025, 38, 35.

  • 122.

    Cheng, R.; Chen, D.; Song, H.; et al. Formal Modeling and Verification Methods for the System Requirement Specifications of Train Control Systems: A Survey. IEEE Trans. Intell. Transp. Syst. 2025, 26, 1419–1440.

  • 123.

    Mao, Y.; Yu, X.; Huang, K.; et al. Green Edge AI: A Contemporary Survey. Proc. IEEE 2024, 112, 880–911.

  • 124.

    Ghasemi, M.; Heidari, S.; Kim, Y.G.; et al. Energy-Efficient, Delay-Constrained Edge Computing of a Network of DNNs. IEEE Trans. Comput. 2025, 74, 569–581.

  • 125.

    Ma, H.; Huang, Z.; Lu, G.; et al. Greening Edge AI: Optimizing Inference Accuracy and Reducing Carbon Emissions with Renewable Energy. IEEE Internet Things J. 2025, 12, 24300–24312.

Share this article:
How to Cite
Bi, J.; Yang, Y.; Yuan, H.; Wang, Z.; Li, J.; Zhai, J.; Zhang, J. Towards a Unified Framework for Large–Small Model Collaboration in Cloud–Edge Systems. Journal of Artificial Intelligence for Automation 2026, 1 (3), 12. https://doi.org/10.53941/jaia.2026.100012.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.