2607004541
  • Open Access
  • Article

Robust Aggregation Oriented Communication Compression via Differential Gradient Sparsification

  • Zhanbo Jin 1,   
  • Haozhao Wang 2,*,   
  • Senyao Li 2,   
  • Zhigang Zuo 2,   
  • Ruixuan Li 2

Received: 19 Apr 2026 | Revised: 03 Jul 2026 | Accepted: 06 Jul 2026 | Published: 23 Jul 2026

Abstract

Communication efficiency and robustness are two important concerns for distributed learning systems. Gradient sparsification that only transmits partial dimensions of the original gradient is one of the major methods to reduce the communication cost and gradient filter that filters out gradients with abnormal value is the mainstream to tolerate Byzantine attackers. However, in this paper, we identify that existing gradient sparsification techniques are not compatible with the gradient filter because the sparsified gradient may be viewed as abnormal and thus be filtered out. To tackle this challenge, we propose DGSB that sparsifies the differential gradient which is the difference between the gradients in two adjacent iterations instead of the original gradient. The principle is that the server can recover the approximately original gradient of each worker by using their sent sparsified differential gradient, and thus the server can apply the gradient filter over these recovered gradients which are close to the original gradients. Experiments conducted on various training tasks demonstrate that, in the presence of Byzantine workers, DGSB achieves over 30-fold communication compression while maintaining stable test accuracy in the evaluated settings, while the traditional gradient sparsification empowered with Byzantine-resilient methods sometimes cannot converge.

References 

  • 1.

    Harlap, A.; Tumanov, A.; Chung, A.; et al. Proteus: Agile ML Elasticity Through Tiered Reliability in Dynamic Resource Markets. In Proceedings of the Twelfth European Conference on Computer Systems (EuroSys), Belgrade, Serbia, 23–26 April 2017; pp. 589–604.

  • 2.

    Konecny, J.; McMahan, H.B.; Yu, F.X.; et al. Federated Learning: Strategies for Improving Communication Efficiency. arXiv 2016, arXiv:1610.05492.

  • 3.

    Li, M.; Andersen, D.G.; Park, J.W.; et al. Scaling Distributed Machine Learning with the Parameter Server. In Proceedings of the 11th USENIX conference on Operating Systems Design and Implementation, Broomfield, CO, USA, 6–8 October 2014; pp. 583–598.

  • 4.

    Aji, A.F.; Heafield, K. Sparse Communication for Distributed Gradient Descent. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP), Copenhagen, Denmark, 7–11 September 2017; pp. 440–445.

  • 5.

    Alistarh, D.; Hoefler, T.; Johansson, M.; et al. The Convergence of Sparsified Gradient Methods. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montr´eal, QC, Canada, 3–8 December 2018; pp. 5977–5987.

  • 6.

    Stich, S.U.; Cordonnier, J.; Jaggi, M. Sparsified SGD with Memory. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montr´eal, QC, Canada, 3–8 December 2018; pp. 4452–4463.

  • 7.

    Blanchard, P.; Mhamdi, E.M.E.; Guerraoui, R.; et al. Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; pp. 118–128.

  • 8.

    Chen, Y.; Su, L.; Xu, J. Distributed Statistical Machine Learning in Adversarial Settings: Byzantine Gradient Descent. In Proceedings of the ACM International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS), Irvine, CA, USA, 18–22 June 2018; p. 96.

  • 9.

    Xia, Q.; Tao, Z.; Hao, Z.; et al. FABA: An Algorithm for Fast Aggregation Against Byzantine Attacks in Distributed Neural Networks. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI), Macao, China, 10–16 August 2019; pp. 4824–4830.

  • 10.

    Ruder, S. An Overview of Gradient Descent Optimization Algorithms. arXiv 2016, arXiv:1609.04747.

  • 11.

    Yin, D.; Chen, Y.; Ramchandran, K.; et al. Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates. In Proceedings of the 35th International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 5650–5659.

  • 12.

    Xie, C.; Koyejo, O.; Gupta, I. Generalized Byzantine-Tolerant SGD. arXiv 2018, arXiv:1802.10116.

  • 13.

    Alistarh, D.; Allen-Zhu, Z.; Li, J. Byzantine Stochastic Gradient Descent. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montr´eal, QC, Canada, 3–8 December 2018; pp. 4618–4628.

  • 14.

    Chen, L.; Wang, H.; Charles, Z.; et al. DRACO: Byzantine-Resilient Distributed Training via Redundant Gradients. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 903–912.

  • 15.

    Li, L.; Xu, W.; Chen, T.; et al. RSA: Byzantine-Robust Stochastic Aggregation Methods for Distributed Learning from Heterogeneous Datasets. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; pp. 1544–1551.

  • 16.

    Damaskinos, G.; Mhamdi, E.M.E.; Guerraoui, R.; et al. Asynchronous Byzantine Machine Learning (the Case of SGD). In Proceedings of the 35th International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1145–1154.

  • 17.

    Yang, Y.; Li, W. BASGD: Buffered Asynchronous SGD for Byzantine Learning. arXiv 2020, arXiv:2003.00937.

  • 18.

    Liu, Y.; Chen, C.; Lyu, L.; et al. Byzantine-Robust Learning on Heterogeneous Data via Gradient Splitting. In Proceedings of the 40th International Conference on Machine Learning (ICML), Honolulu, HI, USA, 23–29 July 2023; Volume 202, pp. 21404–21425.

  • 19.

    Allouah, Y.; Guerraoui, R.; Gupta, N.; et al. Adaptive Gradient Clipping for Robust Federated Learning. In Proceedings of the 13th International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025.

  • 20.

    Hsieh, K.; Harlap, A.; Vijaykumar, N.; et al. Gaia: Geo-Distributed Machine Learning Approaching LAN Speeds. In Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI), Boston, MA, USA, 27–29 March 2017; pp. 629–647.

  • 21.

    Chen, C.; Choi, J.; Brand, D.; et al. AdaComp: Adaptive Residual Gradient Compression for Data-Parallel Distributed Training. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; pp. 2827–2835.

  • 22.

    Shi, S.; Zhao, K.; Wang, Q.; et al. A Convergence Analysis of Distributed SGD with Communication-Efficient Gradient Sparsification. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI), Macao, China, 10–16 August 2019; pp. 3411–3417.

  • 23.

    Sattler, F.; Wiedemann, S.; M¨uller, K.; et al. Robust and Communication-Efficient Federated Learning from Non-i.i.d. Data. IEEE Trans. Neural Netw. Learn. Syst. 2020, 31, 3400–3413.

  • 24.

    Lin, Y.; Han, S.; Mao, H.; et al. Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training. In Proceedings of the 6th International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018.

  • 25.

    Zhao, S.; Gao, H.; Li, W. On the Convergence of Memory-Based Distributed SGD. arXiv 2019, arXiv:1905.12960.

  • 26.

    Wangni, J.; Wang, J.; Liu, J.; et al. Gradient Sparsification for Communication-Efficient Distributed Optimization. In Proceedings of the 32nd Annual Conference on Neural Information Processing Systems (NeurIPS), Montr´eal, QC, Canada, 3–8 December 2018; pp. 1299–1309.

  • 27.

    Li, S.; Xu, W.; Wang, H.; et al. FedBAT: Communication-Efficient Federated Learning via Learnable Binarization. In Proceedings of the 41st International Conference on Machine Learning (ICML), Vienna, Austria, 21–27 July 2024; Volume 235, pp. 29074–29095.

  • 28.

    Kim, D.Y.; Han, D.J.; Seo, J.; et al. Achieving Lossless Gradient Sparsification via Mapping to Alternative Space in Federated Learning. In Proceedings of the 41st International Conference on Machine Learning (ICML), Vienna, Austria, 21–27 July 2024; Volume 235, pp. 23867–23900.

  • 29.

    Li, S.; Luo, X.; Wang, H.; et al. The Panaceas for Improving Low-Rank Decomposition in Communication-Efficient Federated Learning. In Proceedings of the 42nd International Conference on Machine Learning (ICML), Vancouver, BC, Canada, 13–19 July 2025; Volume 267, pp. 35536–35561.

  • 30.

    Wang, H.; Liu, X.; Niu, J.; et al. Why Go Full? Elevating Federated Learning Through Partial Network Updates. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 9–15 December 2024; pp. 99773–99799.

  • 31.

    Rammal, A.; Gruntkowska, K.; Fedin, N.; et al. Communication Compression for Byzantine Robust Learning: New Efficient Algorithms and Improved Rates. In Proceedings of the 27th International Conference on Artificial Intelligence and Statistics (AISTATS), Valencia, Spain, 2–4 May 2024; Volume 238, pp. 1207–1215.

  • 32.

    Zhou, F.; Cong, G. On the Convergence Properties of a K-Step Averaging Stochastic Gradient Descent Algorithm for Nonconvex Optimization. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI), Stockholm, Sweden, 13–19 July 2018; pp. 3219–3227.

  • 33.

    Zhang, X.; Wang, J.; Joshi, G.; et al. Machine Learning on Volatile Instances. In Proceedings of the IEEE Conference on Computer Communications (INFOCOM), Online, 6–9 July 2020; pp. 139–148.

  • 34.

    PyTorch. PyTorch MNIST Example. Available online: https://github.com/pytorch/examples/tree/main/mnist (accessed on 15 June 2026).

  • 35.

    Johnson, R.; Zhang, T. Accelerating Stochastic Gradient Descent Using Predictive Variance Reduction. In Proceedings of the 27th Annual Conference on Neural Information Processing Systems (NeurIPS), Lake Tahoe, NV, USA, 5–10 December 2013; pp. 315–323.

  • 36.

    Lei, L.; Jordan, M.I. Less Than a Single Pass: Stochastically Controlled Stochastic Gradient. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA, 20–22 April 2017; Volume 54, pp. 148–156.

  • 37.

    Zhou, D.; Xu, P.; Gu, Q. Stochastic Nested Variance Reduction for Nonconvex Optimization. In Proceedings of the 32nd Annual Conference on Neural Information Processing Systems (NeurIPS), Montr´eal, QC, Canada, 3–8 December 2018.

  • 38.

    Hosmer, D.W.; Hosmer, T.; Le Cessie, S.; et al. A Comparison of Goodness-of-Fit Tests for the Logistic Regression Model. Stat. Med. 1997, 16, 965–980.

  • 39.

    He, K.; Zhang, X.; Ren, S.; et al. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 770–778.

  • 40.

    LeCun, Y.; Cortes, C.; Burges, C.J.C. MNIST Database of Handwritten Digits [Dataset]. UCI Machine Learning Repository, 1998. Available online: https://archive.ics.uci.edu/dataset/683/mnist (accessed on 28 June 2026).

  • 41.

    Krizhevsky, A.; Hinton, G. Learning Multiple Layers of Features from Tiny Images; Technical Report; University of Toronto: Toronto, ON, Canada, 2009. Available online: https://www.cs.toronto.edu/˜kriz/cifar.html (accessed on 28 June 2026).

Share this article:
How to Cite
Jin, Z.; Wang, H.; Li, S.; Zuo, Z.; Li, R. Robust Aggregation Oriented Communication Compression via Differential Gradient Sparsification. Edge Intelligence and Systems 2026, 1 (1), 5.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.