2609005156
  • Open Access
  • Article

Rethinking Attention: Entropy-Guided Iterative Refinement for Language Models

  • Subham Divakar 1,*,   
  • Rojalina Priyadarshini 2,*

Received: 26 Jul 2026 | Revised: 29 Aug 2026 | Accepted: 10 Sep 2026 | Published: 29 Sep 2026

Abstract

Standard Multi-Head Attention (MHA) computes contextual representations in a single forward pass, providing no mechanism to detect or correct diffuse or suboptimal attention distributions. Existing efforts to improve attention have primarily targeted computational efficiency or sequence length, leaving attention quality itself largely unaddressed. We propose IMHA (Entropy-Guided Iterative Multi-Head Attention), a within-layer attention formulation that progressively refines token representations over T iterations using shared projection weights. At each step, the Shannon entropy of each token’s per-head attention row—normalised against the position-dependent maximum to remove a structural artefact of causal masking—serves as an intrinsic quality signal: tokens with peaked (low-entropy) attention are amplified while those with diffuse (high-entropy) attention are suppressed. A gated residual update incorporating a prefix-causal global context vector ensures stable convergence. The formulation is strictly causal, verified at the gradient level across all ablation configurations. We prove (Proposition 1) that under mild assumptions the tokenspecific content of an attention output decays monotonically with attention diffuseness, providing principled motivation for entropy weighting. We further show (Proposition 2) that if iterative refinement sharpens the attention distribution, discriminative content is non-decreasing across iterations. We evaluate across three model scales (Small 14 M, Medium 45 M, Large 87 M parameters) on WikiText-2, Penn Treebank, and TinyStories, with downstream transfer to SST-2 sentiment classification and AG News topic classification. At medium scale, IMHA reduces WikiText-2 perplexity from 47.30 ± 0.18 to 44.85 ± 0.16 over the MHA baseline (paired t-test, p = 3.2 × 10−6, d = 11.4, after Holm–Bonferroni correction over 18 comparisons), outperforms parameter-matched capacity baselines by 1.40 points, and surpasses attention-quality methods including α-entmax, RealFormer, Universal Transformer (FLOPs-matched), and entropy-regularised MHA. Critically, IMHA outperforms entropy-regularised MHA by 1.18 points, demonstrating that the internal, within-layer delivery of the entropy signal—rather than entropy as an external loss term—is what drives the improvement. The relative gain is stable across scales (4.8%, 4.9%, 4.6% at Small, Medium, and Large respectively), and the pretrained representations transfer to classification: IMHA achieves 93.2 ± 0.3% on SST-2 vs. 91.8 ± 0.4% for MHA, and 90.1 ± 0.4% vs. 88.7 ± 0.5% on AG News. Total overhead is a uniform ∼8% across parameters, FLOPs, training time, and inference latency.

References 

  • 1.

    Vaswani, A.; Shazeer, N.; Parmar, N.; et al. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017.

  • 2.

    Radford, A.; Narasimhan, K.; Salimans, T.; et al. Improving Language Understanding by Generative Pre-Training. 2018. Available online: https://www.mikecaptain.com/resources/pdf/GPT-1.pdf (accessed on 1 January 2026).

  • 3.

    Devlin, J.; Chang, M.W.; Lee, K.; et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. https://doi.org/10.18653/v1/n19-1423.

  • 4.

    Voita, E.; Talbot, D.; Moiseev, F.; et al. Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 5797–5808. https://doi.org/10.18653/v1/p19-1580.

  • 5.

    Jain, S.; Wallace, B.C. Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 3543–3556. https://doi.org/10.18653/v1/n19-1357.

  • 6.

    Katharopoulos, A.; Vyas, A.; Pappas, N.; et al. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In Proceedings of the International Conference on Machine Learning, Virtual, 13–18 July 2020; pp. 5156–5165.

  • 7.

    Choromanski, K.; Likhosherstov, V.; Dohan, D.; et al. Rethinking attention with performers. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021.

  • 8.

    Wang, S.; Li, B.Z.; Khabsa, M.; et al. Linformer: Self-Attention with Linear Complexity. arXiv 2020, arXiv:2006.04768.

  • 9.

    Kitaev, N.; Kaiser, Ł.; Levskaya, A. Reformer: The efficient transformer. In Proceedings of the International Conference on Learning Representations, Virtual, 26 April–1 May 2020.

  • 10.

    Xiong, Y.; Zeng, Z.; Chakraborty, R.; et al. Nystromformer: A Nystr¨om-based Algorithm for Approximating Self-Attention. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021; Volume 35, pp. 14138–14148. https://doi.org/10.1609/aaai.v35i16.17664.

  • 11.

    Dao, T.; Fu, D.Y.; Ermon, S.; et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022.

  • 12.

    Dai, Z.; Yang, Z.; Yang, Y.; et al. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019. https://doi.org/10.18653/v1/p19-1285.

  • 13.

    Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The Long-Document Transformer. arXiv 2020, arXiv:2004.05150.

  • 14.

    Zaheer, M.; Guruganesh, G.; Dubey, A.; et al. Big Bird: Transformers for Longer Sequences. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020.

  • 15.

    Sukhbaatar, S.; Grave, E.; Bojanowski, P.; et al. Adaptive Attention Span in Transformers. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 331–335. https://doi.org/10.18653/v1/p19-1032.

  • 16.

    Dehghani, M.; Gouws, S.; Vinyals, O.; et al. Universal Transformers. In Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA, 6–9 May 2019.

  • 17.

    Bai, S.; Kolter, J.Z.; Koltun, V. Deep equilibrium models. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019.

  • 18.

    Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x.

  • 19.

    Roy, A.; Saffar, M.; Vaswani, A.; et al. Efficient Content-Based Sparse Attention with Routing Transformers. Trans. Assoc. Comput. Linguist. 2021, 9, 53–68. https://doi.org/10.1162/tacl a 00353.

  • 20.

    Tay, Y.; Dehghani, M.; Bahri, D.; et al. Efficient Transformers: A Survey. ACM Comput. Surv. 2022, 55, 109. https://doi.org/10.1145/3530811.

  • 21.

    Correia, G.M.; Niculae, V.; Martins, A.F.T. Adaptively Sparse Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; pp. 2174–2184. https://doi.org/10.18653/v1/d19-1223.

  • 22.

    He, R.; Ravula, A.; Kanagal, B.; et al. RealFormer: Transformer Likes Residual Attention. In Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, Virtual, 1–6 August 2021; pp. 929–943. https://doi.org/10.18653/v1/2021.findings-acl.81.

  • 23.

    Le, H.; Sahoo, D.; Chen, N.; et al. Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 5612–5623.

  • 24.

    Zhang, B.; Titov, I.; Sennrich, R. Sparse Attention with Linear Units. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021; pp. 6507–6520. https://doi.org/10.18653/v1/2021.emnlp-main.523.

  • 25.

    Divakar, S.; Priyadarshini, R. Recursive Additive Attention Based Convolutional Neural Network Model for Plant Disease Detection. In Proceedings of the 2024 IEEE 21st India Council International Conference (INDICON), Kharagpur, India, 19–21 December 2024. https://doi.org/10.1109/indicon63790.2024.10958327.

  • 26.

    Graves, A. Adaptive computation time for recurrent neural networks. arXiv 2016, arXiv:1603.08983.

  • 27.

    Tishby, N.; Zaslavsky, N. Deep learning and the information bottleneck principle. In Proceedings of the 2015 IEEE Information TheoryWorkshop (ITW), Jerusalem, Israel, 26 April–1 May 2015; pp. 1–5. https://doi.org/10.1109/itw.2015.7133169.

  • 28.

    Cho, K.; van Merrienboer, B.; Gulcehre, C.; et al. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; pp. 1724–1734. https://doi.org/10.3115/v1/d14-1179.

  • 29.

    Merity, S.; Xiong, C.; Bradbury, J.; et al. Pointer sentinel mixture models. In Proceedings of the 5th International Conference on Learning Representations, Toulon, France, 24–26 April 2017.

  • 30.

    Marcus, M.P.; Santorini, B.; Marcinkiewicz, M.A. Building a Large Annotated Corpus of English: The Penn Treebank. Comput. Linguist. 1993, 19, 313–330.

  • 31.

    Eldan, R.; Li, Y. TinyStories: How Small Can Language Models Be and Still Speak Coherent English? arXiv 2023, arXiv:2305.07759.

  • 32.

    Paperno, D.; Kruszewski, G.; Lazaridou, A.; et al. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Berlin, Germany, 7–12 August 2016.

  • 33.

    Socher, R.; Perelygin, A.; Wu, J.; et al. Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Seattle, WA, USA, 18–21 October 2013. https://doi.org/10.18653/v1/d13-1170.

  • 34.

    Zhang, X.; Zhao, J.; LeCun, Y. Character-level Convolutional Networks for Text Classification. In Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada, 7–12 December 2015.

Share this article:
How to Cite
Divakar, S.; Priyadarshini, R. Rethinking Attention: Entropy-Guided Iterative Refinement for Language Models. Transactions on Artificial Intelligence 2026, 2 (1), 214–230. https://doi.org/10.53941/tai.2026.100014.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.
Article Metrics
14
Article Views
0
Citations