Co-authored Paper on Explainability of Large Language Models Accepted by Neurocomputing
2026-10-05
- article

A paper co-authored by Associate Professor Hayashi, "Explainability of large language models: Opportunities and challenges toward generating trustworthy explanations," has been accepted by Neurocomputing, an international academic journal published by Elsevier.
Title: Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
Authors: Shahin Atakishiyev, Housam K.B. Babiker, Jiayi Dai, Nawshad Farruque, Teruaki Hayashi, Nafisa Sadaf Hriti, Md Abed Rahman, Iain Smith, Mi-Young Kim, Osmar R. Zaïane, Randy Goebel
Journal: Neurocomputing (Elsevier, 2026)
https://www.sciencedirect.com/science/article/abs/pii/S0925231226026792
This study focuses on the explainability of large language models (LLMs), a topic that has drawn growing attention in recent years, and proposes a theoretical and empirical framework for improving their trustworthiness and effectiveness. As generative AI is increasingly applied to areas that form the foundations of society, such as healthcare and autonomous driving, there is an urgent need to understand "why a model produced a given output" so that users can trust its decisions. However, LLMs have complex internal structures, which makes the basis for their outputs difficult to see through.
To address this, the study systematically organizes how LLMs generate explanations and build trust from two perspectives: "local explainability," which clarifies the reasoning behind individual outputs, and "mechanistic interpretability," which analyzes the internal mechanisms by which the model operates. It further examines high-risk domains such as medical diagnosis support and autonomous driving, analyzing how explanations affect users' understanding and trust, and discusses why the faithfulness and clarity of explanations are essential for building that trust.
This paper offers a methodology for enhancing the transparency and accountability of AI systems, while also pointing toward new directions for implementing explainable AI (XAI). Going forward, advancing techniques for visualizing model structures and standardizing the evaluation of explanations is expected to contribute to realizing AI that is trusted by society.
This research is an outcome of collaborative work with members of the XAI Lab during Associate Professor Hayashi's time as a Visiting Professor at the University of Alberta in Canada. The article introducing the preprint release is available here.

