kirancodes.me
To Proof Maintenance & Beyond!

GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models

Zhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng, Kailong Wang, Ling Shi, Li Li, Jun Wang, Haoyu Wang

Abstract

Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the black-box nature of their internal mechanisms has brought many concerns about their trustworthiness and interpretability. Recent research has discovered a class of abnormal tokens in the model's vocabulary space and named them "glitch tokens". Those tokens, once included in the input, may induce the model to produce incorrect, irrelevant, or even harmful results, drastically undermining the reliability and practicality of LLMs.

BibTeX
@inproceedings{Zhang-al:ASE24,
  author    = {Zhibo Zhang and
               Wuxia Bai and
               Yuxi Li and
               Mark Huasong Meng and
               Kailong Wang and
               Ling Shi and
               Li Li and
               Jun Wang and
               Haoyu Wang},
  title     = {{GlitchProber:} Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models},
  booktitle = {ASE},
  pages     = {643--655},
  publisher = {{ACM}},
  year      = {2024},
}

Related papers