GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
Abstract
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the black-box nature of their internal mechanisms has brought many concerns about their trustworthiness and interpretability. Recent research has discovered a class of abnormal tokens in the model's vocabulary space and named them "glitch tokens". Those tokens, once included in the input, may induce the model to produce incorrect, irrelevant, or even harmful results, drastically undermining the reliability and practicality of LLMs.
BibTeX
@inproceedings{Zhang-al:ASE24,
author = {Zhibo Zhang and
Wuxia Bai and
Yuxi Li and
Mark Huasong Meng and
Kailong Wang and
Ling Shi and
Li Li and
Jun Wang and
Haoyu Wang},
title = {{GlitchProber:} Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models},
booktitle = {ASE},
pages = {643--655},
publisher = {{ACM}},
year = {2024},
}