Towards Interpreting Recurrent Neural Networks through Probabilistic Abstraction
Abstract
Neural networks are becoming a popular tool for solving many real-world problems such as object recognition and machine translation, thanks to its exceptional performance as an end-to-end solution. However, neural networks are complex black-box models, which hinders humans from interpreting and consequently trusting them in making critical decisions. Towards interpreting neural networks, several approaches have been proposed to extract simple deterministic models from neural networks. The results are not encouraging (e.g., low accuracy and limited scalability), fundamentally due to the limited expressiveness of such simple models.
BibTeX
@inproceedings{Dong-al:ASE20,
author = {Guoliang Dong and
Jingyi Wang and
Jun Sun and
Yang Zhang and
Xinyu Wang and
Ting Dai and
Jin Song Dong and
Xingen Wang},
title = {Towards Interpreting Recurrent Neural Networks through Probabilistic Abstraction},
booktitle = {ASE},
pages = {499--510},
publisher = {{IEEE}},
year = {2020},
}