SemantiLog: Log-based Anomaly Detection with Semantic Similarity
Abstract
Logs produced by software applications are invaluable for spotting deviations from expected system behavior. However, automatically detecting anomalies from log data is challenging due to the volume, semi-structured nature, lack of standard formatting, and potential evolution of log records over time. In this work, we approach log-based anomaly detection as a semantic similarity problem. We generate pairwise similarity scores using a general-purpose pre-trained language model and further augment them with ground-truth binary labels. The generated similarity labels supervise an encoder trained for semantic similarity. At inference time, anomalies are detected based on the cosine similarity between the encoded query sequence and the average normal encoding. Our method outperforms contemporary techniques on multiple benchmarks without template extraction or a fixed vocabulary and achieves competitive performance even when provided with limited abnormal examples.
BibTeX
@inproceedings{Shavit-al:ASE24,
author = {Yoli Shavit and
Kathy Razmadze and
Gary Mataev and
Hanan Shteingart and
Eitan Zahavi and
Zachi Binshtock},
title = {{SemantiLog:} Log-based Anomaly Detection with Semantic Similarity},
booktitle = {ASE},
pages = {2438--2439},
publisher = {{ACM}},
year = {2024},
}