2,847 papers · page 12 of 143
Junming Cao, Xuwen Xiang, Mingfei Cheng, Bihuan Chen, Xinyan Wang, You Lu, Chaofeng Sha, Xiaofei Xie + 1 more
Multiple machine learning (ML) models are often incorporated into real-world ML systems. However, updating an individual model in these ML systems frequently results in regression errors, where the new model performs worse than the old model for some inputs. While model-level reg…
Rajrupa Chattaraj, Sridhar Chimalakonda
In the realm of natural language processing (NLP), the rising computational demands of modern models bring energy efficiency to the forefront of sustainable computing. Preprocessing tasks, such as tokenization, stemming, and POS tagging, are critical steps in transforming raw tex…
Wei-Hao Chen, Jia Lin Cheoh, Manthan Keim, Sabine Brunswicker, Tianyi Zhang
Programming is an essential activity in data science (DS). Unlike regular software developers, DS programmers often use Jupyter notebooks instead of conventional IDEs. Moreover, DS programmers focus on statistics, data analytics, and modeling rather than writing production-ready …
Junjie Chen, Xingyu Fan, Chen Yang, Shuang Liu, Jun Sun
The compiler bug duplication problem (where many test failures are caused by the same compiler bug) can lead to huge waste of time and resource in diagnosing test failures produced by compiler testing. It is particularly challenging with regard to the silent compiler bugs that do…
Yuntianyi Chen, Yuqi Huai, Yirui He, Shilong Li, Changnam Hong, Qi Alfred Chen, Joshua Garcia
As autonomous driving systems (ADSes) become increasingly complex and integral to daily life, the importance of understanding the nature and mitigation of software bugs in these systems has grown correspondingly. Addressing the challenges of software maintenance in autonomous dri…
Jia Chen, Yuang He, Peng Wang, Xiaolei Chen, Jie Shi, Wei Wang
For online service systems, alerts are crucial for root cause analysis as they capture symptoms triggered by system faults. In real-world scenarios, a fault can propagate across multiple system components, generating a large volume of alerts. Various approaches have been proposed…
Mengzhuo Chen, Zhe Liu, Chunyang Chen, Junjie Wang, Boyu Wu, Jun Hu, Qing Wang
In software development, similar apps often encounter similar bugs due to shared functionalities and implementation methods. However, current automated GUI testing methods mainly focus on generating test scripts to cover more pages by analyzing the internal structure of the app, …
Zhenpeng Chen, Xinyue Li, Jie M. Zhang, Weisong Sun, Ying Xiao, Tianlin Li, Yiling Lou, Yang Liu
Fairness is a critical requirement for Machine Learning (ML) software, driving the development of numerous bias mitigation methods. Previous research has identified a leveling-down effect in bias mitigation for computer vision and natural language processing tasks, where fairness…
Yujia Chen, Yang Ye, Zhongqi Li, Yuchi Ma, Cuiyun Gao
Large code models (LCMs) have remarkably advanced the field of code generation. Despite their impressive capabilities, they still face practical deployment issues, such as high inference costs, limited accessibility of proprietary LCMs, and adaptability issues of ultra-large LCMs…
Xiao Chen, Hengcheng Zhu, Jialun Cao, Ming Wen, Shing-Chi Cheung
Debugging can be much facilitated if one can identify the evolution commit that introduced the bug leading to a detected failure (aka. bug-inducing commit, BIC). Although one may, in theory, locate BICs by executing the detected failing test on various historical commit versions,…
Gregorio Dalia, Andrea Di Sorbo, Corrado Aaron Visaggio, Gerardo Canfora
Nowadays, although it is widely known which behaviors of a software define a malware, there is a large gray area of invasive behaviors that do not make software necessarily harmful but act pervasively without the user’s perception. Being adequately informed of such behaviors is c…
Farbod Daneshyan, Runzhi He, Jianyu Wu, Minghui Zhou
The release note is a crucial document outlining changes in new software versions. It plays a key role in helping stakeholders recognise important changes and understand the implications behind them. Despite this fact, many developers view the process of writing software release …
Mouna Dhaouadi, Bentley Oakes, Michalis Famelis
Contributors to open source software must deeply understand a project’s history to make coherent decisions which do not conflict with past reasoning. However, inspecting all related changes to a proposed contribution requires intensive manual effort, and previous research has not…
Hridya Dhulipala, Aashish Yadavally, Smit Soneshbhai Patel, Tien N. Nguyen
While LLMs excel in understanding source code and descriptive texts for tasks like code generation, code completion, etc., they exhibit weaknesses in predicting dynamic program behavior, such as code coverage and runtime error detection, which typically require program execution.…
Xiaohu Du, Ming Wen, Haoyu Wang, Zichao Wei, Hai Jin
Code vulnerability detection is crucial to ensure software security. Recent advancements, particularly with the emergence of Code Pre-Trained Models (CodePTMs) and Large Language Models (LLMs), have led to significant progress in this area. However, these models are easily suscep…
Aryaz Eghbali, Felix Burk, Michael Pradel
Python is a dynamic language with applications in many domains, and one of the most popular languages in recent years. To increase code quality, developers have turned to “linters” that statically analyze the source code and warn about potential programming problems. However, the…
Meng Fan, Yuxia Zhang, Klaas-Jan Stol, Hui Liu
Continued contributions of core developers in open source software (OSS) projects are key for sustaining and maintaining successful OSS projects. A major risk to the sustainability of OSS projects is developer turnover. Prior studies have explored developer turnover at the level …
Tanner Finken, Jesse Chen, Sazzadur Rahaman
Protests are public expressions of personal or collective discontent with the current state of affairs. Although traditional protests involve in-person events, the ubiquity of computers and software opened up a new avenue for activism: protestware. Recent events in the Russo-Ukra…
Lukas Fruntke, Jens Krinke
Breaking changes in dependencies are a common challenge in software development, requiring manual intervention to resolve. This study examines how well Large Language Models (LLMs) automate the repair of breaking changes caused by dependency updates in Java projects. Although ear…
Yi Gao, Xing Hu, Xiaohu Yang, Xin Xia
Test smells arise from poor design practices and insufficient domain knowledge, which can lower the quality of test code and make it harder to maintain and update. Manually refactoring of test smells is time-consuming and error-prone, highlighting the necessity for automated appr…