26,098 papers · page 92 of 1,305
Ziyao He, Syed Fatiul Huq, Sam Malek
Websites are integral to people’s daily lives, with billions in use today. However, due to limited awareness of accessibility and its guidelines, developers often release web apps that are inaccessible to people with disabilities, who make up around 16% of the global population. …
Hao He, Bogdan Vasilescu, Christian Kästner
Recent high-profile incidents in open-source software have greatly raised practitioner attention on software supply chain attacks. To guard against potential malicious package updates, security practitioners advocate pinning dependency to specific versions rather than floating in…
Soneya Binta Hossain, Raygan Taylor, Matthew B. Dwyer
Code documentation is a critical artifact of software development, bridging human understanding and machine- readable code. Beyond aiding developers in code comprehension and maintenance, documentation also plays a critical role in automating various software engineering tasks, s…
Katherine Hough, Jonathan Bell
Dynamic taint tracking is a program analysis that traces the flow of information through a program. In the Java virtual machine (JVM), there are two prominent approaches for dynamic taint tracking: “shadowing” and “mirroring”. Shadowing is able to precisely track information flow…
Huimin Hu, Yingying Wang, Julia Rubin, Michael Pradel
Scalable static analyzers are popular tools for finding incorrect, inefficient, insecure, and hard-to-maintain code early during the development process. Because not all warnings reported by a static analyzer are immediately useful to developers, many static analyzers provide a w…
Junjie Huang, Zhihan Jiang, Zhuangbin Chen, Michael R. Lyu
Log parsing serves as an essential prerequisite for various log analysis tasks. Recent advancements in this field have improved parsing accuracy by leveraging the semantics in logs through fine-tuning large language models (LLMs) or learning from in-context demonstrations. Howeve…
Ali Reza Ibrahimzada, Kaiyao Ke, Mrigank Pawagi, Muhammad Salman Abid, Rangeet Pan, Saurabh Sinha, Reyhaneh Jabbarvand
Code translation transforms programs from one programming language (PL) to another. One prominent use case is application modernization to enhance maintainability and reliability. Several rule-based transpilers have been designed to automate code translation between different pai…
Sayem Mohammad Imtiaz, Astha Singh, Fraol Batole, Hridesh Rajan
Not a day goes by without hearing about the impressive feats of large language models (LLMs), and equally, not a day passes without hearing about their challenges. LLMs are notoriously vulnerable to biases in their dataset, leading to issues such as toxicity, harmful responses, a…
Sujin Jang, Yeonhee Ryou, Heewon Lee, Kihong Heo
We present U nit C on , a system for synthesizing targeted unit tests for runtime exceptions in Java programs. Targeted unit tests aim to reveal a bug at a specific location in the program under test. This capability benefits various tasks in software development, such as patch t…
Md Mahir Asef Kabir, Xiaoyin Wang, Na Meng
When building enterprise applications (EAs) on Java frameworks (e.g., Spring), developers often configure application components via metadata (i.e., Java annotations and XML files). It is challenging for developers to correctly use metadata, because the usage rules can be complex…
Nima Karimipour, Erfan Arvan, Martin Kellogg, Manu Sridharan
Null-pointer exceptions are serious problem for Java, and researchers have developed type-based nullness checking tools to prevent them. These tools, however, have a downside: they require developers to write nullability annotations, which is time-consuming and hinders adoption. …
Sayali Kate, Yifei Gao, Shiwei Feng, Xiangyu Zhang
Increasingly popular Robot Operating System (ROS) framework allows building robotic systems by integrating newly developed and/or reused modules, where the modules can use different versions of the framework (e.g., ROS1 or ROS2) and programming language (e.g. C++ or Python). The …
Haris Ali Khan, Yanjie Jiang, Qasim Umer, Yuxia Zhang, Waseem Akram, Hui Liu
It is often valuable to know whether a given piece of source code has or hasn’t been used to train a given deep learning model. On one side, it helps avoid data contamination problems that may exaggerate the performance of evaluated models. Conversely, it facilitates copyright pr…
Myeongsoo Kim, Saurabh Sinha, Alessandro Orso
Modern web services rely heavily on REST APIs, typically documented using the OpenAPI specification. The widespread adoption of this standard has resulted in the development of many black-box testing tools that generate tests based on OpenAPI specifications. Although Large Langua…
Jiaolong Kong, Xiaofei Xie, Shangqing Liu
Large Language Models (LLMs) have achieved remarkable success in various applications, particularly in coderelated tasks such as code generation and program repair, setting new performance benchmarks. However, the extensive use of large training corpora raises concerns about whet…
Ziqiao Kong, Cen Zhang, Maoyi Xie, Ming Hu, Yue Xue, Ye Liu, Haijun Wang, Yang Liu
Billions of dollars are transacted through smart contracts, making vulnerabilities a major financial risk. One focus in the security arms race is on profitable vulnerabilities that attackers can exploit. Fuzzing is a key method for identifying these vulnerabilities. However, curr…
Bruno Kreyssig, Alexandre Bartel
While multiple recent publications on detecting Java Deserialization Vulnerabilities highlight an increasing relevance of the topic, until now no proper benchmark has been established to evaluate the individual approaches. Hence, it has become increasingly difficult to show impro…
Tanakorn Leesatapornwongsa, Fazle Elahi Faisal, Suman Nath
Failure reproduction is a crucial step for debugging software systems, but it is often challenging and timeconsuming, especially when the failures are caused by complex inputs, states, or environments. In this paper, we present ReproCopilot, a tool that leverages program analysis…
Kyla Levin, Nicolas van Kempen, Emery D. Berger, Stephen N. Freund
Debugging is a critical but challenging task for programmers. This paper proposes ChatDBG, an AI-powered debugging assistant. ChatDBG integrates large language models (LLMs) to significantly enhance the capabilities and user-friendliness of conventional debuggers. ChatDBG lets pr…
Haodong Li, Xiao Cheng, Guohan Zhang, Guosheng Xu, Guoai Xu, Haoyu Wang
Learning-based Android malware detection has earned significant recognition across industry and academia, yet its effectiveness hinges on the accuracy of labeled training data. Manual labeling, being prohibitively expensive, has prompted the use of automated methods, such as leve…