3,458 papers · page 11 of 173
Masashi Mizoguchi, Kentaro Yoshimura, Keita Nakazawa, Yasuomi Sato, Takahiro Iida, Fumio Narisawa
Improving the efficiency of software integration testing is a critical challenge in the automotive industry, particularly as Electronic Control Unit (ECU) architectures become increasingly complex. This paper addresses the automation of integration test script generation by lever…
Audris Mockus, Peter C. Rigby, Rui Abreu, Anatoly Akkerman, Yogesh Bhootada, Payal Bhuptani, Gurnit Ghardhora, Lan Hoang Dao + 15 more
The focus on rapid software delivery inevitably results in the accumulation of technical debt, which, in turn, affects quality and slows future development. Our primary aim is to discover how companies keep their codebases maintainable and how code improvements might be automated…
Sudharssan Mohan, Kyeongseok Yang, Zelun Kong, Yonghwi Kwon, Junghwan Rhee, Tyler Summers, Hongjun Choi, Heejo Lee + 1 more
Robotic aerial vehicles (RAVs), particularly drones, are crucial in civil and military sectors. However, researchers have found that adversaries can inject noise into sensor measurements and cause physical impacts on the RAVs like crashes. Although identifying such signal injecti…
Facundo Molina, Nazareno Aguirre, Alessandra Gorla
The effectiveness of testing in uncovering software defects depends not only on the characteristics of the test inputs and how thoroughly they exercise the software, but also on the quality of the oracles used to determine whether the software behaves as expected. Therefore, asse…
Davide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst, Mauro Pezzè
Generation of thorough test oracles is an open problem. Popular test case generators, like EvoSuite and Randoop, rely on implicit, rule-based, and regression oracles that miss failures that depend on the semantics of the program under test. Formal specifications can yield test or…
Akira Mori, Masatomo Hashimoto
Three-way merge tools play crucial roles in modern software development, where a developer forks a branch to make local modifications and requests it to be merged into the main branch via a "pull request." Despite its importance, the task has traditionally been defined in an intu…
Dany Moshkovich, Sergey Zeltyn
Large Language Models (LLMs) are increasingly deployed within agentic systems—collections of interacting, LLM-powered agents that execute complex, adaptive workflows using memory, tools, and dynamic planning. While enabling powerful new capabilities, these systems also introduce …
Sali Moussa
Autonomous Driving Systems (ADS) must reliably perceive and react to complex environments, even when sensor blind spots obscure critical objects. While existing testing methods often focus on dynamic interactions, they significantly underestimate safety risks arising from both dy…
Prasita Mukherjee, Minghai Lu, Benjamin Delaware
We present SynVer — a novel, general purpose synthesizer for C programs equipped with machine-checked proofs of correctness using the Verified Software Toolchain. To do so, SynVer employs two Large Language Models (LLMs): the first generates candidate programs from user-provided …
Yash Mundhra, Max Valk, Maliheh Izadi
Large language models have shown impressive performance in various domains, including code generation across diverse open-source domains. However, their applicability in proprietary industrial settings, where domain-specific constraints and code interdependencies are prevalent, r…
Doha Nam, Jongmoon Baik
The prevalence of software vulnerabilities necessitates accurate and scalable detection techniques. While Pre-trained Language Models (PLMs) have shown strong potential in vulnerability analysis, most existing methods provide no explicit guidance on which parts of the input code …
Monil Narang, Hang Du, James A. Jones
In this work, we characterize a novel test-code smell—Disjoint Assertion Tangle (DAT)—which occurs when a test method verifies multiple, logically unrelated behaviors that can be separated. We propose a program analysis-based approach that automatically detects DAT and refactors …
Noor Nashid, Daniel Ding, Keheliya Gallaba, Ahmed E. Hassan, Ali Mesbah
Multi-hunk bugs, where fixes span disjoint regions of code, are common in practice, yet remain underrepresented in automated repair. Existing techniques and benchmarks predominantly target single-hunk scenarios, overlooking the added complexity of coordinating semantically relate…
Yannic Noller, Erick Chandra, Srinidhi Chandrashekar, Kenny T. W. Choo, Cyrille Jégourel, Oka Kurniawan, Christopher M. Poskitt
Debugging software, i.e., the localization of faults and their repair, is a key activity in software engineering. Therefore, effective and efficient debugging is one of the core skills a software engineer must develop. However, the teaching of debugging techniques is usually very…
Michael Norris, Syed Rafiul Hussain, Gang Tan
In an Internet of Things (IoT) environment, there are several way things can go wrong based on device activity. Poorly defined rules, conflicts between applications, physical interactions between devices, or unintentional interference by user behavior. Since these devices can hav…
Gustavo Ansaldi Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia, Haoxiang Zhang, Yihao Chen, Zhilong Chen, Arthur Leung, Dayi Lin + 2 more
High-quality labeled datasets are crucial for training and evaluating foundation models in software engineering, but creating them is often prohibitively expensive and labor-intensive. We introduce SPICE, a scalable, automated pipeline for labeling SWE-bench-style datasets with a…
Dillon Otto, Tanner Rowlett, Stefan Nagy
Desktop applications represent one of today’s largest software ecosystems, accounting for over 96% of workplace computing and supporting essential operations across critical sectors such as healthcare, commerce, industry, and government. Though modern software is increasingly bei…
Guangsheng Ou, Mingwei Liu, Yuxuan Chen, Yanlin Wang, Xin Peng, Zibin Zheng
Recent advancements in large language models (LLMs) have demonstrated impressive capabilities in code translation, typically evaluated using benchmarks like CodeTransOcean and RepoTransBench. However, dependency-free benchmarks fail to capture real-world complexities by focusing …
Ruwei Pan, Hongyu Zhang, Zhonghao Jiang, Ran Hou
With the increasing prevalence of fraudulent Android applications such as fake and malicious applications, it is crucial to detect them with high accuracy and adaptability. We present AgentDroid, a novel tool for Android fraudulent application detection based on multi-modal analy…
Haolin Pan, Xulin Zhou, Mingjie Xing, Yanjun Wu
Single Instruction, Multiple Data (SIMD) technology is crucial for enhancing computational efficiency in High-Performance Computing (HPC). While C++ SIMD libraries abstract away low-level complexities, their proliferation has led to a fragmented set of libraries, creating signifi…