26,098 papers · page 71 of 1,305
Yash Mundhra, Max Valk, Maliheh Izadi
Large language models have shown impressive performance in various domains, including code generation across diverse open-source domains. However, their applicability in proprietary industrial settings, where domain-specific constraints and code interdependencies are prevalent, r…
Doha Nam, Jongmoon Baik
The prevalence of software vulnerabilities necessitates accurate and scalable detection techniques. While Pre-trained Language Models (PLMs) have shown strong potential in vulnerability analysis, most existing methods provide no explicit guidance on which parts of the input code …
Monil Narang, Hang Du, James A. Jones
In this work, we characterize a novel test-code smell—Disjoint Assertion Tangle (DAT)—which occurs when a test method verifies multiple, logically unrelated behaviors that can be separated. We propose a program analysis-based approach that automatically detects DAT and refactors …
Noor Nashid, Daniel Ding, Keheliya Gallaba, Ahmed E. Hassan, Ali Mesbah
Multi-hunk bugs, where fixes span disjoint regions of code, are common in practice, yet remain underrepresented in automated repair. Existing techniques and benchmarks predominantly target single-hunk scenarios, overlooking the added complexity of coordinating semantically relate…
Yannic Noller, Erick Chandra, Srinidhi Chandrashekar, Kenny T. W. Choo, Cyrille Jégourel, Oka Kurniawan, Christopher M. Poskitt
Debugging software, i.e., the localization of faults and their repair, is a key activity in software engineering. Therefore, effective and efficient debugging is one of the core skills a software engineer must develop. However, the teaching of debugging techniques is usually very…
Michael Norris, Syed Rafiul Hussain, Gang Tan
In an Internet of Things (IoT) environment, there are several way things can go wrong based on device activity. Poorly defined rules, conflicts between applications, physical interactions between devices, or unintentional interference by user behavior. Since these devices can hav…
Gustavo Ansaldi Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia, Haoxiang Zhang, Yihao Chen, Zhilong Chen, Arthur Leung, Dayi Lin + 2 more
High-quality labeled datasets are crucial for training and evaluating foundation models in software engineering, but creating them is often prohibitively expensive and labor-intensive. We introduce SPICE, a scalable, automated pipeline for labeling SWE-bench-style datasets with a…
Dillon Otto, Tanner Rowlett, Stefan Nagy
Desktop applications represent one of today’s largest software ecosystems, accounting for over 96% of workplace computing and supporting essential operations across critical sectors such as healthcare, commerce, industry, and government. Though modern software is increasingly bei…
Guangsheng Ou, Mingwei Liu, Yuxuan Chen, Yanlin Wang, Xin Peng, Zibin Zheng
Recent advancements in large language models (LLMs) have demonstrated impressive capabilities in code translation, typically evaluated using benchmarks like CodeTransOcean and RepoTransBench. However, dependency-free benchmarks fail to capture real-world complexities by focusing …
Ruwei Pan, Hongyu Zhang, Zhonghao Jiang, Ran Hou
With the increasing prevalence of fraudulent Android applications such as fake and malicious applications, it is crucial to detect them with high accuracy and adaptability. We present AgentDroid, a novel tool for Android fraudulent application detection based on multi-modal analy…
Haolin Pan, Xulin Zhou, Mingjie Xing, Yanjun Wu
Single Instruction, Multiple Data (SIMD) technology is crucial for enhancing computational efficiency in High-Performance Computing (HPC). While C++ SIMD libraries abstract away low-level complexities, their proliferation has led to a fragmented set of libraries, creating signifi…
Jai Parera, Nathan Huey, Ben Limpanukorn, Miryung Kim
Metamorphic testing (MT) is a powerful technique for software testing. We introduce Chrysalis, a lightweight, extensible logging and replay-based metamorphic testing framework in Python. Chrysalis allows developers to define custom input transformations and their associated invar…
Sungmin Park
Vulnerabilities in the Node.js ecosystem pose serious security threats. Generating exploits for such vulnerabilities is a critical and essential step for fixing the vulnerabilities and understanding attack vectors. To address this need, prior work has proposed a range of methods,…
Mrigank Pawagi, Lize Shao, Hyeonmin Lee, Yixin Sun, Wenxi Wang
Internet protocol specifications, published as Requests for Comments (RFCs) by the IETF organization, are essential to ensuring the interoperability, security, and reliability of the Internet. However, ambiguities in these specifications, particularly logical ambiguities such as …
Fatih Pehlivan, Arçin Ülkü Ergüzen, Sahand Moslemi Yengejeh, Mayasah Lami, Anil Koyuncu
Traditional static analysis methods struggle to detect semantic design flaws, such as violations of the SOLID principles, which require a strong understanding of object-oriented design patterns and principles. Existing solutions typically focus on individual SOLID principles or s…
Yun Peng, Kisub Kim, Linghan Meng, Kui Liu
Code review is an essential process to ensure the quality of software that identifies potential software issues at an early stage of software development. Among all software issues, security issues are the most important to identify, as they can easily lead to severe software cra…
Zhiyuan Peng, Xin Yin, Zijie Zhou, Chenhao Ying, Chao Ni, Yuan Luo
While Large Language Models (LLMs) have demonstrated remarkable progress in generating functionally correct Solidity code, they continue to face critical challenges in producing gas-efficient and secure code, which are critical requirements for real-world smart contract deploymen…
Tri Minh-Triet Pham, Diego Elias Costa, Weiyi Shang, Jinqiu Yang
Obstacle detection is crucial to the operation of autonomous driving systems, which rely on multiple sensors, such as cameras and LiDARs, combined with code logic and deep learning models to detect obstacles for time-sensitive decisions. Consequently, obstacle detection latency i…
Tri Minh-Triet Pham, Bo Yang, Jinqiu Yang
Autonomous driving systems (ADSs) rely on real-time sensor data, such as cameras and LiDARs, for time-critical decisions using deep neural networks. The accuracy of these decisions is crucial for the widespread adoption of ADSs, as errors can have serious consequences. 3D obstacl…
Kevin Pitstick, Alex Derr, Lihan Zhan, Sebastián Echeverría
Build reproducibility of container images is essential to ensure that deployed systems will work as expected and have not been tampered with. However, bit-by-bit reproducibility of container images is almost never achievable due to external factors, and it is also very slow and l…