2,847 papers · page 3 of 143
Kishan Kumar Ganguly, Tim Menzies
Much of software engineering (SE) research assumes that progress depends on massive datasets and CPU-intensive optimizers. Yet has this assumption been rigorously tested? The counter-evidence presented in this paper suggests otherwise. For over 100 optimization tasks from recent …
Yi Gao, Xing Hu, Xiaohu Yang, Xin Xia
Constructing user interfaces (UIs) is one of the most resource-intensive tasks in mobile development, often consuming more than half of overall effort. Although declarative frameworks such as Jetpack Compose (Android) and SwiftUI (iOS) have become mainstream, the majority of exis…
Yazhuo Gao, Lin Yang, Lianxiao Meng, Ran Zhu, Yining Cao
In large-scale microservice systems, multi-root-cause failures often intertwine, significantly increasing overall system risk and triggering a deluge of cascading alerts that pose serious challenges to fault diagnosis and recovery. Existing root-cause localization techniques rema…
Yi Gao, Ziyuan Zhang, Xing Hu, Xiaohu Yang, Xin Xia
Unit tests capture both functional checks and domain-specific knowledge, but this knowledge remains locked within individual projects and is rarely reused across libraries with overlapping functionality. Existing migration techniques based on structural code mappings (e.g., API s…
Fengjuan Gao, Qingjie Zhu, Yi Zhang, Yu Wang, Xuandong Li, Ke Wang
Control-flow reconstruction is a fundamental yet challenging problem in firmware analysis, particularly for stripped or raw-format binaries that lack symbolic metadata. Existing methods typically rely on syntax heuristics or format-specific patterns, which are often inadequate fo…
Yu Ge, Linna Xie, Zhong Li, Yu Pei, Tian Zhang
Large Language Model Powered Multi-Agent Systems (MASs) are increasingly employed to automate complex real-world tasks, such as programming and scientific discovery. While promising, MASs are not immune to defects or failures. Failure attribution in MASs, i.e., to pinpoint the sp…
Philipp Görz, Joschua Schilling, Nicolai Bissantz, Thorsten Holz
Fuzzing is a widely used technique to automatically test software for potential faults. To fuzz software projects efficiently and effectively, software developers must use fuzz harnesses , i.e., small programs that connect the fuzzer to the project’s code under test. However, as …
Min Gou, Zhiyu Yao, Hualong Ma, Ende Zhang, Jian Zhou, Fei He
Large language models (LLMs) are increasingly applied to code generation in IDEs, CI pipelines, and automated workflows. Existing evaluations, however, have largely focused on functionality, with comparatively limited attention to compliance with established safety standards. Thi…
Jingdong Guo, Chaopeng Dong, Yimo Ren, Siyuan Li, Jie Liu, Hong Li, Hongsong Zhu
Firmware lies at the heart of IoT devices. Its development depends heavily on third-party libraries (TPLs), which greatly accelerate the process but simultaneously introduce associated vulnerabilities. Binary Code Similarity Detection (BCSD) is an effective technique for identify…
Lianghong Guo, Yanlin Wang, Caihua Li, Wei Tao, Pengyu Yang, Jiachi Chen, Haoyu Song, Duyu Tang + 1 more
Constructing large-scale datasets for the GitHub issue resolution task is crucial for both training and evaluating the software engineering capabilities of Large Language Models (LLMs). However, the existing GitHub issue resolution data construction pipeline is challenging and la…
Jiaqi He, Karim Ali
Flow-sensitive pointer analysis offers highly precise results that are essential for various security analyses, bug detection tools, and compiler optimizations. However, its high computational cost often leads to prohibitively long analysis times, especially for large, real-world…
Yirui He, Ziyao He, Syed Fatiul Huq, Sam Malek
With over 60 percent of global Internet traffic originating from mobile devices, Responsive Web Design (RWD) has become essential for ensuring seamless user experiences across diverse screen sizes and resolutions. The Web Content Accessibility Guidelines require that both informa…
Kaifeng He, Mingwei Liu, Chong Wang, Zike Li, Yanlin Wang, Xin Peng, Zibin Zheng
Code generation with large language models (LLMs) is highly sensitive to token selection during decoding, particularly at uncertain decision points that influence program logic. While standard strategies such as greedy decoding treat all tokens uniformly, they overlook code-speci…
Jongchan Hong, Jaewon Kim, Sungjae Hwang
Electric vehicles (EVs) are being rapidly adopted, with over 61,000 publicly accessible charging stations deployed across the United States as of 2024. A core component of this infrastructure is the Charging Station Management System (CSMS), which is responsible for security-crit…
Xiaohui Hu, Ningyu He, Haoyu Wang
Serving as the first touch point for users to the cryptocurrency world, cryptocurrency wallets allow users to manage, receive, and transmit digital assets on blockchains and interact with emerging decentralized finance (DeFi) applications. Unfortunately, cryptocurrency wallets ha…
Wenbo Hu, Jie Lu, Jingting Chen, Feng Li, Chenghang Shi, Xiaonan Shi, Jinchen Wang, Wei Huo
HTTP API specifications are essential for modern web development, yet existing tools fail in production environments due to multi-layer routing. Production deployments employ infrastructure-level and framework-level routing that apply sequential rewrite and dispatch rules, creati…
Chao Hu, Wenhao Zeng, Yuling Shi, Beijun Shen, Xiaodong Gu
Repository-level code generation has attracted growing attention in recent years. Unlike function-level code generation, it requires the model to understand the entire repository and reason over complex dependencies across functions, classes, and modules. However, existing approa…
Haocheng Huang, Yuchen Chen, Weisong Sun, Peizhuo Lv, Yuan Xiao, Chunrong Fang, Yang Liu, Xiaofang Zhang
Constructing and curating high-quality code datasets requires significant resources, making them valuable intellectual property. Unfortunately, these datasets currently face severe risks of unauthorized use. Although digital watermarking offers a post hoc mechanism for copyright …
Yifan Huang, Xiaojun Jia, Wenbo Guo, Yuqiang Sun, Yihao Huang, Chong Wang, Yang Liu
Large language models (LLMs) have revolutionized software development through AI-assisted coding tools, enabling developers with limited programming expertise to create sophisticated applications. This democratization of software development has significantly lowered the barriers…
Yuan Huang, Yukang Zhou, Xiangping Chen, Zibin Zheng
With the rapid development of large language models in code generation, AI-powered editors such as GitHub Copilot and Cursor are revolutionizing software development practices. At the same time, studies have identified potential defects in the generated code. Previous research ha…