3,458 papers · page 14 of 173
Ethan Stanley, Eric Eide
The WebAssembly Component Model is an emerging standard for assembling applications from parts that are implemented in WebAssembly. Unlike ordinary WebAssembly modules, components implement their interfaces in a standardized way, enabling the interoperation of software parts comp…
Tiezhu Sun, Marco Alecci, Yewei Song, Xunzhu Tang, Kisub Kim, Jordan Samhi, Tegawendé F. Bissyandé, Jacques Klein
Android malware detection and family classification have been extensively studied, yet localizing the exact malicious payloads within a detected sample remains a challenging and labor-intensive task. We propose RAML, a novel Retrieval-Augmented Malicious payload Localization pipe…
Zeyu Sun, Jingjing Liang, Weiyi Wang, Chenyao Suo, Junjie Chen, Fanjiang Xu
MLIR (Multi-Level Intermediate Representation) has rapidly become a foundational technology for modern compiler frameworks, enabling extensibility across diverse domains. However, ensuring the correctness and robustness of MLIR itself remains challenging. Existing fuzzing approac…
Yongqian Sun, Yu Luo, Xidao Wen, Yuan Yuan, Xiaohui Nie, Shenglin Zhang, Tong Liu, Xi Luo
Automated incident management plays a pivotal role in large-scale microservice systems. However, many existing methods rely solely on single-modal data (e.g., metrics, logs, and traces) and struggle to simultaneously address multiple downstream tasks, including anomaly detection …
Yongqian Sun, Mengyao Li, Xiao Xiong, Lei Tao, Yimin Zuo, Wenwei Gu, Shenglin Zhang, Junhua Kuang + 4 more
Timely detection of performance regression issues is critical to ensuring the stability and user experience of software systems. Traditional methods often rely on high-quality annotated data or data distribution assumptions, which cannot effectively adapt to performance changes i…
Ruoyu Sun, Da Song, Jiayang Song, Yuheng Huang, Lei Ma
As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in safety and robustness. To address these challenges, we introduce TrustVis, an automated evaluation fram…
Xinyu Sun, Fugen Tang, Yu Zhang, Han Shen, Chengru Song, Di Zhang
The growing demand for high-performance tensor programs on GPUs, especially for large language models (LLMs), necessitates advanced compilation and optimization techniques. However, the critical task of analyzing optimized, low-level PTX code for performance tuning or understandi…
Kairan Sun, Zhengzi Xu, Kaixuan Li, Lyuye Zhang, Yebo Feng, Daoyuan Wu, Yang Liu
Smart contract vulnerabilities continue to cause significant financial losses, despite the implementation of security measures such as manual audits and bug bounty platforms. A critical component often required by these security measures is the proof-of-concept (PoC) exploit, whi…
Kairan Sun, Zhengzi Xu, Kaixuan Li, Lyuye Zhang, Yuqiang Sun, Liwei Tan, Yang Liu
Web3 applications, particularly decentralized finance (DeFi) protocols, have grown rapidly with over $100 billion locked in smart contracts, attracting sophisticated attacks causing billions in losses. When attack occur, security analysts need to perform fault localization to ide…
Zhensu Sun, Chengran Yang, Xiaoning Du, Zhou Yang, Li Li, David Lo
Large language models (LLMs) have shown exceptional performance in code generation and understanding tasks, yet their high computational costs hinder broader adoption. One important factor is the inherent verbosity of programming languages, such as unnecessary formatting elements…
Fanny Febriani Susilo
As software systems grow in complexity—especially in cloud-based microservice architectures, automated testing has become crucial for reliability and security. While academic REST API fuzzing tools (e.g., EvoMaster, Schemathesis) show strong fault detection, their adoption in ind…
Aman Swaraj, Harsh Goyal, Sumit Chadgal, Sandeep Kumar
Identifying AI code plagiarism on technical forums like Stack Overflow (SO) is critical, as it can directly impact the platform’s trust and credibility. While previous studies have explored AI-generated code detection, they have focused on long, standalone samples from repositori…
Mahzabin Tamanna, Yash Chandrani, Matthew Burrows, Brandon Wroblewski, Laurie A. Williams, Dominik Wermke
Build scripts automate the process of compiling source code, managing dependencies, running tests, and packaging software into deployable artifacts. These scripts are ubiquitous in modern software development pipelines for streamlining testing and delivery. While developing build…
Honghao Tan, Haibo Wang, Diany Pressato, Yisen Xu, Shin Hwei Tan
Harmful content embedded in program elements within source code may have detrimental impact on mental health of software developers, and promote harmful behavior. Our key insight is that software developers may introduce harmful content into source code via diverse semantic-prese…
Kishanthan Thangarajah, Boyuan Chen, Shi Chang, Ahmed E. Hassan
AI-assisted coding tools powered by Code Large Language Models (CodeLLMs) are increasingly integrated into modern software development workflows. To address concerns around privacy, latency, and model customization, many enterprises opt to self-host these models. However, the div…
Alexi Turcotte, Neev Nirav Mehta
Data analysts need to be careful when they apply statistical inference techniques to data, as misuse of statistical inference methods can lead an analyst to draw the wrong conclusions. They need to be careful because, in the general case, misuse of statistics does not result in o…
Tahir Ullah, Waseem Akram, Fiza Khaliq, Hui Liu
Null Pointer Exceptions (NPEs) are one of the leading causes of software crashes and runtime errors. Although existing methods attempt to detect and classify NPE fixes, they often fall short due to irrelevant or noisy data, a lack of contextual understanding, and inefficiency in …
Marco Vieira, Priyam Ashish Shah, Bhavain Shah, Rrezarta Krasniqi
Large Language Models (LLMs) show great potential for automating code-related tasks. However, sound assessments are necessary to understand their true capabilities, particularly in code translation, where reliability is crucial. We introduce Polyglot, an automated, multi-language…
Zhonghan Wang
The Model-Constructing Satisfiability Calculus (MCSAT) framework has been applied to SMT problems over various arithmetic theories. NLSAT, an implementation using cylindrical algebraic decomposition (CAD) for explanation, is especially competitive for nonlinear real arithmetic (N…
Wei-Ji Wang
Assembly-oriented software architecture is transforming how modern systems are developed—shifting focus from writing code to composing business-aligned capabilities. While modular components are increasingly common, most remain confined within organizational boundaries, limiting …