3,458 papers · page 13 of 173
Benjamin Rombaut, Sogol Masoumzadeh, Kirill Vasilevski, Dayi Lin, Ahmed E. Hassan
Large language models (LLMs) are increasingly integrated into autonomous systems, giving rise to a new class of software known as Agentware, where LLM-powered agents perform complex, open-ended tasks in domains such as software engineering, customer service, and data analysis. Ho…
Bonan Ruan, Zhiwei Lin, Jiahao Liu, Chuqi Zhang, Kaihang Ji, Zhenkai Liang
Identifying the impact scope and scale is critical for software supply chain vulnerability assessment. However, existing studies face substantial limitations. First, prior studies either work at coarse package-level granularity—producing many false positives—or fail to accomplish…
Chandan Kumar Sah
Large Language Models (LLMs) have revolutionized AI, yet their inherent uncertainties pose significant challenges to reliable deployment. This paper presents a comprehensive systematic review of uncertainty in LLMs, bridging theoretical foundations and cutting-edge methodologies.…
Vasil Sarafov, David Markvica, Stefan Brunthaler
Fuzz testing has proven effective in discovering nontrivial bugs in complex, real-world systems, with coverage-guided greybox fuzzing being a key contributor to this success. Existing research has largely focused on developing new heuristics to increase code coverage, and current…
Benjamin Schmitz
Testing Graphical User Interface applications in non-graphical environments is challenging, especially for auto-graders in large-scale programming courses, where traditional approaches often depend on a graphical operating system. This paper presents a headless testing approach t…
Souhaila Serbout
Large Language Models (LLMs) increasingly power critical business processes, yet prompt robustness remains under-explored. Small variations—such as synonym changes or instruction reordering—can cause significant output shifts, undermining reliability in domains like customer serv…
Padmanabha Venkatagiri Seshadri, Harikrishnan Balagopal, Mehant Kammakomati, Ashok Pon Kumar, Dushyant Behl
Training-as-a-service platforms facilitate users to deploy pre-configured Generative AI training jobs as batch workloads. The immutability of configuration offers minimal flexibility to dynamically adapt to training progress. Existing approaches invariably involve manually monito…
Aarsh Shah, Cleyton V. C. de Magalhães, Kiev Gama, Ronnie de Souza Santos
Equity, diversity, and inclusion in software engineering often overlook neurodiversity, particularly the experiences of developers with Attention Deficit Hyperactivity Disorder (ADHD). Despite the growing awareness about that population in SE, few tools are designed to support th…
Kaveh Shahedi, Matthew Khouzam, Heng Li, Maxime Lamothe, Foutse Khomh
System tracing has become essential for understanding complex software behavior in modern systems, yet sophisticated trace analysis tools face significant adoption gaps in industrial settings. Through a year-long collaboration with Ericsson Montreal, developing TMLL (Trace-Server…
Mingyu Shao, Zhao Liu, Weihong Han, Cuiyun Gao, Jiachen Liu, Qing Liao
Network topology construction in this paper refers to designing the structural layouts and configuration rules among network devices according to natural language requirements in network simulation. Relatedly, Infrastructure as Code (IaC) enables the configuration and management …
Zhuoxiang Shen, Jiarun Dai, Yuan Zhang, Min Yang
The advantages of large language models (LLMs) in content comprehension and question answering have led to the rapid emergence of LLM agent. Developers across diverse domains are actively building their own agent applications (apps), as these apps can streamline workflows, boost …
Jingyi Shi, Yufeng Chen, Yang Xiao, Yuekang Li, Zhengzi Xu, Sihao Qiu, Chi Zhang, Keyu Qi + 5 more
Binary Function Similarity Detection (BFSD) is a foundational technique in software security, underpinning a wide range of applications including vulnerability detection, malware analysis. Recent advances in AI-based BFSD tools have led to significant performance improvements. Ho…
Yuling Shi, Yichun Qian, Hongyu Zhang, Beijun Shen, Xiaodong Gu
Code generation under long contexts is becoming increasingly critical as Large Language Models (LLMs) are required to reason over extensive information in the code-base. While recent advances enable code LLMs to process long inputs, high API costs and generation latency remain su…
Iti Shree, Karine Even-Mendoza, Tomasz Radzik
Existing LLM-based compiler fuzzers often produce syntactically or semantically invalid test programs, limiting their effectiveness in exercising compiler optimisations and backend components. We introduce ReFuzzer, a framework for refining LLM-generated test programs by systemat…
Weipeng Shuai, Jie Liu, Zhirou Ma, Liangyi Kang, Zehua Wang, Shuai Wang, Dan Ye, Hui Li + 2 more
Build failures are a major obstacle in RISC-V software migration, often involving complex interactions across logs, configurations, and environments. Traditional diagnostic tools struggle with the unstructured, multi-phase nature of build logs and lack semantic reasoning.We propo…
Sebastian Simon, Alina Mailach, Johannes Dorn, Norbert Siegmund
Configuration dependencies arise when multiple technologies in a software system require coordinated settings for correct interplay. Existing approaches for detecting such dependencies often yield high false-positive rates, require additional validation mechanisms, and are typica…
Vishal Singh, Ravi Shankar Das, Prajwal H. G, Subhajit Roy
We present our tool, AndroFL, that provides an infrastructure for an evolutionary algorithm-based test-suite generation backed by a statistical fault localization module for diagnosing faults. AndroFL’s evolutionary test-generator supports configurable fitness functions (e.g., co…
Marius Smytzek, Martin Eberlein, Tural Mammadov, Lars Grunske, Andreas Zeller
Program fixes must preserve passing tests while fixing failing ones. Validating these properties requires test oracles that distinguish passing from failing runs.We introduce BASHIRI, a tool that learns failure oracles from test suites with labeled outcomes using execution featur…
Yewei Song, Tiezhu Sun, Xunzhu Tang, Prateek Rajput, Tegawendé F. Bissyandé, Jacques Klein
Assessing the stability of code generation from large language models (LLMs) is essential for judging their reliability in real-world development. We extend prior "structural-entropy" concepts to the program domain by pairing entropy with abstract-syntax-tree (AST) analysis. For …
Yi Song, Dongchen Xie, Lin Xu, He Zhang, Chunying Zhou, Xiaoyuan Xie
For a vulnerability reported as an item of platforms such as CVE or NVD, software maintainers need to submit patches (in the form of code commit) to fix it, which is often performed silently for the sake of keeping products’ reputation or avoiding malicious attacks. But such a si…