ICSE 2025
245 papers
- "Get Me in the Groove": a Mixed Methods Study on Supporting Adhd Professional Programmers
- $ZTD_{\text{JAVA}}$: Mitigating Software Supply Chain Vulnerabilities via Zero-Trust Dependencies
- $\mu \text{PRL}$: A Mutation Testing Pipeline for Deep Reinforcement Learning Based on Real Faults
- 3DGen: AI-Assisted Generation of Provably Correct Binary Format Parsers
- A Catalog of Micro Frontends Anti-Patterns
- A Differential Testing Framework to Identify Critical AV Failures Leveraging Arbitrary Inputs
- A First Look at Conventional Commits Classification
- A Large-Scale Study of Model Integration in ML-Enabled Software Systems
- A Little Goes a Long Way: Tuning Configuration Selection for Continuous Kernel Fuzzing
- A Multi-Agent Approach for REST API Testing with Semantic Graphs and LLM-Driven Inputs
- A Multiple Representation Transformer with Optimized Abstract Syntax Tree for Efficient Code Clone Detection
- A Study of Undefined Behavior Across Foreign Function Boundaries in Rust Libraries
- A Tale of Two DL Cities: When Library Tests Meet Compiler
- A Test Oracle for Reinforcement Learning Software Based on Lyapunov Stability Control Theory
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service Systems
- Accessibility Issues in Ad-Driven Web Applications
- Accounting for Missing Events in Statistical Information Leakage Analysis
- Aligning the Objective of LLM-Based Program Repair
- An Empirical Study of Proxy Contracts at the Ethereum Ecosystem Scale
- An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far are We?
- An Empirical Study on Commit Message Generation Using LLMs via In-Context Learning
- An Empirical Study on Package-Level Deprecation in Python Ecosystem
- An Empirical Study on Reproducible Packaging in Open-Source Ecosystems
- An Exploratory Study of ML Sketches and Visual Code Assistants
- An Exploratory Study on the Engineering of Security Features
- An Extensive Empirical Study of Nondeterministic Behavior in Static Analysis Tools
- An LLM-Based Agent-Oriented Approach for Automated Code Design Issue Localization
- Analyzing the Feasibility of Adopting Google's Nonce-Based CSP Solutions on Websites
- Answering User Questions About Machine Learning Models Through Standardized Model Cards
- Are LLMs Correctly Integrated into Software Systems?
- Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection Solutions
- AssetHarvester: A Static Analysis Tool for Detecting Secret-Asset Pairs in Software Artifacts
- Automated Accessibility Analysis of Dynamic Content Changes on Mobile Apps
- Automated Generation of Accessibility Test Reports from Recorded User Transcripts
- Automated Test Generation For Smart Contracts via On-Chain Test Case Augmentation and Migration
- Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly Detection
- Automating a Complete Software Test Process Using LLMs: An Automotive Case Study
- BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks
- BSan: A Powerful Identifier-Based Hardware-Independent Memory Error Detector for COTS Binaries
- Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers
- Boosting Code-line-level Defect Prediction with Spectrum Information and Causality Analysis
- Boosting Path-Sensitive Value Flow Analysis Via Removal of Redundant Summaries
- Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention Inference
- COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge
- Calibration and Correctness of Language Models for Code
- Can an LLM Find Its Way Around a Spreadsheet?
- Chatgpt Inaccuracy Mitigation During Technical Report Understanding: Are we There Yet?
- Chatgpt-Based Test Generation for Refactoring Engines Enhanced by Feature Analysis on Examples
- Chord: Towards a Unified Detection of Blockchain Transaction Parallelism Bugs
- Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE
- Clozemaster: Fuzzing Rust Compiler by Harnessing Llms for Infilling Masked Real Programs
- Code Cloning in Solidity Smart Contracts: Prevalence, Evolution, and Impact on Development
- Code Comment Inconsistency Detection and Rectification Using a Large Language Model
- Code Today, Deadline Tomorrow: Procrastination Among Software Developers
- CodeImprove: Program Adaptation for Deep Code Models
- Combining Fine-Tuning and LLM-Based Agents for Intuitive Smart Contract Auditing with Justifications
- Coni: Detecting Database Connector Bugs via State-Aware Test Case Generation
- ConsCS: Effective and Efficient Verification of Circom Circuits
- Constrained LTL Specification Learning from Examples
- Context Conquers Parameters: Outperforming Proprietary Llm in Commit Message Generation
- Cooperative Software Verification via Dynamic Program Splitting
- Critical Variable State-Aware Directed Greybox Fuzzing
- DPFuzzer: Discovering Safety Critical Vulnerabilities for Drone Path Planners
- Datalog-Based Language-Agnostic Change Impact Analysis for Microservices
- Decictor: Towards Evaluating the Robustness of Decision-Making in Autonomous Driving Systems
- Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
- Decoding the Issue Resolution Process in Practice via Issue Report Analysis: a Case Study of Firefox
- Deep Learning-based Code Reviews: A Paradigm Shift or a Double-Edged Sword?
- Definition and Detection of Centralization Defects in Smart Contracts
- Demystifying and Detecting Cryptographic Defects in Ethereum Smart Contracts
- DesignRepair: Dual-Stream Design Guideline-Aware Frontend Repair with Large Language Models
- Dissecting Global Search: A Simple Yet Effective Method to Boost Individual Discrimination Testing and Repair
- Distilled Lifelong Self-Adaptation for Configurable Systems
- Diversity Drives Fairness: Ensemble of Higher Order Mutants for Intersectional Fairness of Machine Learning Software
- Dockerfile Flakiness: Characterization and Repair
- Does GenAI Make Usability Testing Obsolete?
- EP-Detector: Automatic Detection of Error-Prone Operation Anomalies in Android Applications
- Early Detection of Performance Regressions by Bridging Local Performance Data and Architectural Models
- EffBT: An Efficient Behavior Tree Reactive Synthesis and Execution Framework
- Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
- Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding
- Enhancing Fault Localization in Industrial Software Systems via Contrastive Learning
- Enhancing the Open Network: Definition and Automated Detection of Smart Contract Defects
- Evaluating Garbage Collection Performance Across Managed Language Runtimes
- Execution Trace Reconstruction Using Diffusion-Based Generative Models
- Exploring the Robustness of the Effect of EVO on Intention Valuation Through Replication
- Exposing the Hidden Layer: Software Repositories in the Service of Seo Manipulation
- FIXDRIVE: Automatically Repairing Autonomous Vehicle Driving Behaviour for $0.08 per Violation
- FairChecker: Detecting Fund-Stealing Bugs in DeFi Protocols via Fairness Validation
- FairSense: Long-Term Fairness Analysis of ML-Enabled Systems
- Fairness Testing Through Extreme Value Theory
- Fairquant: Certifying and Quantifying Fairness of Deep Neural Networks
- Famos: Fault Diagnosis for Microservice Systems Through Effective Multi-Modal Data Fusion
- Faster Configuration Performance Bug Testing with Neural Dual-Level Prioritization
- Feature-Driven End-to-End Test Generation
- Fidelity of Cloud Emulators: The Imitation Game of Testing Cloud-Based Software
- Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
- Fork State-Aware Differential Fuzzing for Blockchain Consensus Implementations
- Formally Verified Binary-Level Pointer Analysis
- Formally Verified Cloud-Scale Authorization
- From Bugs to Benefits: Improving User Stories by Leveraging Crowd Knowledge with CrUISE-AC
- Fuzzing MLIR Compilers with Custom Mutation Synthesis
- GARL: Genetic Algorithm-Augmented Reinforcement Learning to Detect Violations in Marker-Based Autonomous Landing Systems
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability Detectors
- GenC2Rust: Towards Generating Generic Rust Code from C
- Gpass: A Goal-Adaptive Neural Theorem Prover Based on Coq for Automated Formal Verification
- HIFI: Explaining and Mitigating Algorithmic Bias Through the Lens of Game-Theoretic Interactions
- Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code Search
- Hetrify: Efficient Verification of Heterogeneous Programs on RISC-V
- Hints Help Finding and Fixing Bugs Differently in Python and Text-Based Program Representations
- How Scientists Use Jupyter Notebooks: Goals, Quality Attributes, and Opportunities
- HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code Generation
- Hyperion: Unveiling DApp Inconsistencies Using LLM and Dataflow-Guided Symbolic Execution
- INTERTRANS: Leveraging Transitive Intermediate Translations to Enhance LLM-Based Code Translation
- IRFuzzer: Specialized Fuzzing for LLVM Backend Code Generation
- Improved Detection and Diagnosis of Faults in Deep Neural Networks Using Hierarchical and Explainable Classification
- Increasing the Effectiveness of Automatically Generated Tests by Improving Class Observability
- Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
- Instrumentation-Driven Evolution-Aware Runtime Verification
- Insvdf: Interface-State-Aware Virtual Device Fuzzing
- Intention is All you Need: Refining your Code from your Intention
- Interactive Cross-Language Pointer Analysis for Resolving Native Code in Java Programs
- Investigating the Impact of Interpersonal Challenges on Feeling Welcome in OSS
- Invivo Fuzzing by Amplifying Actual Executions
- Iterative Generation of Adversarial Example for Deep Code Models
- Janus: Detecting Rendering Bugs in Web Browsers via Visual Delta Consistency
- Knowledge-Enhanced Program Repair for Data Science Code
- LLM Assistance for Memory Safety
- LLM Based Input Space Partitioning Testing for Library APIs
- LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems
- LLM-Aided Automatic Modeling for Security Protocol Verification
- LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-Based Code Completion
- LWDIFF: an LLM-Assisted Differential Testing Framework for Webassembly Runtimes
- Large Language Models as Configuration Validators
- Large Language Models for Safe Minimization
- Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
- Leveraging Large Language Models to Detect NPM Malicious Packages
- Leveraging Propagated Infection to Crossfire Mutants
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented Generation
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language Models
- Licoeval: Evaluating LLMs on License Compliance in Code Generation
- Lightweight Concolic Testing via Path-Condition Synthesis for Deep Learning Libraries
- MAGIKA: AI-Powered Content-Type Detection
- MARQ: Engineering Mission-Critical AI-Based Software with Automated Result Quality Adaptation
- Measuring the Runtime Performance of C++ Code Written by Humans Using Github Copilot
- Metamorphic-Based Many-Objective Distillation of LLMs for Code-Related Tasks
- Mobile Application Coverage: The 30% Curse and Ways Forward
- Mock Deep Testing: Toward Separate Development of Data and Models for Deep Learning
- Model Editing for LLMs4Code: How Far are we?
- Module-Aware Context Sensitive Pointer Analysis
- Moye: A Wallbreaker for Monolithic Firmware
- NIODebugger: A Novel Approach to Repair Non-Idempotent-Outcome Tests with LLM-Based Agent
- Navigating the Testing of Evolving Deep Learning Systems: An Exploratory Interview Study
- Neurosymbolic Modular Refinement Type Inference
- No Harness, No Problem: Oracle-guided Harnessing for Auto-generating C API Fuzzing Harnesses
- On Prescription or Off Prescription? An Empirical Study of Community-Prescribed Security Configurations for Kubernetes
- On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
- PUPPY: Finding Performance Degradation Bugs in DBMSs via Limited-Optimization Plan Construction
- PacDroid: A Pointer-Analysis-Centric Framework for Security Vulnerabilities in Android Apps
- PairSmell: A Novel Perspective Inspecting Software Modular Structure
- Parametric Falsification of Many Probabilistic Requirements Under Flakiness
- Patch Synthesis for Property Repair of Deep Neural Networks
- Pattern-Based Generation and Adaptation of Quantum Workflows
- Planning a Large Language Model for Static Detection of Runtime Errors in Code Snippets
- Practical Object-Level Sanitizer with Aggregated Memory Access and Custom Allocator
- Preserving Privacy in Software Composition Analysis: A Study of Technical Solutions and Enhancements
- Prompt-to-SQL Injections in LLM-Integrated Web Applications: Risks and Defenses
- QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement Learning
- RLCoder: Reinforcement Learning for Repository-Level Code Completion
- ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation
- ROSA: Finding Backdoors with Fuzzing
- Rango: Adaptive Retrieval-Augmented Proving for Automated Software Verification
- Ranking Relevant Tests for Order-Dependent Flaky Tests
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?
- RediI: Test Infrastructure to Enable Deterministic Reproduction of Failures for Distributed Systems
- Reduce Dependence for Sound Concurrency Bug Prediction
- Relationship Status: "It's Complicated" Developer-Security Expert Dynamics in Scrum
- RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
- Repository-Level Graph Representation Learning for Enhanced Security Patch Detection
- Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models
- Rug: Turbo Llm for Rust Unit Test Generation
- RustAssistant: Using LLMs to Fix Compilation Errors in Rust Code
- SECRET: Towards Scalable and Efficient Code Retrieval via Segmented Deep Hashing
- SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents
- Sand: Decoupling Sanitization from Fuzzing for Low Overhead
- Scenario-Driven and Context-Aware Automated Accessibility Testing for Android Apps
- Search-Based LLMs for Code Optimization
- SeeAction: Towards Reverse Engineering How-What-Where of HCI Actions from Screencasts for UI Automation
- Selecting Initial Seeds for Better JVM Fuzzing
- Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
- Similar but Patched Code Considered Harmful: The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove Them
- Smartreco: Detecting Read-Only Reentrancy via Fine-Grained Cross-DApp Analysis
- Software Model Evolution with Large Language Models: Experiments on Simulated, Public, and Industrial Datasets
- Source Code Summarization in the Era of Large Language Models
- SpecGen: Automated Generation of Formal Program Specifications via Large Language Models
- SpecRover: Code Intent Extraction via LLMs
- Static Analysis of Remote Procedure Call in Java Programs
- Studying Programmers Without Programming: Investigating Expertise Using Resting State fMRI
- Synthesizing Document Database Queries Using Collection Abstractions
- TIGER: A Generating-Then-Ranking Framework for Practical Python Type Inference
- TOGLL: Correct and Strong Test Oracle Generation with LLMS
- TacDroid: Detection of Illicit Apps Through Hybrid Analysis of UI-Based Transition Graphs
- Template-Guided Program Repair in the Era of Large Language Models
- Test Intention Guided LLM-Based Unit Test Generation
- Testing and Understanding Deviation Behaviors in FHE-Hardened Machine Learning Models
- Thanos: DBMS Bug Detection via Storage Engine Rotation Based Differential Testing
- The Design Smells Breaking the Boundary between Android Variants and AOSP
- The Fact Selection Problem in LLM-Based Program Repair
- The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
- The Product Beyond the Model - An Empirical Study of Repositories of Open-Source ML Products
- The Same Only Different: On Information Modality for Configuration Performance Analysis
- The Seeds of the Future Sprout from History: Fuzzing for Unveiling Vulnerabilities in Prospective Deep-Learning Libraries
- Tiver: Identifying Adaptive Versions of C/C++ Third-Party Open-Source Components Using a Code Clustering Technique
- Topseed: Learning Seed Selection Strategies for Symbolic Execution from Scratch
- Toward a Better Understanding of Probabilistic Delta Debugging
- Towards Better Answers: Automated Stack Overflow Post Updating
- Towards High-Strength Combinatorial Interaction Testing for Highly Configurable Software Systems
- Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
- Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language Models
- TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance
- TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code
- Treefix: Enabling Execution with a Tree of Prefixes
- Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications
- Tumbling Down the Rabbit Hole: How do Assisting Exploration Strategies Facilitate Grey-Box Fuzzing?
- UML is Back. Or is it? Investigating the Past, Present, and Future of UML in Open Source Software
- Unavoidable Boundary Conditions: a Control Perspective on Goal Conflicts
- Understanding Architectural Complexity, Maintenance Burden, and Developer Sentiment -A Large-Scale Study
- Understanding Compiler Bugs in Real Development
- Understanding and Detecting Peer Dependency Resolving Loop in npm Ecosystem
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
- Understanding the Response to Open-Source Dependency Abandonment in the npm Ecosystem
- Unleashing the True Potential of Semantic-Based Log Parsing with Pre-Trained Language Models
- Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar
- Unveiling the Energy Vampires: A Methodology for Debugging Software Energy Consumption
- User Personas Improve Social Sustainability by Encouraging Software Developers to Deprioritize Antisocial Features
- Vulnerability Detection with Code Language Models: How Far are We?
- WDD: Weighted Delta Debugging
- Weakly-Supervised Log-Based Anomaly Detection with Inexact Labels via Multi-Instance Learning
- What Guides Our Choices? Modeling Developers' Trust and Behavioral Intentions Towards Genai
- What You See is What You Get: Attention-Based Self-Guided Automatic Unit Test Generation
- When Quantum Meets Classical: Characterizing Hybrid Quantum-Classical Issues Discussed in Developer Forums
- Who's Pushing the Code? An Exploration of GitHub Impersonation
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language Models
- exLong: Generating Exceptional Behavior Tests with Large Language Models