FSE 2025
132 papers
- 10 Years Later: Revisiting How Developers Search for Code
- A Causal Learning Framework for Enhancing Robustness of Source Code Models
- A Comprehensive Study of Bug-Fix Patterns in Autonomous Driving Systems
- A Knowledge Enhanced Large Language Model for Bug Localization
- A Mixed-Methods Study of Model-Based GUI Testing in Real-World Industrial Settings
- A New Approach to Evaluating Nullability Inference Tools
- Adaptive Random Testing with Q-grams: The Illusion Comes True
- Alert Summarization for Online Service Systems by Validating Propagation Paths of Faults
- AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation
- An Adaptive Language-Agnostic Pruning Method for Greener Language Models for Code
- An Empirical Study of Bugs in Data Visualization Libraries
- An Empirical Study of Code Clones from Commercial AI Code Generators
- An Empirical Study of Suppressed Static Analysis Warnings
- An Empirical Study on Release-Wise Refactoring Patterns
- Automated Extraction and Analysis of Developer's Rationale in Open Source Software
- Automated Recognition of Buggy Behaviors from Mobile Bug Reports
- Automated Soap Opera Testing Directed by LLMs and Scenario Knowledge: Feasibility, Challenges, and Road Ahead
- Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
- Automated Unit Test Refactoring
- Automated and Accurate Token Transfer Identification and Its Applications in Cryptocurrency Security
- Automatically Detecting Numerical Instability in Machine Learning Applications via Soft Assertions
- Automatically Fixing Dependency Breaking Changes
- Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models
- Beyond PEFT: Layer-Wise Optimization for More Effective and Efficient Large Code Model Tuning
- Blended Analysis for Predictive Execution
- Bridging Operator Semantic Inconsistencies: A Source-Level Cross-Framework Model Conversion Approach
- CAShift: Benchmarking Log-Based Cloud Attack Detection under Normality Shift
- CKTyper: Enhancing Type Inference for Java Code Snippets by Leveraging Crowdsourcing Knowledge in Stack Overflow
- COFFE: A Code Efficiency Benchmark for Code Generation
- CRISPE: Semantic-Guided Execution Planning and Dynamic Reasoning for Enhancing Code Coverage Prediction
- CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
- Calibration of Large Language Models on Code Summarization
- ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
- ChatDBG: Augmenting Debugging with Large Language Models
- Clone Detection for Smart Contracts: How Far Are We?
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLM
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming Tasks
- Core Developer Turnover in the Rust Package Ecosystem: Prevalence, Impact, and Awareness
- CoverUp: Effective High Coverage Test Generation for Python
- Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-Learning
- De-duplicating Silent Compiler Bugs via Deep Semantic Representation
- DeclarUI: Bridging Design and Development with Automated Declarative UI Code Generation
- Demystifying LLM-Based Software Engineering Agents
- Demystifying Memorization in LLM-Based Program Repair via a General Hypothesis Testing Framework
- Detecting Metadata-Related Bugs in Enterprise Applications
- Detecting Smart Contract State-Inconsistency Bugs via Flow Divergence and Multiplex Symbolic Execution
- Detecting and Handling WoT Violations by Learning Physical Interactions from Device Logs
- Detecting and Reducing the Factual Hallucinations of Large Language Models with Metamorphic Testing
- DiSCo: Towards Decompiling EVM Bytecode to Source Code using Large Language Models
- Directed Testing in MLIR: Unleashing Its Potential by Overcoming the Limitations of Random Fuzzing
- Dissecting Real-World Cross-Language Bugs
- Divide-and-Conquer: Generating UI Code from Screenshots
- Doc2OracLL: Investigating the Impact of Documentation on LLM-Based Test Oracle Generation
- DuoReduce: Bug Isolation for Multi-layer Extensible Compilation
- DyLin: A Dynamic Linter for Python
- Dynamic Taint Tracking for Modern Java Virtual Machines
- Element-Based Automated DNN Repair with Fine-Tuned Masked Language Model
- Eliminating Backdoors in Neural Code Models for Secure Code Understanding
- Empirically Evaluating the Impact of Object-Centric Breakpoints on the Debugging of Object-Oriented Programs
- Enhancing Web Accessibility: Automated Detection of Issues with Generative AI
- Error Delayed Is Not Error Handled: Understanding and Fixing Propagated Error-Handling Bugs
- Expressing and Checking Statistical Assumptions
- Gleipner: A Benchmark for Gadget Chain Detection in Java Deserialization Vulnerabilities
- Hallucination Detection in Large Language Models with Metamorphic Relations
- Has My Code Been Stolen for Model Training? A Naturalness Based Approach to Code Contamination Detection
- HornBro: Homotopy-Like Method for Automated Quantum Program Repair
- How Do Programming Students Use Generative AI?
- IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
- Impact of Request Formats on Effort Estimation: Are LLMs Different Than Humans?
- Incorporating Verification Standards for Security Requirements Generation from Functional Specifications
- Integrating Large Language Models and Reinforcement Learning for Non-linear Reasoning
- It's Acting Odd! Exploring Equivocal Behaviors of Goodware
- LLM-Based Method Name Suggestion with Automatically Generated Context-Rich Prompts
- LLMDroid: Enhancing Automated Mobile App GUI Testing Coverage with Large Language Model Guidance
- Large Language Models for In-File Vulnerability Localization Can Be "Lost in the End"
- Less Is More: On the Importance of Data Quality for Unit Test Generation
- Liberating Libraries through Automated Fuzz Driver Generation: Striking a Balance without Consumer Code
- LlamaRestTest: Effective REST API Testing with Small Language Models
- LookAhead: Preventing DeFi Attacks via Unveiling Adversarial Contracts
- Medusa: A Framework for Collaborative Development of Foundation Models with Automated Parameter Ownership Assignment
- MendelFuzz: The Return of the Deterministic Stage
- MiSum: Multi-modality Heterogeneous Code Graph Learning for Multi-intent Binary Code Summarization
- Mitigating Emergent Malware Label Noise in DNN-Based Android Malware Detection
- Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
- Mystique: Automated Vulnerability Patch Porting with Semantic and Syntactic-Enhanced LLM
- NLP Libraries, Energy Consumption and Runtime: An Empirical Study
- No More Labelled Examples? An Unsupervised Log Parser with LLMs
- On the Characteristics and Impacts of Protestware Libraries
- On the Unnecessary Complexity of Names in X.509 and Their Impact on Implementations
- On-Demand Scenario Generation for Testing Automated Driving Systems
- One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)
- PDCAT: Preference-Driven Compiler Auto-tuning
- Pinning Is Futile: You Need More Than Local Dependency Versioning to Defend against Supply Chain Attacks
- Prompts Are Programs Too! Understanding How Developers Build Software Containing Prompts
- QSF: Multi-objective Optimization Based Efficient Solving for Floating-Point Constraints
- ROSCallBaX: Statically Detecting Inconsistencies in Callback Function Setup of Robotic Systems
- Ransomware Detection through Temporal Correlation between Encryption and I/O Behavior
- RePurr: Automated Repair of Block-Based Learners' Programs
- Recasting Type Hints from WebAssembly Contracts
- RegTrieve: Reducing System-Level Regression Errors for Machine Learning Systems via Retrieval-Enhanced Ensemble
- ReproCopilot: LLM-Driven Failure Reproduction with Dynamic Refinement
- Revisiting Optimization-Resilience Claims in Binary Diffing Tools: Insights from LLVM Peephole Optimization Analysis
- Revolutionizing Newcomers' Onboarding Process in OSS Communities: The Future AI Mentor
- Scene Flow Specifications: Encoding and Monitoring Rich Temporal Safety Properties of Autonomous Systems
- Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software
- SemBIC: Semantic-Aware Identification of Bug-Inducing Commits
- Smaller but Better: Self-Paced Knowledge Distillation for Lightweight yet Effective LCMs
- Smart Contract Fuzzing Towards Profitable Vulnerabilities
- SmartNote: An LLM-Powered, Personalised Release Note Generator That Just Works
- SmartShot: Hunt Hidden Vulnerabilities in Smart Contracts using Mutable Snapshots
- Software Fairness Dilemma: Is Bias Mitigation a Zero-Sum Game?
- Standing on the Shoulders of Giants: Bug-Aware Automated GUI Testing via Retrieval Augmentation
- Statement-Level Adversarial Attack on Vulnerability Detection Models via Out-of-Distribution Features
- Teaching AI the 'Why' and 'How' of Software Vulnerability Fixes
- The Landscape of Toxicity: An Empirical Investigation of Toxicity on GitHub
- The Struggles of LLMs in Cross-Lingual Code Clone Detection
- Today's Cat Is Tomorrow's Dog: Accounting for Time-Based Changes in the Labels of ML Vulnerability Detection Approaches
- Towards Diverse Program Transformations for Program Simplification
- Towards Understanding Docker Build Faults in Practice: Symptoms, Root Causes, and Fix Patterns
- Towards Understanding Fine-Grained Programming Mistakes and Fixing Patterns in Data Science
- Towards Understanding Performance Bugs in Popular Data Science Libraries
- TracePicker: Optimization-Based Trace Sampling for Microservice-Based Systems
- Understanding Debugging as Episodes: A Case Study on Performance Bugs in Configurable Software Systems
- Understanding Industry Perspectives of Static Application Security Testing (SAST) Evaluation
- Understanding and Characterizing Mock Assertions in Unit Tests
- UnitCon: Synthesizing Targeted Unit Tests for Java Runtime Exceptions
- Unlocking Optimal ORM Database Designs: Accelerated Tradeoff Analysis with Transformers
- VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
- VulPA: Detecting Semantically Recurring Vulnerabilities with Multi-object Typestate Analysis
- Who Will Stop Contributing to OSS Projects? Predicting Company Turnover Based on Initial Behavior
- Why the Proof Fails in Different Versions of Theorem Provers: An Empirical Study of Compatibility Issues in Isabelle
- Zero-Shot Cross-Domain Code Search without Fine-Tuning