FSE 2026
212 papers
- A Grounded Theory of Debugging in Professional Software Engineering Practice
- A Tuple-Oriented Sampling Method for Generating Small Pairwise Covering Arrays in Configurable Software Systems
- A Wily Hare Has Three Havens: Combating Programmable Logic Controller Attacks via Virtualization Redundancy
- ACME: Automated Clause Mapping Engine for Testing Emerging Database Systems
- Accelerating Policy Synthesis in Large-Scale MDPs via Hierarchical Adaptive Refinement
- AccessDroid: Detecting Screen Reader Accessibility Issues in Android Applications via Semantics Trees
- AccessRefinery: Fast Mining Concise Access Control Intents on Public Cloud
- Active Learning of Symbolic Automata for Reactive Programs via Dynamic Symbolic Mapper
- AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
- Adaptive Mutation Scheduling with Deep Reinforcement Learning for Smart Contract Fuzzing
- AgentBound: Securing Execution Boundaries of AI Agents
- Agentic Verification of Software Systems
- Aligning with Human Coding Preferences for Improving Code Generation
- An Empirical Study of Fuzz Harness Degradation
- Automated Detection of Configuration-Specific Security Vulnerabilities via Patch Analysis
- Automated Knowledge-Aware Test Reuse
- Automated Repair of Requirements for Cyber-Physical Systems in Simulink Requirements Tables
- Automated Repair of TEE Partitioning Issues via DSL-Guided and LLM-Assisted Patching
- Automating Dockerfile Refactoring to Multi-stage Builds
- BackportBench: A Multilingual Benchmark for Automated Patch Backporting
- Balancing Latency and Accuracy of Code Completion via Local-Cloud Model Cascading
- Bash-Commenter: Leveraging Syntax-Aware Preference Optimization to Reinforce Large Language Model for Bash Code Comment Generation
- Behind Defective Mobile AR Apps: Studying Reviews and Bugs of Android AR Software with Comparison to Prior Bug Studies
- Beyond Language Boundaries: Uncovering Programming Language Families for Code Language Models
- Binvariants: Enhancing Fuzzing of Closed-Source Binary Executables via Register-Level Likely Invariants
- Boosting LLMs for Mutation Generation
- Break to Adapt: Knowledge-Based Updates of Breaking Dependencies in JavaScript
- Bringing Managed Language Support to WebAssembly with External Library Linking
- Building Software by Rolling the Dice: A Qualitative Study of Vibe Coding
- CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test Generation
- Can Old Tests Do New Tricks for Resolving SWE Issues?
- Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models
- Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing
- Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-Breaking
- CertiCoder: Towards MISRA-Compliant C Code Generation with LLMs
- ChainDelta: Automatic Patch-Based Exploit Generation for Ethereum with Fuzzing Agents
- Characterizing Trust Boundary Vulnerabilities in TEE Container Systems: An Empirical Study
- Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel
- Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code Generation
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- Co-evolution of Types and Dependencies: Towards Repository-Level Type Inference for Python Code
- CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
- Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs during Code Adaptation
- Comment Traps: How Defective Commented-Out Code Augment Defects in AI-Assisted Code Generation
- Compiling Code LLMs into Lightweight Executables
- Corrigendum: SmartNote: An LLM-Powered, Personalised Release Note Generator That Just Works
- Cost-Effective Testing of MPC Compilers
- Cross-Refactoring-Type Test Program Migration for Refactoring Engines
- CrossFit: Demystifying VM Callback Bugs in Interpreters
- CrypFormBench: Benchmarking Formal Analysis Capability of Large Language Models for Cryptographic Schemes
- CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries
- DECODE: Dynamic Exploration for Constraint-Guided Vulnerability Discovery in Deep Learning Operators
- Debugging Engine Enhanced by Prior Knowledge: Can We Teach LLM How to Debug?
- Denoising Fault Localization with Test Line Proximity
- Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
- Detecting Bugs in Rust Compiler Fix Suggestions via Constraint-Violation-Guided Mutation
- Detecting Code-Comment Inconsistencies in Smart Contracts by Combining LLM and Program Analysis
- DiverFPS: Generating Diverse Solutions for Floating-Point SMT Formulas
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond
- DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
- DualCodeDetect: Zero-Shot LLM-Generated Code Detection via Dual-Channel Perturbation
- EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code Generation
- Eidolon: Perform Noise-Aware Fuzzing on FHE Libraries via Equivalence Expression Transformation
- Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts
- Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis
- Evaluating LLM-Based Regression Test Generation
- Evaluating Risk and Confidence in Performance Bounds of Configuration Sampling Strategies
- Event-B Agent: Towards LLM Agent for Formal Model Synthesis and Repair
- EventADL: Open-Box Anomaly Detection and Localization Framework for Events in Cloud-Based Service Systems
- Exorcist: Enabling Atomic-Level Runtime Detection of Spectre Attacks using Precise Event Based Sampling
- ExpeRepair: Dual-Memory Enhanced LLM-Based Repository-Level Program Repair
- Failing with Purpose: Dangling Coverage-Guided Negative Test Generation from a Mechanized P4 Type System
- Failure-Based Testing for Deep Reinforcement Learning Agents
- Fairness Testing of Large Language Models in Role-Playing
- Feature Slice Matching for Precise Bug Detection
- Flash: Query-Efficient Black-Box Static Malware Evasion through Transferable GAN-Guided Modification Sequences
- Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-Preserving Transformations
- From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing
- From Specifications to Implementation in the Gen-AI Era: Lessons from a Project-Based Software Engineering Course
- From Suspicious Signals to Crashes: Guiding Bug-Driven GUI Testing via Code-Inspired Tracing
- GAER: Graph Auto-encoders for Unsupervised Software Architecture Recovery
- GPU-Accelerated Flow-Sensitive Pointer Analysis for C/C++ Programs
- GREClue: Failure Indexing with Graph-Based Failure Representation and Entropy-Based Deep Clustering
- GUIMigrator: Semantics-Preserving Transpilation from Android XML to Compose and SwiftUI
- GadgetHunter: Region-Based Neuro-symbolic Detection of Java Deserialization Vulnerabilities
- Generalizing Test Cases for Comprehensive Test Scenario Coverage
- GraphLocator: Graph-Guided Causal Reasoning for Issue Localization
- GraphQLify: Automated and Type Safety-Preserving GraphQL API Adoption
- Hallucinations in LLM-Based Code Summarization: Unveiling, Detection, and Mitigation
- How Do Developers Interact with AI? An Exploratory Study on Modeling Developer Programming Behavior
- How Low Can You Go? The Data-Light SE Challenge
- Improving Data Leakage Detection in Machine Learning Notebooks through Static Slicing and Structured LLM Prompts
- In Bugs We Trust? On Measuring the Randomness of a Fuzzer Benchmarking Outcome
- In Line with Context: Repository-Level Code Generation via Context Inlining
- InDe-LLM: Defending against Jailbreak Attacks in LLM-Powered Systems via Intention Disentangling
- Influence-Aware Bayesian-Inspired Token Reweighting for Improved Code Generation
- IntentTester: Intent-Driven Multi-agent Framework for Cross-Library Test Migration
- Interrogation Testing of CHC Solvers
- It Takes Two: Option-Aware Directed Greybox Fuzzing for Vulnerability PoC Generation
- JavaScript Pointer Analysis with Adaptive Heap Abstraction
- Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study
- LLM-Assisted Input-Requirement-Aware Differential Testing of Array Programming Frameworks
- Large Language Models for Opaque Predicate Resolution: A Universal Control Flow Deobfuscation Framework
- LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery
- LoCaL: Countering Surface Bias in Code Evaluation Metrics
- Look Before You Leap: Context-Sensitive GUI Grounding for Boosting Automated Extended Reality (XR) Testing
- MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling Analysis
- MetaRCA: A Generalizable Root Cause Analysis Framework for Cloud-Native Systems Powered by Meta Causal Knowledge
- Mining Long Tail Bugs: Identifying Rare and Overlooked Issues in Code
- Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering
- Mitigating the Risk of Defects and Improving Knowledge Distribution with Code Reviewer Recommenders
- Multi-LLM Persona Generation for Virtual Focus Groups in Software Engineering: A Controlled, Multi-domain Study of Emotional Requirements Elicitation
- NESA: Relational Neuro-Symbolic Static Program Analysis
- Natural Language-Focused Software Engineering via Code-Documentation Equivalence
- Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?
- Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks
- OCPPuzz: Specification-Driven Fuzzing of Charging Station Management Systems with Large Language Model
- OdoTest: An Automated Testing Approach for Odometry Systems
- Odyssey: Hunting Smart Contract Vulnerabilities with Fine-Grained State Modeling and Exploration
- On the Road to Personalized Code Intelligence: Portraiting and Assisting Developers Based on Their In-IDE Behaviors
- One Size Does Fit All: Exploring Model Fusion for Software Engineering Tasks
- One Size Does Not Fit All: Revisiting Code Context Engineering for Repository-Level Code Generation
- PROGnosticator: Testing Source-to-Source Code Translators via Construct-Oriented Fuzzing
- Phantom Rendering Detection: Identifying and Analyzing Unnecessary UI Computations
- Pig: Leveraging Large Language Models for Python Library Migrations
- PlayCoder: Making LLM-Generated GUI Code Playable
- PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages
- Precondition Synthesis for Deep Neural Networks with Statistical Guarantees
- Project-Level C-to-Rust Translation via Pointer Knowledge Graphs
- ProofFusion: Improving Neural Theorem Proving via Adaptive Retrieval-Augmented Reasoning
- Property Refinement in Linear Temporal Logic: Formal Semantics and Algorithms for Software Verification
- Protocol Reverse Engineering via Deep Transfer Learning
- PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion Models
- QuanForge: A Mutation Testing Framework for Quantum Neural Networks
- RAT: Retrieval-Augmented Testing of Certificate Revocation List Parsers in TLS Implementations
- ReDef: Do Code Language Models Truly Understand Code Changes for Just-in-Time Software Defect Prediction?
- ReFLAIR: Detecting Responsive Layout Reflow Issues using Multimodal Generative AI
- ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
- Recommending Usability Improvements with Multimodal Large Language Models
- Red Teaming LLMs via Linguistic-Aware Fuzzing
- Reducing Cost of LLM Agents with Trajectory Reduction
- Reducing Coverage-Equivalent Inputs in Grammar-Based Fuzzing by Avoiding Recurrent Rule Sequences
- Reducing the TCB of SGX-Oriented LibOSes at Runtime
- RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware Benchmark
- Revealing Regressions: A Comparative Study of State-Capture Strategies in Validating Program Behavior
- Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
- SBridge: Identifying Source-to-Binary Function Similarity via Cross-Domain Control Block Matching
- SQLiFuzz: Uncovering SQL Injection in Any Web Applications
- SWE Data Construction, Automatically!
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation
- Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
- ScanCoder: Leveraging Human Attention Patterns to Enhance LLMs for Code
- Semantics-Guided Control-Flow Reconstruction for Firmware Binaries via Static Analysis
- Small Is Beautiful: A Practical and Efficient Log Parsing Framework
- SmarTrim: Symbolic Execution for Smart Contracts Powered by Redundant Transaction-Sequence Pruning
- SmartCoder-R1: Towards Secure and Explainable Smart Contract Generation with Security-Aware Group Relative Policy Optimization
- SmartDispatch: Dynamic Substitution of NumPy-Style APIs on Heterogeneous CPU-GPU Systems
- SmartIFSyn: Automated Information Flow Security Policy Synthesis for Smart Contracts
- SnakeCharmer: Automatic Fuzzing Harness Generation for Pure and Hybrid Python Libraries
- Sound Termination and Non-termination Analysis of C Programs with Bit-Precise Bounded Semantics and Advanced Constructs
- SpecWeaver: End-to-End HTTP API Specification Inference across Multi-layer Routing in Production Web Services
- Spectrum-Based Failure Attribution for Multi-agent Systems
- Speculate: Generating REST API Specifications using LLMs
- StepFly: Agentic Troubleshooting Guide Automation for Incident Diagnosis
- Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
- Structure-Aware Delta Debugging with Geometric-Information Weights
- SwarmBox: A Plug-and-Play Drone Swarm Framework for Streamlined Development and Comprehensive Analysis
- TLR: Codebase-Level C Memory Management Error Repair with Large Language Models
- TORAI: Multi-source Root Cause Analysis for Blind Spots in Microservice Service Call Graph
- TSGuard: Automated User-Centric Incident Diagnosis for AI Workloads in the Cloud
- TUSR: A Test Unit-Based Framework for Repairing Obsolete GUI Test Scripts
- TestTailor: Generating High-Coverage Tests via Path-Proximal Tests with LLMs
- The Effect of Complexity and Provenance on Code Review Decisions: Evidence from a Controlled Experiment
- Thought Is All You Need: Smart Contract Vulnerability Detection with Thought-Augmented Large Language Model
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability Detection
- Towards Automated Crowdsourced Testing via Personified-LLM
- Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
- Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs
- Towards the Localization of Multi-Root-Cause Failures in Microservice Systems: An Active Intervention Framework
- ToxiShield: Promoting Inclusive Developer Communication through Real-Time Toxicity Filtering
- TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment
- TransLibEval: Demystify Large Language Models' Capability in Third-Party Library-Targeted Code Translation
- Two-Level Adaptation for Budget-Constrained Continuous Dynamic Dependence Analysis
- TypePro: Boosting LLM-Based Type Inference via Inter-Procedural Slicing
- UNICS: Multilingual Code Search via Unified Pseudocode and Contrastive Transfer Learning
- Uncovering Similar but Different Packages in PyPI and Potential Security Threats
- Understanding Binary Code Similarity for Real-World Vulnerability Detection: A Large-Scale Empirical Study
- Understanding Code Similarity across Instruction Set Architectures: An Empirical Study
- Understanding Performance Problems in CUDA Programs
- Understanding and Predicting Accepted Code Suggestions in AI-Assisted Programming
- Understanding the Limitations of C/C++ Binary Third-Party Library Detection Tool: An Empirical Study at Scale
- Understanding, Detecting, and Repairing Real-World In-Context-Learning-Based Text-to-SQL Errors
- Unfulfilled Promises: LLM-Based Detection of OS Compatibility Issues in Infrastructure as Code
- Unleashing HPC Application Performance through Software Deployment: A Joint Model of Software Parallelism and Co-location
- Unveiling AI-Driven Web Applications: Insights into Characteristics, Functionality, and Compliance
- Unveiling the Fragility of Binary Code Similarity Detection via Targeted Attacks with Model Explanations
- V2E: Validating Smart Contract Vulnerabilities through Profit-Driven Exploit Generation and Execution
- Validating LLM-Generated SQL Queries through Metamorphic Prompting
- Verifying Smart Contract Security against Re-entrancy Attacks through Relational Value Analysis
- Verifying Structural Robustness of Deep Neural Network
- VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement Framework
- ViBR: Automated Bug Replay from Video-Based Reports using Vision-Language Models
- VisionScratch: LLM-Based Automated Feedback Generation using Code-Produced Videos for Scratch Programs
- VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications
- VulKey: Automated Vulnerability Repair Guided by Domain-Specific Repair Patterns
- WalleTruth: Visual-Oriented Software Testing for Web3 Wallet Browser Extensions
- WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements
- When Shared Worlds Break: Demystifying Defects in Multi-user Extended Reality Software Systems
- iCoRe: An Iterative Correlation-Aware Retriever for Bug Reproduction Test Generation
- pPatch: Automated Vulnerability Unpatching