350 papers · page 1 of 18
Nadav Amit
While huge pages can dramatically reduce address translation overhead, their use for executable code remains limited by structural barriers in binary formats, page cache management, and loader mechanisms. Existing solutions copy code into anonymous memory, abandoning file-backed …
Anthony Arnold, Mark Marron
Garbage Collectors (GCs) are a critical component of a modern application stack. Long pauses, large memory consumption, and high CPU usage can unexpectedly occur with certain workloads or series of events. These behaviors can make systems unresponsive, make it impossible to run t…
Soham Bagchi, Sanya Srivastava, Reese Levine, Tyler Sorensen, Ryan Stutsman, Vijay Nagarajan
Modern heterogeneous processors like the NVIDIA Grace-Hopper Superchip tightly integrate CPU and GPU cores across a cache-coherent interconnect, with an implicit assumption that independently compiled CPU and GPU code can safely interact via shared memory. Yet the memory consiste…
Edoardo D'Alessio, Mohamed Husain Noor Mohamed, Xiaoguang Wang, Binoy Ravindran
Recent Linux memory-management interfaces make it practical to revisit distributed shared memory (DSM) as a deployable runtime substrate for conventional multithreaded software. We present Stretch, a userspace fault-driven page-granularity DSM runtime that combines userfaultfd-ba…
Kai Feng, Huanting Wang, Jeremy Singer, Zheng Wang
Memory safety is a critical issue in embedded systems. Although high-level languages like MicroPython simplify IoT development, their C-based runtimes remain vulnerable to memory errors triggered by Python code or native extensions. The CHERI (Capability Hardware Enhanced RISC In…
Nicolas van Kempen, Emery D. Berger
Programmers using native languages such as C, C++, or Rust can implement custom memory allocation strategies to improve execution time. In their paper titled "Reconsidering Custom Memory Allocation" almost 25 years ago, Berger et al. showed that while per-class allocators provide…
Hayley Patton, Stephen M. Blackburn
Offset-Vector Compaction (OVC) algorithms, such as the Compressor, are among the most widely deployed garbage collectors today, yet they have remained largely unexplored by the literature since the earliest algorithms were published two decades ago. Although an implementation in …
Yunqi Shen, Dimitrios Nikolopoulos
GPU unified memory simplifies programming by automatically migrating pages between CPUs and GPUs, but page faults trigger migrations with hundreds of microseconds to millisecond-scale latency, stalling thousands of threads. We target this bottleneck with a page prefetching framew…
Bijan Tabatabai, Eishan Mirakhur, Ravi Shankar Jonnalagadda, Vinicius Petrucci, Rohit Sehgal, Jus Singh, Michael M. Swift
CXL memory devices increase the memory capacity and bandwidth available to a server, at the cost of higher access latency. Prior research focused on how to make use of the expanded memory capacity provided by CXL while minimizing the impact of its higher access latency. However, …
Luís Eduardo de Souza Amorim, Yi Lin, Stephen M. Blackburn, Diogo Netto, Gabriel Baraldi, Nathan Daly, Antony L. Hosking, Kiran Pamnany + 1 more
Julia is a dynamically-typed garbage-collected language designed for high performance. Julia has a non-moving tracing collector, which, while performant, is subject to the same unavoidable fragmentation and lack of locality as all other non-moving collectors. In this work, we ref…
Stephen Dolan
The effectiveness of generational garbage collection is usually explained through the generational hypothesis, that “most objects die young”.
Despite its simplicity, the generational hypothesis leaves some things to be desired: it is not obvious how it can be measured as a prope…
Parth Gangar, Ashish Panwar, K. Gopinath
Recent processors rely on huge pages to reduce the cost of virtual-to-physical address translation. However, huge pages are notorious for creating memory bloat – a phenomenon wherein the OS ends up allocating more physical memory to an application than its actual requirement. Thi…
Frédéric Lahaie-Bertrand, Léonard Oest O'Leary, Olivier Melançon, Marc Feeley, Stefan Monnier
Reclaiming cyclic garbage has been a long-standing challenge in automatic memory management. Common approaches to this problem often involve extending reference counting with an asynchronous background task to reclaim cycles. While this ensures that cycles are eventually collecte…
Ryu Morimoto, Kazuki Ichinose, Tomoharu Ugawa
Processing-in-memory (PIM) is a promising approach to overcome the performance bottleneck caused by the gap between CPU speed and memory speed, known as the memory wall problem. The UPMEM PIM-enabled memory is the first commercialized general-purpose PIM accelerator, to which the…
Sai Dhawal Phaye, Gregory J. Duck, Roland H. C. Yap, Trevor E. Carlson
Memory errors continue to be a critical concern for programs written in low-level programming languages such as C and C++. Many different memory error defenses have been proposed, each with varying trade-offs in terms of overhead, compatibility, and attack resistance. Some defens…
Yun Joon Soh, Sihang Liu, Steven Swanson, Jishen Zhao
Writing crash-consistent programs for memory-semantic storage such as persistent memory (PMEM) is error-prone and cumbersome. Programmers must implement both the main logic and the recovery logic to ensure data consistency after unexpected power failures. Prior work has reduced t…
Sathvik Swaminathan, Sandeep Kumar, Aravinda Prasad, Sreenivas Subramoney
Deep neural networks (DNNs) are one of the popular models for learning relationships between complex data. Training a DNN model is a compute- and memory-intensive operation. The size of modern DNN models spans into the terabyte region, requiring multiple accelerators to train -- …
Kunshan Wang, Stephen M. Blackburn, Peter Zhu, Matthew Valentine-House
Ruby is a dynamic programming language that was first released in 1995 and remains heavily used today. Ruby underpins Ruby on Rails, one of the most widely deployed web application frameworks. The scale at which Rails is deployed has placed increasing pressure on the underlying C…
Huanting Wang, Dejice Jacob, David Kelly, Yehia Elkhatib, Jeremy Singer, Zheng Wang
Large language models (LLMs) hold great promise for automating software vulnerability detection and repair, but ensuring their correctness remains a challenge. While recent work has developed benchmarks for evaluating LLMs in bug detection and repair, existing studies rely on han…
Tomer Cory, Erez Petrank
Compaction algorithms alleviate fragmentation by relocating heap objects to compacted areas. A full heap compaction eliminates fragmentation by compacting all objects into the lower addresses of the heap. Following a marking phase that marks all reachable (live) objects, the comp…