350 papers · page 2 of 18
Nathaniel Wesley Filardo, Matthew J. Parkinson
Modern, high-performance memory allocators must scale to a wide array of uses, including producer-consumer workloads. In such workloads, objects are allocated by one thread and deallocated by another, which we call remote deallocations. These remote deallocations lead to contenti…
Taekyung Heo, Seunghyo Kang, Sanghyeon Lee, Soojin Hwang, Joongun Park, Jaehyuk Huh
Although recent studies have been improving the performance of RDMA-based memory disaggregation systems, their security aspect has not been thoroughly investigated. For secure disaggregated memory, the memory-providing node must protect its memory from memory-requesting nodes, an…
Akifumi Imanishi, Zijian Xu
Neural network training requires immense GPU memory, and memory optimization methods such as recomputation are being actively researched. Recent improvements in recomputation have reduced more than 90 % of the peak allocated size. However, because it produces complex irregular al…
Akira Inoue, Tomoharu Ugawa, Shigeru Chiba
This paper presents a managed memory system for micro controllers with only a small amount of memory but with NOR flash memory. This system is targeted at a device such as Raspberry Pi Pico, which is equipped with ARM Coretex M0+, on-chip 264KB SRAM, and 2MB flash memory. To exte…
Sebastian Jordan-Montaño, Guillermo Polito, Stéphane Ducasse, Pablo Tesone
Using object lifetime information enables performance improvement through memory optimizations such as pretenuring and tuning garbage collector parameters. However, profiling object lifetimes is nontrivial and often requires a specialized virtual machine to instrument object allo…
Chaitanya S. Koparkar, Vidush Singhal, Aditya Gupta, Mike Rainey, Michael Vollmer, Artem Pelenitsyn, Sam Tobin-Hochstadt, Milind Kulkarni + 1 more
Over the years, traditional tracing garbage collectors have accumulated assumptions that may not hold in new language designs. For instance, we usually assume that run-time objects do not hold addressable sub-parts and have a size of at least one pointer. These fail in systems st…
Matthew J. Parkinson, Sylvan Clebsch, Tobias Wrigstad
Immutable data structures are a powerful tool for building concurrent programs. They allow the sharing of data without the need for locks or other synchronisation mechanisms. This makes it much easier to reason about the correctness of the program. In this paper, we focus on what…
Kunal Sareen, Stephen M. Blackburn, Sara S. Hamouda, Lokesh Gidra
The performance of mobile devices directly affects billions of people worldwide. Yet, despite memory management being key to their responsiveness, energy efficiency, and cost, mobile devices are understudied in the literature. A paucity of suitable methodologies and benchmarks is…
Susav Shrestha, A. L. Narasimha Reddy, Zongwang Li
Recent advances in large language models have demonstrated remarkable effectiveness in information retrieval (IR) tasks. While many neural IR systems encode queries and documents into single-vector representations, multi-vector models elevate the retrieval quality by producing mu…
Xiaofan Sun, Rajiv Gupta
Concolic testing combines concrete execution with symbolic execution to automatically generate test inputs that exercise different program paths and deliver high code coverage. This approach has been extended to multithreaded programs for exposing data races. Multithreaded progra…
Jacob Bramley, Dejice Jacob, Andrei Lascu, Jeremy Singer, Laurence Tratt
Several open-source memory allocators have been ported to CHERI, a hardware capability platform. In this paper we examine the security and performance of these allocators when run under CheriBSD on Arm's prototype Morello platform. We introduce a number of security attacks and sh…
Maria Carpen-Amarie, Georgios Vavouliotis, Konstantinos Tovletoglou, Boris Grot, René Müller
The garbage collector (GC) is a crucial component of language runtimes, offering correctness guarantees and high productivity in exchange for a run-time overhead. Concurrent collectors run alongside application threads (mutators) and share CPU resources. A likely point of content…
Sayak Chakraborti, Zhizhou Zhang, Noah Bertram, Chen Ding, Sandhya Dwarkadas
Cache replacement policies typically use some form of statistics on past access behavior. As a common limitation, however, the extent of the history being recorded is limited to either just the data in cache or, more recently, a larger but still finite-length window of accesses, …
Aditya Chilukuri, Shoaib Akram
Managed search engines, such as Apache Solr and Elastic- search, host huge inverted indices in main memory to offer fast response times. This practice faces two challenges. First, limited DRAM capacity necessitates search engines aggres- sively compress indices to reduce their st…
Akshay Gopalakrishnan, Clark Verbrugge, Mark Batty
Memory consistency models traditionally specify the behavior of shared memory concurrent hardware. Hardware behavior drifts away from traditional sequential reasoning, thus exhibiting behaviors that are termed as "weak". Weaker consistency models allow for more concurrent behavio…
Brandon Kammerdiener, J. Zach McMichael, Michael R. Jantz, Kshitij A. Doshi, Terry R. Jones
Computing platforms that package multiple types of memory, each with their own performance characteristics, are quickly becoming mainstream. To operate efficiently, heterogeneous memory architectures require new data management solutions that are able to match the needs of each a…
Gurneet Kaur, Rajiv Gupta
Partitioning and processing of large graphs on a single machine with limited memory is a challenge. While many custom solutions for out-of-core processing have been developed, limited work has been done on out-of-core partitioning that can be far more memory intensive than proces…
Christos Panagiotis Lamprakos, Sotirios Xydis, Francky Catthoor, Dimitrios Soudris
Two-dimensional rectangular bin packing (2DBP) is a known abstraction of dynamic storage allocation (DSA). We argue that such abstractions can aid practical purposes. 2DBP algorithms optimize their placements’ makespan, i.e., the size of the used address range. At first glance mo…
Linsen Ma, Rui Xie, Tong Zhang
This paper studies how to mitigate the speed performance loss caused by integrating block data compression into in-memory key-value (KV) stores. Despite extensive prior research on in-memory KV stores, little focus has been given to memory usage reduction via block data compress…
Christian Navasca, Martin Maas, Petros Maniatis, Hyeontaek Lim, Guoqing Harry Xu
Memory allocators and runtime systems can leverage dynamic properties of heap allocations – such as object lifetimes, hotness or access correlations – to improve performance and resource consumption. A significant amount of work has focused on approaches that collect this informa…