14,842 papers · page 14 of 743
Matthias Hetzenberger, Georg Moser, Florian Zuleger
Probabilistic algorithms and data structures are widely used to obtain favourable expected performance guarantees. While their mathematical analysis is often well understood, mechanising expected-cost analyses remains challenging, requiring reasoning about probability distributio…
Alex Hobbs, Alex Dixon
Functional programming is a core part of many undergraduate-level courses in Computer Science. Pedagogical interest in Haskell continues to grow as functional idioms become commonplace in popular multi-paradigm languages. The Glasgow Haskell Compiler (GHC) is by far Haskell’s mos…
Alperen Keles, George Miao, Leonidas Lampropoulos
Property-based testing frameworks rely on shrinking to turn noisy random failures into counterexamples that developers can debug. Although bug-finding performance is routinely measured, shrinking itself is rarely evaluated quantitatively. We present an experience report on evalua…
Robert Krook, Lennart Augustsson
Cloud Haskell brings Erlang-style distributed programming to Haskell, but its treatment of mobile code exposes a difficult boundary in the source-level API. Remote processes must be expressed as static closures, messages must satisfy serialisation constraints, and participating n…
Stephanie Weirich
Dependent type theory is having a moment as the foundation for interactive provers, such as Lean, Rocq, and Agda. But what does dependent type theory offer to programmers, who just want to get work done? While Haskell is not a full spectrum dependently-typed language, its type sy…
Nadav Amit
While huge pages can dramatically reduce address translation overhead, their use for executable code remains limited by structural barriers in binary formats, page cache management, and loader mechanisms. Existing solutions copy code into anonymous memory, abandoning file-backed …
Anthony Arnold, Mark Marron
Garbage Collectors (GCs) are a critical component of a modern application stack. Long pauses, large memory consumption, and high CPU usage can unexpectedly occur with certain workloads or series of events. These behaviors can make systems unresponsive, make it impossible to run t…
Soham Bagchi, Sanya Srivastava, Reese Levine, Tyler Sorensen, Ryan Stutsman, Vijay Nagarajan
Modern heterogeneous processors like the NVIDIA Grace-Hopper Superchip tightly integrate CPU and GPU cores across a cache-coherent interconnect, with an implicit assumption that independently compiled CPU and GPU code can safely interact via shared memory. Yet the memory consiste…
Edoardo D'Alessio, Mohamed Husain Noor Mohamed, Xiaoguang Wang, Binoy Ravindran
Recent Linux memory-management interfaces make it practical to revisit distributed shared memory (DSM) as a deployable runtime substrate for conventional multithreaded software. We present Stretch, a userspace fault-driven page-granularity DSM runtime that combines userfaultfd-ba…
Kai Feng, Huanting Wang, Jeremy Singer, Zheng Wang
Memory safety is a critical issue in embedded systems. Although high-level languages like MicroPython simplify IoT development, their C-based runtimes remain vulnerable to memory errors triggered by Python code or native extensions. The CHERI (Capability Hardware Enhanced RISC In…
Nicolas van Kempen, Emery D. Berger
Programmers using native languages such as C, C++, or Rust can implement custom memory allocation strategies to improve execution time. In their paper titled "Reconsidering Custom Memory Allocation" almost 25 years ago, Berger et al. showed that while per-class allocators provide…
Hayley Patton, Stephen M. Blackburn
Offset-Vector Compaction (OVC) algorithms, such as the Compressor, are among the most widely deployed garbage collectors today, yet they have remained largely unexplored by the literature since the earliest algorithms were published two decades ago. Although an implementation in …
Yunqi Shen, Dimitrios Nikolopoulos
GPU unified memory simplifies programming by automatically migrating pages between CPUs and GPUs, but page faults trigger migrations with hundreds of microseconds to millisecond-scale latency, stalling thousands of threads. We target this bottleneck with a page prefetching framew…
Bijan Tabatabai, Eishan Mirakhur, Ravi Shankar Jonnalagadda, Vinicius Petrucci, Rohit Sehgal, Jus Singh, Michael M. Swift
CXL memory devices increase the memory capacity and bandwidth available to a server, at the cost of higher access latency. Prior research focused on how to make use of the expanded memory capacity provided by CXL while minimizing the impact of its higher access latency. However, …
Yuchen Ma, Bin Ren, Andreas Stathopoulos
Distributed matrix-block-vector multiplication (Matvec) algorithm is a critical component of many applications, but can be computationally challenging for dense matrices of dimension O(10^6–10^7) and blocks of O(10–100) vectors. We present performance analysis, implementation, an…
Bing Lu, Zedong Liu, Hairui Zhao, Dejun Luo, Wenjing Huang, Yida Gu, Jinyang Liu, Guangming Tan + 1 more
With the exponential growth of computing power, large-scale scientific simulations are producing massive volumes of data, leading to critical storage and I/O challenges. Error-bounded lossy compression has become one of the most effective solutions for reducing data size while pr…
Ajay Singh, Nikos Metaxakis, Panagiota Fatourou
We present a new blocking linearizable stack implementation which utilizes sharding and fetch&increment to achieve significantly better performance than all existing concurrent stacks. The proposed implementation is based on a novel elimination mechanism and a new combining appro…
Yida Li, Siwei Zhang, Yiduo Niu, Yang Du, Qingxiao Sun, Zhou Jin, Weifeng Liu
Sparse direct solvers are critical building blocks in a range of scientific applications on heterogeneous supercomputers. However, existing sparse direct solvers have not been able to well leverage the high bandwidth and floating-point performance of modern GPUs. The primary chal…
Zhiyuan Zhang, Yanxin Cai, Wenhao Yin, Xueyu Wu, Yi Wang, Lei Ju, Zhuoran Ji
Zero-knowledge proofs (ZKPs) are cryptographic protocols that allow verification of statements without disclosing the underlying information. Among them, PLONK-based ZKPs are particularly notable for offering succinct, non-interactive proofs of knowledge with a universal trusted …
Junyao Zhang, Zhuo Wang, Zhe Zhou
In high-performance applications, critical sections often become performance bottlenecks due to contention among multiple cores. Critical section delegation mitigates this overhead by consistently executing critical sections on the same core, thereby reducing contention. Traditio…