PPoPP 2017
47 papers
- A Multicore Path to Connectomics-on-Demand
- An Efficient Abortable-locking Protocol for Multi-level NUMA Systems
- Checking Concurrent Data Structures Under the C/C++11 Memory Model
- Combining SIMD and Many/Multi-core Parallelism for Finite State Machines with Enumerative Speculation
- Contention in Structured Concurrency: Provably Efficient Dynamic Non-Zero Indicators for Nested Parallelism
- EffiSha: A Software Framework for Enabling Effficient Preemptive Scheduling of GPU
- Eunomia: Scaling Concurrent Search Trees under Contention Using HTM
- Exploiting Vector and Multicore Parallelism for Recursive, Data- and Task-Parallel Programs
- Function Call Re-Vectorization
- Grammar-aware Parallelization for Scalable XPath Querying
- Groute: An Asynchronous Multi-GPU Programming Model for Irregular Computations
- Isoefficiency in Practice: Configuring and Understanding the Performance of Task-based Applications
- It's Time for a New Old Language
- KiWi: A Key-Value Map for Scalable Real-Time Analytics
- Layout Lock: A Scalable Locking Paradigm for Concurrent Data Layout Modifications
- Model-based Iterative CT Image Reconstruction on GPUs
- Noise Injection Techniques to Expose Subtle and Unintended Message Races
- Optimizing the Four-Index Integral Transform Using Data Movement Lower Bounds Analysis
- POSTER: A GPU-Friendly Skiplist Algorithm
- POSTER: A Wait-Free Queue with Wait-Free Memory Reclamation
- POSTER: An Architecture and Programming Model for Accelerating Parallel Commutative Computations via Privatization
- POSTER: An Infrastructure for HPC Knowledge Sharing and Reuse
- POSTER: Automated Load Balancer Selection Based on Application Characteristics
- POSTER: Cache-Oblivious MPI All-to-All Communications on Many-Core Architectures
- POSTER: Distributed Control: The Benefits of Eliminating Global Synchronization via Effective Scheduling
- POSTER: HythTM: Extending the Applicability of Intel TSX Hardware Transactional Support
- POSTER: IOGP: An Incremental Online Graph Partitioning for Large-Scale Distributed Graph Databases
- POSTER: MAPA: An Automatic Memory Access Pattern Analyzer for GPU Applications
- POSTER: On the Problem of Consistency Exceptions in the Context of Strong Memory Models
- POSTER: Poor Man's URCU
- POSTER: Provably Efficient Scheduling of Cache-Oblivious Wavefront Algorithms
- POSTER: Recovering Performance for Vector-based Machine Learning on Managed Runtime
- POSTER: Reuse, don't Recycle: Transforming Algorithms that Throw Away Descriptors
- POSTER: STAR (Space-Time Adaptive and Reductive) Algorithms for Real-World Space-Time Optimality
- POSTER: State Teleportation via Hardware Transactional Memory
- Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks
- Processor-Oblivious Record and Replay
- S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters
- SC-Haskell: Sequential Consistency in Languages That Minimize Mutable Shared Heap
- Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL
- Silent Data Corruption Resilient Two-sided Matrix Factorizations
- Simple, Accurate, Analytical Time Modeling and Optimal Tile Size Selection for GPGPU Stencils
- Synchronized-by-Default Concurrency for Shared-Memory Systems
- Tapir: Embedding Fork-Join Parallelism into LLVM's Intermediate Representation
- Thread Data Sharing in Cache: Theory and Measurement
- Understanding the GPU Microarchitecture to Achieve Bare-Metal Performance Tuning
- Using Butterfly-Patterned Partial Sums to Draw from Discrete Distributions