CGO 2015
24 papers
- A graph-based higher-order intermediate representation
- A parallel abstract interpreter for JavaScript
- Approximating flow-sensitive pointer analysis using frequent itemset mining
- Automatic data placement into GPU on-chip memory resources
- Branch prediction and the performance of interpreters: don't trust folklore
- Characterizing and enhancing global memory data coalescing on GPUs
- Checking correctness of code generator architecture specifications
- Data provenance tracking for concurrent programs
- EMEURO: a framework for generating multi-purpose accelerators via deep learning
- Getting in control of your control flow with control-data isolation
- HELIX-UP: relaxing program semantics to unleash parallelization
- HERMES: a fast cross-ISA binary translator with post-optimization
- Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS
- Locality aware concurrent start for stencil applications
- Locality-centric thread scheduling for bulk-synchronous programming models on CPU architectures
- MemorySanitizer: fast detector of uninitialized memory use in C++
- On performance debugging of unnecessary lock contentions on multicore processors: a replay-based approach
- Optimizing and auto-tuning scale-free sparse matrix-vector multiplication on Intel Xeon Phi
- Optimizing binary translation of dynamically generated code
- Optimizing the flash-RAM energy trade-off in deeply embedded systems
- PSLP: padded SLP automatic vectorization
- Reactive tiling
- Scalable conditional induction variables (CIV) analysis
- Snapshot-based loading-time acceleration for web applications