CGO 2017
26 papers
- A collaborative dependence analysis framework
- A space- and energy-efficient code Compression/Decompression technique for coarse-grained reconfigurable architectures
- Automatic detection of extended data-race-free regions
- Automatic generation of fast BLAS3-GEMM: a portable compiler approach
- Characterizing data organization effects on heterogeneous memory architectures
- Clairvoyance: look-ahead compile-time scheduling
- Cross-ISA machine emulation for multicores
- Discovery and exploitation of general reductions: a constraint based approach
- Dynamic buffer overflow detection for GPGPUs
- FinePar: irregularity-aware fine-grained workload partitioning on integrated architectures
- Formalizing the concurrency semantics of an LLVM fragment
- Incremental whole program optimization and compilation
- Legato: end-to-end bounded region serializability using commodity hardware transactional memory
- Lift: a functional data-parallel IR for high-performance GPU code generation
- Minimizing the cost of iterative compilation with active learning
- Optimistic loop optimization
- Optimizing function placement for large-scale data-center applications
- Parallel associative reductions in halide
- Phase-aware optimization in approximate computing
- Pointer disambiguation via strict inequalities
- Removing checks in dynamically typed languages through efficient profiling
- Software prefetching for indirect memory accesses
- Synthesizing benchmarks for predictive modeling
- Taming warp divergence
- ThinLTO: scalable and incremental LTO
- TwinKernels: an execution model to improve GPU hardware scheduling at compile time