Parallel-RePlAce

Improvement of performance on the state of the art (at the time) VLSI placer. Analysis and reduction of memory footprint leading to 2x improvement, and then acceleration of 3 main bottlenecks with shared memory parallelism. VTune analysis of optimized code. Study of other techniques like SIMD parallelism and columnar-style data structures (and algorithms).


Written by Frédéric Gessler Deep learning, hardware accelerators and compilers enthusiast. Hacker. Builder. EPFL and Amazon alumni