My first CUDA kernel ran slower than a single-threaded CPU loop.
Not slightly slower. Noticeably slower. I had thousands of cores available, and I was getting beaten by one. That was my first real CUDA lesson: GPUs are not magic fast machines. They are throughput machines. If you do not give them enough work, you are not accelerating anything. You are just paying overhead. The Ferrari and the Cargo Train A CPU core is a Ferrari. One passenger. Very fast trip. Great handling. Smart suspension. Branch prediction, out-of-order execution, deep caches. Everything is built to
评论
?
参与讨论