GPU Padding Is Not One Thing
Padding on a GPU is easy to misunderstand because the word is used for several different mechanisms. In CPU code, I usually think of padding as adding unused bytes to a struct or array for alignment. In GPU programming, the same word can mean alignment padding, tensor padding, convolution boundary padding, shared-memory padding, row-stride padding, warp-tail masking, or library-internal layout padding.
Those cases have different semantics. Some padding contains zeros. Some padding is uninitialized memory. Some padding is never read. Some padding is read intentionally to simplify a kernel. Treating all of them as "the GPU adds zeros" is the source of many bugs.
A better mental model is this:
Padding is the difference between the logical domain of the computation and the physical domain used to execute it efficiently.
The logical domain is what the physics, model, image, matrix, or histogram means. The physical domain is what the GPU kernel actually touches in memory or maps onto threads.
CUDA does not automatically add zeros
CUDA does not generally pad your data for you. If you allocate
float * x;
cudaMalloc ( & x, n * sizeof ( float ) );
you get space for exactly n floats. CUDA does not add extra usable elements at the end. If a kernel reads x[n], that is an out-of-bounds read, even if the address happens to fall inside a larger allocation nearby.
Zero-padding is explicit. You either write the zeros yourself, use a library mode that defines zero-padding, or rely on a kernel branch that returns zero for out-of-domain indices.
For example, this is logical zero-padding:
device float load_with_zero_padding (
const float * x,
int i,
int n
) {
return ( i > = 0 & & i < n )? x [ i ]: 0.0f;
}
No extra memory is needed. The padding exists in the indexing rule.
This version uses physical zero-padding:
cudaMemset ( x_padded, 0, padded_n * sizeof ( float ) );
cudaMemcpy ( x_padded, x, n * sizeof ( float ), cudaMemcpyDeviceToDevice );
Here the extra elements e…