In CUDA, each thread knows its block index in the grid and thread index within the block. But two important values do not seem to be explicitly available to it:
- Its index as a lane within its warp (its "lane id")
- The index of the warp of which it is a lane within the block (its "warp id")
Assuming the grid is 1-dimensional(a.k.a. linear, i.e. blockDim.y and blockDim.z are 1), one can obviously obtain these as follows:
enum : unsigned { warp_size = 32 };
auto lane_id = threadIdx.x % warp_size;
auto warp_id = threadIdx.x / warp_size;
and if you don't trust the compiler to optimize that, you could rewrite it as:
enum : unsigned { warp_size = 32, log_warp_size = 5 };
auto lane_id = threadIdx.x & (warp_size - 1);
auto warp_id = threadIdx.x >> log_warp_size;
is that the most efficient thing to do? It still seems like a lot of waste for every thread to have to compute this.
(inspired by this question.)