How to cope with shared memory size and number of blocks in CUDA

Viewed 58

I managed to sort different arrays with CPU or GPU (without using shared memory) implementation of Odd-Even sort algorithm, but I'm having issues with using shared memory in CUDA.

This is my invocation of the kernel ->

main.cu:

//shared memory size
sizeVSM = nThreadPerBlock.x * sizeof(float);

for (int i = 0; i < array_size; i++)
    oeSortGPUSM << < nBlocks, nThreadPerBlock, sizeVSM >> > (array, i % 2, array_size);

My kernel function ->

kernel.cu:

__global__ void oeSortGPUSM(float * array, int index, int array_size) 
{
   extern __shared__ float shMem[];
   int localIndexT = threadIdx.x;
   int globalIndexT = threadIdx.x + blockIdx.x * blockDim.x;

   shMem[localIndexT] = array[globalIndexT];
   __syncthreads();

   if (index == 0 && ((localIndexT * 2 + 1) < array_size))
      //check if I have to swap elements, if so do it
      checkThenSwap(shMem, localIndexT * 2, localIndexT * 2 + 1);

}
__syncthreads();

if (index == 1 && ((localIndexT * 2 + 2) < array_size))
   //check if I have to swap elements, if so do it
   checkSwap(shMem, localIndexT * 2 + 1, localIndexT * 2 + 2, orderType);

__syncthreads();
array[globalIndexT] = shMem[localIndexT];
}

Now the issue is this: I already know that the maximum shared memory size is equal to the number of threads per block (which is 1024 for my GPU), but even if I respect this limit I'm having issues if the number of blocks is greater then 1.

Example 1: [array_size = 1024; blocks = 1; threads per block = 1024] no problems, the array is sorted.

Example 2: [array_size = 1024; blocks = 2; threads per block = 512] array is not sorted.

How can I manage all this with more than 1 block?

0 Answers
Related