I managed to sort different arrays with CPU or GPU (without using shared memory) implementation of Odd-Even sort algorithm, but I'm having issues with using shared memory in CUDA.
This is my invocation of the kernel ->
main.cu:
//shared memory size
sizeVSM = nThreadPerBlock.x * sizeof(float);
for (int i = 0; i < array_size; i++)
oeSortGPUSM << < nBlocks, nThreadPerBlock, sizeVSM >> > (array, i % 2, array_size);
My kernel function ->
kernel.cu:
__global__ void oeSortGPUSM(float * array, int index, int array_size)
{
extern __shared__ float shMem[];
int localIndexT = threadIdx.x;
int globalIndexT = threadIdx.x + blockIdx.x * blockDim.x;
shMem[localIndexT] = array[globalIndexT];
__syncthreads();
if (index == 0 && ((localIndexT * 2 + 1) < array_size))
//check if I have to swap elements, if so do it
checkThenSwap(shMem, localIndexT * 2, localIndexT * 2 + 1);
}
__syncthreads();
if (index == 1 && ((localIndexT * 2 + 2) < array_size))
//check if I have to swap elements, if so do it
checkSwap(shMem, localIndexT * 2 + 1, localIndexT * 2 + 2, orderType);
__syncthreads();
array[globalIndexT] = shMem[localIndexT];
}
Now the issue is this: I already know that the maximum shared memory size is equal to the number of threads per block (which is 1024 for my GPU), but even if I respect this limit I'm having issues if the number of blocks is greater then 1.
Example 1: [array_size = 1024; blocks = 1; threads per block = 1024] no problems, the array is sorted.
Example 2: [array_size = 1024; blocks = 2; threads per block = 512] array is not sorted.
How can I manage all this with more than 1 block?