CUDA: Forgetting kernel launch configuration does not result in NVCC compiler warning or error

Viewed 870

When I try to call a CUDA kernel (a __global__ function) using a function pointer, everything appears to work just fine. However, if I forget to provide launch configuration when calling the kernel, NVCC will not result in an error or warning, but the program will compile and then crash if I attempt to run it.

__global__ void bar(float x) { printf("foo: %f\n", x); }

typedef void(*FuncPtr)(float);

void invoker(FuncPtr func)
{
    func<<<1, 1>>>(1.0);
}

invoker(bar);
cudaDeviceSynchronize();

Compile and run the above. Everything will work just fine. Then, remove the kernel's launch configuration (i.e., <<<1, 1>>>). The code will compile just fine but it will crash when you try to run it.

Any idea what is going on? Is this a bug, or I am not supposed to pass around pointers of __global__ functions?

CUDA version: 8.0

OS version: Debian (Testing repo) GPU: NVIDIA GeForce 750M

1 Answers
Related