task scheduling of NVIDIA GPU

Viewed 2087

I have some doubt about the task scheduling of nvidia GPU.

(1) If a warp of threads in a block(CTA) have finished but there remains other warps running, will this warp wait the others to finish? In other words, all threads in a block(CTA) release their resource when all threads are all finished, is it ok? I think this point should be right,since threads in a block share the shared memory and other resource, these resource allocated in a CTA size manager.

(2) If all threads in a block(CTA) hang-up for some long latency such as global memory access? will a new CTA threads occupy the resource which method like CPU? In other words, if a block(CTA) has been dispatched to a SM(Streaming Processors), if it will take up the resource until it has finished?

I would be appreciate if someone recommend me some book or articles about the architecture of GPU.Thanks!

2 Answers
Related