GPUs, when used for general-purpose computing, put a lot of emphasis on fine-grained parallelism with SIMD and SIMT. They perform best on regular numbercrunching workloads with high arithmetic intensity.
Nonetheless, to be applicable to as many workloads as they have been applied to, they must also be capable of coarse-grained MIMD parallelism, where different cores execute different instruction streams on different chunks of data.
This means different cores on the GPU must synchronize with each other after executing different instruction streams. How do they do it?
On a CPU the answer would be that there is cache coherence plus a set of communication primitives chosen to work well with that such as CAS or LL/SC. But as I understand it, GPUs do not have cache coherence - avoiding the overhead of such is the biggest reason they are more efficient than CPUs in the first place.
So what method do GPU cores use for synchronizing with each other? If the answer to how they exchange data is by writing to shared main memory, then how do they synchronize so the sender can inform the recipient when to read the data?
If the answer depends on the particular architecture, then I'm particularly interested in modern Nvidia GPUs that support CUDA.
Edit: From the document Booo linked, here is my understanding so far:
They seem to use the word 'stream' for a quantity of stuff that gets done synchronously (including fine-grained parallelism like SIMD); the problem is then how to synchronize/communicate between multiple streams.
As I surmised, this is much more explicit than it is on CPUs. in particular, they talk about:
- Page-locked memory
- cudaDeviceSynchronize ()
- cudaStreamSynchronize ( streamid )
- cudaEventSynchronize ( event )
So streams can communicate by writing data to main memory (or L3 cache?) and there is nothing like the cache coherence there is on CPUs, instead there is locking pages of memory, and/or an explicit synchronization API.