The "cooperative groups" mechanism has appeared in recent versions of CUDA. Some of it involves actual hardware features which are less obvious (?) to utilize otherwise; but a lot of it is basically just library code; and it's difficult to discern where the hardware actually assists with some special functionality.
This question is about GPUs with Compute Capability 7.x. In the CUDA programming guide, I notice the following features dependencies on compute capabilities:
| Cooperative group feature | Required minimum Compute Capability |
|---|---|
labeled_partition() |
7.0 |
binary_partition() |
7.0 |
async_memcpy() actually being asynchronous |
8.0 |
Some kind of asynchronicity of wait() |
8.0 |
"acceleration" of reduce() |
8.0 |
| use of intrinsics in reductions for: plus, less, greater, bitwise and, bitwise or, bitwise_xor | 8.0 |
What hardware features, specifically, were introduced with CC 7.0 and with CC 8.0 which enable this functionality? What are their exact semantics? And are they all explicitly exposed via PTX, or are some of them only visible in SASS?