I am profiling a kernel (nsight 2021.2.1, compute capability 8.3, cuda 11.4) and looking at the metric Avg thread executed for a source line. It was my understanding that this value can be between 0 and 32. However, in my profiling, it is much higher.
Clearly I have a poor understanding of the predicated-on instructions metric and therefore avg thread executed means. How should I interpret this value, and can I draw any conclusions from it?
