I am trying to find some advantage of tensor cores by comparing the performance of cublasGemmEx and cublasDgemm on A100. As the doc describes, cublasGemmEx on cuda11.2+A100 support FP64, and I think cublasDgemm adapts the old algorithm(CUBLAS_GEMM_DEFAULT?), cublasGemmEx should be faster than cublasDgemm. But my experiment show that both have the same performance. does cublasDgemm already adapt tensor cores? btw, for m=n=k=5440, both have about 10 TFLOPS.