below is an implementation of a matrix multiply in AVX2. The machine I am using only supports AVX so I am trying to implement the same configuration with AVX.
However, I am having trouble deciphering really what the differences are, and what would needed to be changed! What in this implementation is specific to AVX2 that would not work with a machine only able to process AVX?
This is a link to all the commands for AVX as well as AVX2 https://software.intel.com/sites/landingpage/IntrinsicsGuide/#techs=AVX
Thank you for any insight at all!
for (uint64_t i = 0; i < M; i++)
{
for (uint64_t j = 0; j < N; j++)
{
__m256 X = _mm256_setzero_ps();
for (uint64_t k = 0; k < L; k+= 8) {
const __m256 AV = _mm256_load_ps(A+i*L+k);
const __m256 BV = _mm256_load_ps(B+j*L+k);
X = _mm256_fmadd_ps(AV,BV,X);
}
C[i*N+j] = hsum_avx(X);
}
}