In AVX/AVX2 I could only find _mm256_stream_load_si256() , which is for __m256i. Is there no way to stream-load __m256d and why? (I would like to load it without polluting CPU cache)
Is there any obstacle for doing the following (aggressive casting)?
__m256d *pDest = /* ... */;
__m256d *pSrc = /* ... */;
/* ... */
const __m256i iWeight = _mm256_stream_load_si256(reinterpret_cast<const __m256i*>(pSrc));
const __m256d prior = _mm256_div_pd(*reinterpret_cast<const __m256d*>(&iWeight), divisor);
_mm256_stream_pd(reinterpret_cast<double*>(pDest), prior);