Is it always ok to simply use float shuffles + casts as substitute for missing integer shuffle intrinsics in SSE/AVX, like this:
__m128i x = _mm_castps_si128( _mm_shuffle_ps ( _mm_castsi128_ps(y), ...
In theory this should, of course, work with instructiuons that do not interpret the binary bit patterns of the vector elements and thus agnostic wrt. they contain floats or integers. However, I remember one post by (IIRC) @Peter Cordes who wrote that using float shuffles for integer registers works on "some" CPUs only.