In my c++ scripts, I have many for loops to compute linear algebra operations. I am wondering what is the best way to make the loops parallel? One example is the following function which computes the kronecker product of two matrices.
void Kronecker(const gsl_matrix *K, const gsl_matrix *V, gsl_matrix *H)
{
for (size_t i=0; i<K->size1; i++) {
for (size_t j=0; j<K->size2; j++) {
gsl_matrix_view H_sub=gsl_matrix_submatrix (H, i*V->size1, j*V->size2, V->size1, V->size2);
gsl_matrix_memcpy (&H_sub.matrix, V);
gsl_matrix_scale (&H_sub.matrix, gsl_matrix_get (K, i, j));
}
}
return;
}
How can I improve the computation time of my code, when I have for loops which can be parallel?