void reserve1(vector<vector<int>>& vec, size_t siz) {
for (auto& v : vec) {
v.reserve(siz);
}
}
void reserve2(vector<vector<int>>& vec, size_t siz) {
#pragma omp parallel default(shared) num_threads(vec.size())
{
vec[omp_get_thread_num()].reserve(vec);
}
}
void doSomething(vector<int>& v) {
// Do something write (and read) heavy on v.
}
int main() {
vector<vector<int>> data(omp_get_num_threads());
size_t theSize = 1000000;
reserve1(data, theSize); OR reserve2(data, theSize); <===================================
// Do some stuff here (not parallel)
#pragma omp parallel default(shared) num_threads(vec.size())
{
doSomething(data[omp_get_thread_num()]);
}
return 0;
}
In the code above is it better to use reserve2 rather than reserve1? Here are some pros and cons I could think of (I am not sure if they are correct):
Pros:
1- Allocation is done in parallel (and with appropriate version of malloc it might be faster)
2- Each thread is working with data allocated by itself which would have faster access (I have heard this from a colleague, but have no idea if that is correct).
Con:
1- Creating a parallel region has an overhead that might affect the performance.
Here are my questions and I would appreciate your help:
a) Are these pros and con correct or incorrect?
b) Are there other benefits/disadvantages in using parallel vs serial allocation here?
c) Do "number of threads" or "theSize" affect the pros/cons?
Edit:
Some clarification:
1- data always has the size equal to the number of threads that is given.
2- the size of data[i] vectors is not known in advance, but there is an estimate (the upper-bound is very large (almost nthreads*estimate) and the lower-bound is 0.)
3- doSomething is the method resizing data[i] (using e.g. push_back). It runs in parallel and the elements it collects is coming from a code I cannot change much.
4- In reality, data is a vector<vector>, where someObject is a small struct with multiple fields.
5- I can use something other than a vector, but I should be able to add elements one by one, and I do not know the exact final size, as mentioned before.
6- The code in main() is going to run thousands of times.
The main question for me is this: If thread i is working (reading/writing) on data[i] and data[i] only, would it affect the performance if data[i] is allocated by thread "i" or if it is allocated by a different thread (e.g. main thread in a non-parallel region)?